<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/scaling-token-factory-revenue-and-ai-efficiency-by-maximizing-performance-per-watt-2qzgn7uql" -->

---
title: Scaling Token Factory Revenue and AI Efficiency by...
description: NVIDIA outlines how its GPU architectures and AI factory software maximize performance per watt to increase token throughput and revenue within fixed power...
canonical: https://daily.dev/posts/scaling-token-factory-revenue-and-ai-efficiency-by-maximizing-performance-per-watt-2qzgn7uql
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: Scaling Token Factory Revenue and AI Efficiency by Maximizing Performance per Watt | daily.dev
og:description: NVIDIA outlines how its GPU architectures and AI factory software maximize performance per watt to increase token throughput and revenue within fixed power...
og:url: https://daily.dev/posts/scaling-token-factory-revenue-and-ai-efficiency-by-maximizing-performance-per-watt-2qzgn7uql
og:image: https://api.daily.dev/og/posts/2qZgN7UQL.png
og:image:alt: Scaling Token Factory Revenue and AI Efficiency by Maximizing Performance per Watt
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Scaling Token Factory Revenue and AI Efficiency by Maximizing Performance per Watt

**[NVIDIA Developer](https://daily.dev/sources/nvidiadev)** · 9 min read · 0 upvotes · 0 comments

## Summary

NVIDIA outlines how its GPU architectures and AI factory software maximize performance per watt to increase token throughput and revenue within fixed power envelopes. Across six architecture generations, NVIDIA claims a 1,000,000x improvement in inference throughput per megawatt. Key highlights include: Blackwell Ultra GB300 NVL72 delivering up to 50x higher throughput per megawatt and 35x lower token cost versus Hopper for DeepSeek-R1; the upcoming Vera Rubin platform achieving up to 10x higher inference throughput per megawatt versus Blackwell; 100% liquid cooling enabling 1.1 PUE; and the NVIDIA DSX system enabling AI factories to operate up to 30% more GPUs within the same power envelope. The post also covers efficiency gains in chip manufacturing via cuLitho and cuEST, GPU-accelerated EDA, and how operators should track revenue per megawatt as the core business metric.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://developer.nvidia.com/blog/scaling-token-factory-revenue-and-ai-efficiency-by-maximizing-performance-per-watt/>

## Similar posts on daily.dev

- [New SemiAnalysis InferenceX Data Shows NVIDIA Blackwell Ultra Delivers up to 50x Better Performance and 35x Lower Costs for Agentic AI](https://daily.dev/posts/new-semianalysis-inferencex-data-shows-nvidia-blackwell-ultra-delivers-up-to-50x-better-performance--xlkxksrjj) · NVIDIA · 0 upvotes · 0 comments
- [NVIDIA Vera Rubin and Blackwell Set a New Standard for Agentic AI Performance per Watt](https://daily.dev/posts/nvidia-vera-rubin-and-blackwell-set-a-new-standard-for-agentic-ai-performance-per-watt-r1iqkqtkn) · NVIDIA Developer · 0 upvotes · 0 comments
- [Maximize AI Factory Energy Efficiency Through Full-Stack Inference and Training Optimizations](https://daily.dev/posts/maximize-ai-factory-energy-efficiency-through-full-stack-inference-and-training-optimizations-lwusm4jtd) · NVIDIA Developer · 0 upvotes · 0 comments

---

Tags: [#nvidia](https://daily.dev/tags/nvidia), [#ai-infrastructure](https://daily.dev/tags/ai-infrastructure)

[View this post on daily.dev](https://daily.dev/posts/scaling-token-factory-revenue-and-ai-efficiency-by-maximizing-performance-per-watt-2qzgn7uql)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"Scaling Token Factory Revenue and AI Efficiency by Maximizing Performance per Watt","url":"https://daily.dev/posts/scaling-token-factory-revenue-and-ai-efficiency-by-maximizing-performance-per-watt-2qzgn7uql","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/scaling-token-factory-revenue-and-ai-efficiency-by-maximizing-performance-per-watt-2qzgn7uql"},"datePublished":"2026-03-25T11:05:27.732Z","dateModified":"2026-03-25T11:05:54.786Z","description":"NVIDIA outlines how its GPU architectures and AI factory software maximize performance per watt to increase token throughput and revenue within fixed power...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/d801fbaac9eb3cde944d4284954d47df?_a=AQAEuop","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/d801fbaac9eb3cde944d4284954d47df?_a=AQAEuop","isAccessibleForFree":true,"articleSection":"NVIDIA Developer","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"NVIDIA Developer","logo":"https://media.daily.dev/image/upload/t_logo,f_auto/v1/logos/86e45aab42ba48ce83103d01b1119910","url":"https://daily.dev/sources/nvidiadev"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/scaling-token-factory-revenue-and-ai-efficiency-by-maximizing-performance-per-watt-2qzgn7uql","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":0},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"nvidia,ai-infrastructure","timeRequired":"PT9M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"NVIDIA Developer","item":"https://daily.dev/sources/nvidiadev"},{"@type":"ListItem","position":3,"name":"Scaling Token Factory Revenue and AI Efficiency by Maximizing Performance per Watt"}]}
```

