AI Infrastructure Analysis
How Many AI Tokens Can 1 kWh Produce? Cost and Revenue Analysis
Updated August 2026 · Approximately 15 minutes to read
As generative AI becomes more widely adopted, a new “token economy” is beginning to emerge. AI tokens are not cryptocurrencies. They are the basic units models use to process and generate text, code, images, and other forms of content—and they are also a common unit for pricing AI services. NVIDIA founder and CEO Jensen Huang describes the next generation of data centers as “AI factories”: electricity and data enter the factory, GPUs run AI models, and the resulting tokens are assembled into AI products and services. He has highlighted tokens per watt as an important measure of AI infrastructure efficiency and revenue potential.
If tokens are the products manufactured by AI factories, how many AI tokens can one kilowatt-hour of electricity generate—and how much revenue could that output represent?
It sounds like a simple calculation. In reality, the answer depends on GPU performance, model architecture, inference precision, server utilization, data center PUE, input-output ratios, cache hit rates, API prices, and local electricity costs.
This article uses a simplified model to estimate:
- How many tokens 1 kWh could theoretically generate;
- How much production may remain after real-world engineering losses;
- How much those tokens could generate in China and the United States;
- How much gross contribution remains after electricity costs;
- How much of an advantage low-cost electricity can provide.
In this analysis
- Model assumptions
- Theoretical token output
- Engineering scenario
- China power-cost scenario
- U.S. power-cost scenario
- China–U.S. comparison
- Limits of the model
- Frequently asked questions
1. Model and Infrastructure Assumptions
This calculation uses the following assumptions:
- Model: DeepSeek V4 Flash
- Active model parameters: approximately 13 billion
- Inference precision: FP4
- Server: an HGX B300 system with eight NVIDIA B300 GPUs
- Theoretical FP4 peak performance: 144 PFLOPS
- Assumed total server power consumption: approximately 14.5 kW
- Data center PUE: 1.15
- Average effective GPU utilization: 35%
- Input tokens: 80% of total tokens
- Output tokens: 20% of total tokens
- Input-token cache hit rate: 80%
According to NVIDIA’s published specifications, an eight-GPU HGX B300 system can deliver up to 144 PFLOPS of FP4 peak performance.
However, that number represents theoretical peak performance under specific conditions. It does not mean the server can sustain the same level of useful performance in a real AI inference workload.
2. How Long Can 1 kWh Power the Server?
One kilowatt-hour, or 1 kWh, is the amount of energy used by a 1 kW device operating for one hour.
If the AI server consumes approximately 14.5 kW, then 1 kWh can power it for:
Converted into seconds:
In other words, 1 kWh can power a 14.5 kW AI server at full load for only about 4.1 minutes.
3. How Many AI Tokens Could 1 kWh Generate in Theory?
Assume the server delivers its full 144 PFLOPS and that generating one token requires floating-point operations equal to roughly twice the model’s active parameter count.
The theoretical calculation is:
or 1.375 Billion Tokens
This is a highly idealized mathematical ceiling.
It assumes that the GPUs continuously operate at peak performance while ignoring:
- GPU memory bandwidth;
- Multi-GPU communication overhead;
- KV cache reads and writes;
- Request scheduling delays;
- Variations in input and output length;
- Batching efficiency;
- Sampling and decoding overhead;
- Server idle time;
- Inference framework efficiency.
A real AI server is extremely unlikely to sustain this theoretical result. The 1.375 billion figure should therefore be treated as an upper bound, not as a production forecast.
4. How Much Revenue Could the Theoretical Output Generate?
DeepSeek V4 Flash currently lists the following API prices:
| Token category | Price per 1M tokens |
|---|---|
| Cached input | $0.0028 |
| Uncached input | $0.14 |
| Output | $0.28 |
DeepSeek charges different rates for input tokens, output tokens, and cached input tokens. Context caching is enabled by default, but cache hits depend on factors such as whether requests contain qualifying identical prefixes. A cache hit is not guaranteed for every input token.
Assume that the 1.375 billion tokens consist of:
- 80% input tokens;
- 20% output tokens;
- An 80% cache hit rate for input tokens;
- A 20% cache miss rate for input tokens.
| Token category | Token volume | Revenue |
|---|---|---|
| Cached input | 880M | $2.46 |
| Uncached input | 220M | $30.80 |
| Output | 275M | $77.00 |
| Total | 1,375M | $110.26 |
That result sounds extraordinary. But it is not a realistic profit estimate, nor does it mean an operator can consistently turn a few cents of electricity into more than $100.
5. Estimated AI Token Production in an Engineering Scenario
There is usually a large gap between theoretical peak performance and useful inference output.
To create a more realistic operating scenario, this model applies two adjustments:
- Average effective GPU utilization of 35%;
- Data center PUE of 1.15.
PUE, or power usage effectiveness, measures total data center energy consumption relative to the energy consumed by IT equipment.
A PUE of 1.15 means that in addition to the electricity used by the servers, the facility consumes extra energy for cooling, power conversion, and other supporting infrastructure.
418.5 Million Tokens per kWh
That is about 70% below the theoretical ceiling. It is closer to an engineering-style estimate, but it is still not a benchmark result. Actual throughput depends on model implementation, inference software, request structure, context length, and batch size.
6. China Low-Cost Electricity Scenario: AI Token Revenue per kWh
The original Chinese model estimated that 418.5 million tokens would generate approximately RMB 240 in revenue under its assumed domestic token prices.
To make the China-U.S. comparison easier to understand, all revenue and electricity figures are converted into U.S. dollars.
Using a reference exchange rate of approximately RMB 6.74 per U.S. dollar:
approximately $35.61 per kWh
Electricity Cost in a Low-Cost Chinese Power Region
China does not have a single national electricity rate for data centers. Actual costs vary according to province, voltage level, electricity market arrangements, time-of-use pricing, transmission and distribution charges, and long-term power purchase agreements.
Inner Mongolia provides one example of a region with abundant energy resources and a growing computing industry.
According to the Inner Mongolia Energy Administration, registered multiyear power purchase agreements involving local big-data companies had an average transaction price of:
= RMB 0.2196 per kWh
≈ $0.033 per kWh
After subtracting this direct electricity cost:
approximately $35.58 per kWh after electricity
7. United States Electricity Scenario: AI Token Revenue per kWh
The U.S. scenario uses the same engineering output of 418.5 million tokens per kWh, but revenue is calculated using DeepSeek’s published U.S. dollar API prices.
Cached Input Tokens
267.84 × $0.0028 = $0.75
Uncached Input Tokens
66.96 × $0.14 = $9.37
Output Tokens
83.7 × $0.28 = $23.44
| Token category | Token volume | Revenue |
|---|---|---|
| Cached input | 267.84M | $0.75 |
| Uncached input | 66.96M | $9.37 |
| Output | 83.70M | $23.44 |
| Total | 418.50M | $33.56 |
approximately $33.56 per kWh
8. How Much Remains After U.S. Electricity Costs?
According to the U.S. Energy Information Administration, average U.S. retail electricity prices in 2025 were:
| Customer category | Average 2025 electricity price |
|---|---|
| Industrial | $0.0862/kWh |
| Commercial | $0.1341/kWh |
| All sectors | $0.1363/kWh |
These figures represent nationwide average retail prices charged to end users. Actual data center electricity costs can vary significantly depending on state, load size, voltage level, power purchase agreements, demand charges, and taxes.
Using the average U.S. industrial electricity price:
approximately $33.47 per kWh after electricity
If the facility paid the national average commercial rate instead:
9. China vs. the United States: AI Token Revenue Comparison
| Metric | China low-cost power scenario | U.S. industrial-rate scenario |
|---|---|---|
| Token output per kWh | 418.5M | 418.5M |
| Token revenue per kWh | $35.61 | $33.56 |
| Direct electricity cost | $0.033 | $0.086 |
| Gross contribution after electricity | $35.58 | $33.47 |
| Difference versus U.S. scenario | +$2.11 | — |
United States: $33.47/kWh after electricity
Using each market’s respective token-pricing assumptions and reference electricity costs, the China scenario retains approximately $2.11 more per kWh.
However, the entire $2.11 difference does not come from cheaper electricity.
10. Where Does the $2.11 Difference Come From?
The reference electricity costs are:
- China low-cost power example: approximately $0.033/kWh;
- U.S. average industrial rate: approximately $0.086/kWh.
Of the total $2.11 difference:
- Approximately $0.053 comes directly from the lower reference electricity price;
- Approximately $2.06 comes from differences in token pricing, exchange-rate conversion, and rounding in the original Chinese model.
Using each market’s respective token prices, gross contribution after electricity is approximately $35.58 per kWh in the China low-cost power scenario and $33.47 in the U.S. scenario. The total difference is about $2.11, of which approximately $0.053 comes directly from the reference electricity-cost difference.
11. What If We Compare Electricity Costs Alone?
To isolate the effect of electricity prices, assume both locations generate exactly the same token revenue:
| Location | Token revenue | Direct electricity cost | Gross contribution |
|---|---|---|---|
| China low-cost power example | $33.56 | $0.033 | $33.53 |
| U.S. average industrial rate | $33.56 | $0.086 | $33.47 |
With identical token revenue, the China low-cost power example retains approximately:
Assume a data center maintains an average facility load of 100 MW and operates throughout the year:
At a savings of approximately $0.053 per kWh, the theoretical annual energy-cost advantage would be:
For a large AI data center, a difference of only a few cents per kilowatt-hour can become tens of millions of dollars per year.
12. Why AI Token Revenue per kWh Is Not Net Profit
The $35.58 China figure and the $33.47 U.S. figure represent token revenue minus direct electricity cost only.
They do not include:
- GPU server acquisition;
- Hardware depreciation;
- Financing costs;
- Data center construction or colocation fees;
- Liquid cooling and heat rejection;
- High-speed networking and storage;
- CPUs, memory, and supporting hardware;
- Inference software development and optimization;
- Operations and technical support;
- Data security and regulatory compliance;
- Marketing and customer acquisition;
- Taxes and insurance;
- Redundancy and equipment failures;
- Idle capacity and demand volatility.
The ability to generate tokens does not guarantee that customers will buy them at published API prices.
API retail prices are also not the same as the revenue received by the underlying server operator. An API provider must operate the software platform, networking, billing, security, support, and service infrastructure surrounding the model.
For these reasons, “gross contribution after electricity” should not be interpreted as net profit.
13. Bigger Models Are Not Automatically More Profitable
This model uses DeepSeek V4 Flash.
If it were replaced by the more computationally demanding V4 Pro, the compute required per token could increase, reducing the number of tokens generated per kilowatt-hour.
V4 Pro commands a higher API price, but the higher selling price may not fully offset the additional inference cost.
The result would depend on:
- Actual model inference efficiency;
- Active parameter count;
- Input and output length;
- Context length;
- Batch size;
- Cache hit rate;
- Sustained GPU utilization;
- Whether customers are willing to pay a premium for better model performance.
A larger model does not automatically produce a higher profit. A more useful metric is: How much sellable token revenue can be generated per dollar of total cost?
14. Falling Token Prices Could Change the Entire Calculation
The AI industry is experiencing a seemingly contradictory trend.
Technology companies are investing heavily in data centers, GPUs, and power capacity. At the same time, more efficient models, faster hardware, and competition among providers are driving down the price per million tokens.
For customers, cheaper tokens make it more affordable to experiment with AI support systems, automation, coding, data analysis, and AI agents.
For infrastructure operators, however, lower token prices mean that existing compute capacity may lose economic value quickly.
A server that appears profitable today could become uncompetitive because of:
- New generations of GPUs;
- More efficient model architectures;
- Better inference software;
- Additional API price cuts.
Low-cost electricity can provide:
- A lower marginal production cost per token;
- More room to reduce API prices;
- Better tolerance for low utilization;
- A longer economic life for existing hardware;
- Greater pricing competitiveness;
- More resilient long-term margins.
15. Conclusion: What 1 kWh Can Produce in the AI Token Economy
Under the engineering assumptions used in this article:
- 1 kWh of facility electricity produces approximately 418.5M tokens;
- Token revenue in the China scenario is approximately $35.61 per kWh;
- Gross contribution after the reference electricity cost is approximately $35.58 per kWh;
- Token revenue in the U.S. scenario is approximately $33.56 per kWh;
- Gross contribution after the average industrial electricity cost is approximately $33.47 per kWh.
| Key result | China low-cost power scenario | U.S. industrial-rate scenario |
|---|---|---|
| Token revenue | $35.61/kWh | $33.56/kWh |
| Gross contribution after electricity | $35.58/kWh | $33.47/kWh |
When GPU efficiency, token prices, server utilization, and other operating costs are equal, lower electricity prices reduce the marginal cost of producing each token, increase pricing flexibility, and improve an AI infrastructure project’s ability to withstand falling token prices and volatile demand.
Electricity can power GPUs, and GPUs can generate tokens. But only real, stable, paying demand can turn those tokens into a sustainable business.
Related AI Infrastructure Research
Frequently Asked Questions About AI Tokens and Electricity
1. What Is an AI Token?
An AI token is a basic unit used by a model to process and generate content. Instead of reading an entire sentence as a single object, a language model divides text into smaller pieces. A user’s prompt creates input tokens, while the model’s response creates output tokens.
2. Are AI Tokens the Same as Cryptocurrency Tokens?
No. AI tokens are units of computation and billing. Cryptocurrency tokens are generally digital assets recorded on a blockchain.
| Category | AI token | Cryptocurrency token |
|---|---|---|
| Primary purpose | AI processing and billing | Blockchain transactions or digital assets |
| Can it normally be held? | No | Often |
| Can it be traded? | Usually not | Sometimes |
3. Why Do AI Companies Charge by the Token?
Token usage reflects how much information a model processes and generates. Token-based pricing allows API providers to charge according to usage and helps developers estimate the AI cost of each task, user, or product.
4. Why Are Output Tokens More Expensive Than Input Tokens?
Output tokens require the model to generate new content sequentially, which often involves more real-time computation. Repeated input content may qualify for lower cached-input pricing.
5. Do More Tokens Always Produce a Better AI Answer?
No. More tokens can support longer context and more detailed responses, but they can also increase repetition, latency, and cost. The important question is whether those tokens produce a useful result.
6. How Can Individuals Participate in the Token Economy?
Practical options include:
- Using AI to improve writing, coding, design, research, or data analysis;
- Building AI tools, industry assistants, or agents with existing APIs;
- Helping businesses create knowledge bases and automated workflows;
- Providing AI-assisted content, marketing, education, or consulting services;
- Combining professional expertise with AI to solve industry-specific problems.
7. Can Individuals Buy AI Tokens and Earn a Return?
Usually not. Most AI tokens are units of API consumption, not assets that can be freely held, traded, or expected to appreciate. The words “AI” and “token” do not automatically indicate a connection to genuine AI infrastructure revenue.
8. Can You Make Money by Buying GPUs and Producing Tokens?
It is theoretically possible, but buying GPUs only provides the ability to produce tokens. It does not guarantee customer demand or profitability. Operators must account for hardware, electricity, cooling, networking, maintenance, utilization, customer acquisition, and depreciation.
9. What Is the Biggest Opportunity in the Token Economy?
The biggest opportunity may not be producing the most tokens, but creating more value from each token. The key question is how many tokens are required to complete a valuable task and how much revenue or cost savings that task creates.
10. What Are the Biggest Risks in the Token Economy?
Major risks include falling API prices, rapid GPU depreciation, low data center utilization, rising infrastructure costs, unstable customer demand, regulatory changes, and failing to translate token consumption into real business value.
Sources
- NVIDIA: Vera Rubin DSX AI Factory Reference Design
- NVIDIA HGX AI Factory Components
- DeepSeek API Models and Pricing
- DeepSeek Context Caching Documentation
- U.S. Energy Information Administration: Electricity Prices
- U.S. Energy Information Administration: Electric Power Monthly, Table 5.3
- Inner Mongolia Energy Administration: Multiyear Power Purchase Agreement Data
- Federal Reserve H.10 Foreign Exchange Rates