Open any model provider's pricing page and the same pattern repeats: output costs more than input, sometimes by a wide margin, and the number resets every few months. It reads like an arbitrary markup. It is not. Input and output are two different computational jobs running on the same chip, and the price sheet is a fairly direct readout of how expensive each one actually is to run. Once that mechanism is visible, the second pattern, a Chinese lab selling a comparable answer for two orders of magnitude less, stops looking like a mystery and starts looking like a predictable result of a different set of engineering and business choices.
What a token actually is
A token is not a word. It is a chunk of text a model's tokenizer splits language into, roughly three to four characters in English on average, sometimes a whole common word, sometimes a fragment of a rare one. Every dollar figure on a pricing page is quoted per million tokens, because a single request is a rounding error and providers need a unit that scales.
Why input and output are priced differently
The two directions are not the same operation. When you send a prompt, the model reads the entire thing in one parallel pass called prefill. Every input token gets processed simultaneously, the hardware runs near its peak throughput, and the per-token cost of that pass is low. That is why context, the transcript, the reference document, the system prompt, is comparatively cheap.
Generating a response works the opposite way. The model produces one token, appends it to the sequence, then runs a fresh forward pass to decide the next one, and repeats that for every token it writes. A 1,000-token answer is roughly 1,000 sequential passes, not one. That process is bottlenecked by memory bandwidth rather than raw compute, and the key-value cache it has to keep re-reading grows with every token generated, which makes it progressively harder to batch efficiently with other requests in flight. Sequential, memory-bound, and hard to batch is a worse position on every axis that determines cost, and it is why output tokens run two to eight times the price of input tokens across every major lab, not a policy choice by any one of them.
Output isn't priced higher because a vendor decided your answer is worth more than your question. It's priced higher because writing token 500 means re-reading everything that came before it, one more time, in sequence, with no way to parallelize the wait.
What the price sheet actually says, September 2026
Anthropic's current published rates: Claude Haiku 4.5 at $1 per million input tokens and $5 per million output, Claude Sonnet 5 at $2 in and $10 out, Claude Opus 5 at $5 in and $25 out. OpenAI's GPT-5.5 lists at $5 in and $30 out, with the reasoning-heavier GPT-5.5 Pro tier at $30 in and $180 out. The ratio holds everywhere: output consistently runs 5x to 6x the input price on frontier models, which tracks almost exactly with the compute asymmetry above.
Output runs 5x to 6x the input price
List price in dollars per million tokens, September 2026
| Model | Input, per million tokens | Output, per million tokens |
|---|---|---|
| Claude Haiku 4.5 | $1 | $5 |
| Claude Sonnet 5 | $2 | $10 |
| Claude Opus 5 | $5 | $25 |
| GPT-5.5 | $5 | $30 |
Two mechanisms bend those numbers in practice. Prompt caching lets a provider re-serve a prompt prefix it has already processed at roughly a tenth of the standard input rate, which matters enormously for any workflow that reuses the same long system prompt or reference document across many calls. And batch processing, queuing requests for asynchronous completion instead of a live response, runs at a flat 50% discount on both input and output, because the provider can schedule it into idle capacity instead of serving it on demand. Neither changes the underlying compute cost per token. Both change who absorbs the cost of idle time.
The gap with China
Set those numbers next to DeepSeek's current pricing and the comparison stops looking like the same market. DeepSeek V4 Flash lists at $0.14 per million input tokens and $0.28 output. Its larger V4 Pro tier runs $0.435 in and $0.87 out. Against GPT-5.5's $5 and $30, that is not a modest discount, it is roughly one to two orders of magnitude cheaper on the output side, the exact number depending on which tier you line up against which.
The gap with China, on one scale
List price in dollars per million tokens
| Model | Input, per million tokens | Output, per million tokens |
|---|---|---|
| GPT-5.5 | $5 | $30 |
| DeepSeek V4 Pro | $0.435 | $0.87 |
That gap is real, and it is not a single cause.
The largest piece is architectural. Chinese frontier labs lean hard into sparse mixture-of-experts designs, models with enormous total parameter counts that activate only a small slice of themselves per token. DeepSeek V4 Pro holds roughly 1.6 trillion parameters but activates about 49 billion of them, close to 3%, for any given token. Combined with compressed and sparse attention, lower-precision arithmetic and cache-aware serving, the effect compounds on long-context requests specifically: DeepSeek reports V4 Pro using about 27% of the compute and 10% of the cache that its own previous generation needed for a million-token context. That is a genuine engineering gap, not a pricing trick, and it shows up directly in what a provider can afford to charge and still cover its costs.
The rest is business strategy layered on top of that engineering advantage. Several Chinese labs are pricing at a loss to buy market share and usage data while the market is still being decided: MiniMax reported roughly $251 million in losses against $79 million in revenue in 2025, a ratio that only makes sense if the company is optimizing for position rather than margin. Alibaba cut Qwen-Long pricing by 97% within a matter of weeks, a move that tracks a price war more than a cost curve. Local government compute subsidies, Shanghai alone allocated roughly CNY 600 million in compute vouchers, lower the floor further, though subsidies alone cannot explain away an architecture that is doing measurably less work per token. And because several of these models ship as open weights, any hosting provider can compete to serve them, which pushes retail prices toward the marginal cost of inference rather than whatever the lab that trained the model would prefer to charge.
None of that makes the models equivalent. Independent capability benchmarks put DeepSeek V4 Pro meaningfully behind GPT-5.6-class models on hard reasoning tasks, at something like a fortieth of the per-task cost. The honest read is not "China is 40x more efficient," it is: cheaper architecture, a market willing to eat losses for share, subsidized power and compute, and open weights driving retail margins toward zero, stacked on top of each other, buys a real but not free discount, and the quality gap is part of what you're trading for the price gap.
What it means for budgeting production work
For a studio, the practical takeaway is not "switch to the cheapest model." It is that input and output are genuinely different cost centers, and a workflow that generates long outputs against a short prompt, drafting, iteration, brainstorming, sits in a completely different cost bracket than one that reasons over a long input and returns a short answer, review, extraction, QA. Route each job to the pricing shape that fits it, use caching aggressively on anything with a repeated system prompt or reference document, and treat a frontier-cheap Chinese model as a real option for high-volume, lower-stakes work, not as a drop-in replacement for the model doing your hardest reasoning. The price sheet is not noise. It is a fairly accurate map of what each kind of work actually costs to run, on whichever side of the Pacific you're running it.