- The blended cost of AI inference fell to about $1.16 per million tokens in early August, a year-to-date low, down from $2.04 at the end of May, according to Jefferies citing Silicon Data.
- In the same window, DeepSeek launched V4-Pro and raised API prices by as much as 1,100%, detaching its premium tier from the commodity floor its own V4-Flash model helped set.
- DeepSeek's cheap models now account for the largest single share of tokens processed on OpenRouter, at roughly 27%, ahead of Google at 25%.
The blended price of a token hit a year-to-date low in August
The average cost of running a query through a large language model dropped to roughly $1.16 to $1.18 per million tokens between August 6 and 8, the lowest level so far in 2026. That is down from $1.45 in late July and $2.04 at the end of May, based on OpenRouter volume analyzed by Silicon Data and flagged by Jefferies. Two of the price cuts that pushed the blended rate down came from the top of the market. OpenAI reduced GPT-5.6 prices by as much as 80%, and Anthropic now offers Claude Opus 5 at roughly half the price of its previous flagship tier.
None of this is an anomaly. It is the continuation of a curve that Andreessen Horowitz named LLMflation, the observation that the cost of a model at a fixed quality level falls by roughly a factor of ten every year. By a16z's own accounting, a GPT-3 quality model went from $60 per million tokens in late 2021 to about $0.06 three years later, a thousandfold decline that outpaced Moore's Law.
For an LLM of equivalent performance, the cost is decreasing by a factor of 10 every year.Guido Appenzeller, Andreessen Horowitz
DeepSeek raised prices 1,100% the same month the floor fell
Here is the part that does not fit the story that everything is getting cheaper. In mid August, DeepSeek, the company most responsible for setting the commodity floor, launched V4-Pro with stronger agent capabilities and raised API pricing on its V4 line by as much as 1,100%, according to reporting from Reuters and Caixin. The same firm whose V4-Flash model processed more tokens on OpenRouter than any other now charges a steep premium for its top tier.
| Blended price per million tokens, early August | ~$1.16 |
| Blended price per million tokens, end of May | $2.04 |
| Peak DeepSeek V4 price increase | Up to 1,100% |
| DeepSeek share of tokens on OpenRouter | ~27%, ahead of Google at ~25% |
The two facts are not in tension once you stop treating inference as one product. Cheap, high volume tokens and expensive, high stakes tokens are becoming different markets with different economics, and DeepSeek is now selling into both at the same time.
The market is splitting into a commodity floor and a reasoning premium
At the bottom, the floor is collapsing toward zero. Simple classification, extraction, summarization, and retrieval are high volume, low margin, and increasingly interchangeable across providers. When a task can run on a small open weight model, price is the only variable that matters, and every provider is racing the same curve downward.
Inference did not just get cheaper. It split into two markets moving in opposite directions.
At the top, a different logic applies. Long horizon agents, verified reasoning, and tool use that has to work on the first try are where reliability commands pricing power. That is the tier DeepSeek carved out with V4-Pro, and it is why a premium can rise 1,100% in a market where the average price just hit an annual low. When a single failed step can break an entire agent workflow, buyers will pay far more per token to avoid it, and that willingness to pay is what a premium tier monetizes. The value is no longer in the token. It is in the confidence that the token is correct.
What bifurcation means for anyone building on top of these APIs
For companies building on these APIs, single vendor loyalty is now the expensive choice. The rational architecture routes each task to the cheapest model that clears the quality bar, sending bulk work to the floor and reserving the premium tier for the small fraction of calls where a mistake is costly. That makes the routing layer, the software that decides which model gets which request, one of the most strategically valuable pieces of the stack, which is precisely why it is now being bought at model lab prices.
It also resets how founders should read a price cut. A cheaper commodity token lowers the cost of experimentation to almost nothing, which is good for anyone starting out. But durable margin and customer lock in are migrating to the premium tier, where reliability, not raw capability, is the product being sold.
The claim that inference is getting cheaper is true and incomplete. The floor is collapsing toward zero while the ceiling pulls away, and the more useful question is no longer what a token costs. It is which tokens are worth paying for.
Santage is committed to independent, transparent journalism. This article is produced in accordance with Santage's Editorial Standards and aims to provide accurate and timely information. Readers are encouraged to verify information independently.