ANALYSIS

AI Token Prices Hit a Yearly Low, DeepSeek Hiked 1,100%

A white robotic hand reaching toward a glowing blue AI interface surrounded by digital icons on a dark background
AI inference prices fell to a 2026 low even as premium reasoning tiers climbed. Source: Stock image
TLDR

The blended price of a token hit a year-to-date low in August

The average cost of running a query through a large language model dropped to roughly $1.16 to $1.18 per million tokens between August 6 and 8, the lowest level so far in 2026. That is down from $1.45 in late July and $2.04 at the end of May, based on OpenRouter volume analyzed by Silicon Data and flagged by Jefferies. Two of the price cuts that pushed the blended rate down came from the top of the market. OpenAI reduced GPT-5.6 prices by as much as 80%, and Anthropic now offers Claude Opus 5 at roughly half the price of its previous flagship tier.

Line chart showing blended AI inference price per million tokens falling from 2.04 dollars on May 31 to 1.45 dollars in late July to 1.16 dollars in early August 2026
The blended price of a million tokens fell to a 2026 low in early August. Source: Jefferies, citing Silicon Data.

None of this is an anomaly. It is the continuation of a curve that Andreessen Horowitz named LLMflation, the observation that the cost of a model at a fixed quality level falls by roughly a factor of ten every year. By a16z's own accounting, a GPT-3 quality model went from $60 per million tokens in late 2021 to about $0.06 three years later, a thousandfold decline that outpaced Moore's Law.

For an LLM of equivalent performance, the cost is decreasing by a factor of 10 every year.
Guido Appenzeller, Andreessen Horowitz

DeepSeek raised prices 1,100% the same month the floor fell

Here is the part that does not fit the story that everything is getting cheaper. In mid August, DeepSeek, the company most responsible for setting the commodity floor, launched V4-Pro with stronger agent capabilities and raised API pricing on its V4 line by as much as 1,100%, according to reporting from Reuters and Caixin. The same firm whose V4-Flash model processed more tokens on OpenRouter than any other now charges a steep premium for its top tier.

Two markets, one month
Blended price per million tokens, early August~$1.16
Blended price per million tokens, end of May$2.04
Peak DeepSeek V4 price increaseUp to 1,100%
DeepSeek share of tokens on OpenRouter~27%, ahead of Google at ~25%
Source: Jefferies and Silicon Data; Reuters; Caixin. August 2026.

The two facts are not in tension once you stop treating inference as one product. Cheap, high volume tokens and expensive, high stakes tokens are becoming different markets with different economics, and DeepSeek is now selling into both at the same time.

The market is splitting into a commodity floor and a reasoning premium

At the bottom, the floor is collapsing toward zero. Simple classification, extraction, summarization, and retrieval are high volume, low margin, and increasingly interchangeable across providers. When a task can run on a small open weight model, price is the only variable that matters, and every provider is racing the same curve downward.

Inference did not just get cheaper. It split into two markets moving in opposite directions.

At the top, a different logic applies. Long horizon agents, verified reasoning, and tool use that has to work on the first try are where reliability commands pricing power. That is the tier DeepSeek carved out with V4-Pro, and it is why a premium can rise 1,100% in a market where the average price just hit an annual low. When a single failed step can break an entire agent workflow, buyers will pay far more per token to avoid it, and that willingness to pay is what a premium tier monetizes. The value is no longer in the token. It is in the confidence that the token is correct.

What bifurcation means for anyone building on top of these APIs

For companies building on these APIs, single vendor loyalty is now the expensive choice. The rational architecture routes each task to the cheapest model that clears the quality bar, sending bulk work to the floor and reserving the premium tier for the small fraction of calls where a mistake is costly. That makes the routing layer, the software that decides which model gets which request, one of the most strategically valuable pieces of the stack, which is precisely why it is now being bought at model lab prices.

It also resets how founders should read a price cut. A cheaper commodity token lowers the cost of experimentation to almost nothing, which is good for anyone starting out. But durable margin and customer lock in are migrating to the premium tier, where reliability, not raw capability, is the product being sold.

The claim that inference is getting cheaper is true and incomplete. The floor is collapsing toward zero while the ceiling pulls away, and the more useful question is no longer what a token costs. It is which tokens are worth paying for.

Santage is committed to independent, transparent journalism. This article is produced in accordance with Santage's Editorial Standards and aims to provide accurate and timely information. Readers are encouraged to verify information independently.