- V4.1-Flash posts 74.2 on DeepSWE v1.1, edging Claude Opus 5 (74.0) and GPT-5.6 Sol (73.0), though both US models still lead on Terminal-Bench and GPQA Diamond.
- Off-peak pricing runs $0.15 per million input tokens and $0.60 per million output, with cached input as low as $0.003 per million, well below frontier rates.
- The 552-billion-parameter mixture-of-experts model activates only 8 to 16 billion parameters per call and ships with a 1 million-token context window.
What DeepSeek shipped
DeepSeek published V4.1-Flash on September 10 as a direct replacement for its V4 Flash tier, with V4 Pro requests set to route to the new model from September 14 until a V4.1 Pro arrives. The architecture is a 552-billion-parameter mixture of experts with a causal encoder-decoder design that keeps only 8 billion parameters active while reading a prompt and 16 billion while writing the answer. That sparsity is what lets DeepSeek serve a model this capable at the price it is quoting.
On the headline number, V4.1-Flash takes a narrow lead. On the tests that measure sustained agentic work and hard reasoning, it does not. Opus 5 still leads on Terminal-Bench, and Sol holds an edge on GPQA Diamond, so the fair reading is parity on coding rather than a clean sweep.
Why the price is the headline
The benchmark gap is small enough to argue about. The price gap is not. DeepSeek's model card lists the model, its context window, and its routing plan in plain terms.
Legacy V4 Flash requests are now served by V4.1-Flash. V4 Pro will route to V4.1-Flash after September 14, until a future V4.1 Pro arrives.DeepSeek, V4.1-Flash model card, September 10, 2026
Off-peak, developers pay $0.15 per million input tokens, $0.60 per million output, and as little as $0.003 per million on cached input, with peak rates simply doubling those figures during two windows on weekdays. For teams running coding agents that read large repositories on every call, cached-input pricing at that level changes what is affordable to run continuously rather than in short bursts.
What it means for the frontier
The competitive pressure lands on the paid frontier tiers. When an open-weight model matches a flagship on the benchmark buyers cite most for coding, the case for paying frontier prices narrows to the tasks where Opus 5 and Sol still pull ahead, namely long-horizon agentic reliability and the hardest reasoning. DeepSeek is betting that most production coding work does not need that headroom, and that price will decide the rest. The launch continues a pattern visible across the MBZUAI K2 open-model fleet and rival releases, where the gap between open and closed keeps closing on capability while staying wide on cost. It arrives days after OpenAI's GPT-6 Astra launch reset the top of the market, and it answers that move from the bottom.
Santage is committed to independent, transparent journalism. This article is produced in accordance with Santage's Editorial Standards and aims to provide accurate and timely information. Readers are encouraged to verify information independently.