NEWS

DeepSeek's V4.1-Flash Tops Claude Opus 5 and GPT-5.6 Sol on Coding, at a Fraction of the Price

The DeepSeek logo on a dark background, representing the V4.1-Flash model release
DeepSeek's V4.1-Flash matches the paid frontier on one coding benchmark while pricing far below it. Source: DeepSeek
Quick answer: DeepSeek released V4.1-Flash on September 10, 2026, a low-cost model that scores 74.2 on the DeepSWE v1.1 agentic-coding benchmark, narrowly ahead of Claude Opus 5 at 74.0 and GPT-5.6 Sol at 73.0. Off-peak, it charges as little as $0.15 per million input tokens and $0.60 per million output, a fraction of what frontier labs ask. It replaces V4 Flash immediately and absorbs V4 Pro traffic from September 14, moving frontier-level coding to commodity prices.
TLDR

What DeepSeek shipped

DeepSeek published V4.1-Flash on September 10 as a direct replacement for its V4 Flash tier, with V4 Pro requests set to route to the new model from September 14 until a V4.1 Pro arrives. The architecture is a 552-billion-parameter mixture of experts with a causal encoder-decoder design that keeps only 8 billion parameters active while reading a prompt and 16 billion while writing the answer. That sparsity is what lets DeepSeek serve a model this capable at the price it is quoting.

Bar chart comparing DeepSWE v1.1 agentic-coding scores: DeepSeek V4.1-Flash at 74.2, Claude Opus 5 at 74.0, and GPT-5.6 Sol at 73.0, with DeepSeek off-peak pricing of 0.15 dollars per million input tokens and 0.60 per million output
V4.1-Flash edges the frontier on one agentic-coding test while pricing well below it. Source: Santage analysis; DeepSeek model card, September 10, 2026.

On the headline number, V4.1-Flash takes a narrow lead. On the tests that measure sustained agentic work and hard reasoning, it does not. Opus 5 still leads on Terminal-Bench, and Sol holds an edge on GPQA Diamond, so the fair reading is parity on coding rather than a clean sweep.

Why the price is the headline

The benchmark gap is small enough to argue about. The price gap is not. DeepSeek's model card lists the model, its context window, and its routing plan in plain terms.

Legacy V4 Flash requests are now served by V4.1-Flash. V4 Pro will route to V4.1-Flash after September 14, until a future V4.1 Pro arrives.
DeepSeek, V4.1-Flash model card, September 10, 2026

Off-peak, developers pay $0.15 per million input tokens, $0.60 per million output, and as little as $0.003 per million on cached input, with peak rates simply doubling those figures during two windows on weekdays. For teams running coding agents that read large repositories on every call, cached-input pricing at that level changes what is affordable to run continuously rather than in short bursts.

What it means for the frontier

The competitive pressure lands on the paid frontier tiers. When an open-weight model matches a flagship on the benchmark buyers cite most for coding, the case for paying frontier prices narrows to the tasks where Opus 5 and Sol still pull ahead, namely long-horizon agentic reliability and the hardest reasoning. DeepSeek is betting that most production coding work does not need that headroom, and that price will decide the rest. The launch continues a pattern visible across the MBZUAI K2 open-model fleet and rival releases, where the gap between open and closed keeps closing on capability while staying wide on cost. It arrives days after OpenAI's GPT-6 Astra launch reset the top of the market, and it answers that move from the bottom.

In short: DeepSeek's V4.1-Flash scores 74.2 on DeepSWE v1.1, narrowly ahead of Claude Opus 5 and GPT-5.6 Sol, at off-peak prices from $0.15 per million input tokens. It replaces V4 Flash now and V4 Pro from September 14, pressuring the paid frontier on cost while the top labs keep their lead on the hardest reasoning and long-horizon tasks.

Santage is committed to independent, transparent journalism. This article is produced in accordance with Santage's Editorial Standards and aims to provide accurate and timely information. Readers are encouraged to verify information independently.