NEWS

Top AI Models Fix Under 40% of Real Enterprise Code

The Real-SWE benchmark title graphic from Specific Labs, with a purple and black particle vortex on a white background
Specific Labs' Real-SWE benchmark evaluates frontier models on private enterprise codebases. Source: Y Combinator
Quick answer: Specific Labs released Real-SWE on September 12, 2026, a benchmark that tests frontier AI models on private, licensed enterprise codebases they were never trained on. The best model, Anthropic's Fable 5.1, resolved 38.8% of tasks, ahead of OpenAI's GPT-6 Astra at 33.8% and Google's Gemini 3.8 Flash at 31.2%, and six of ten tasks stayed below 15%. Because the code is private, the scores fall well below the two-thirds-plus that leading models post on public benchmarks like SWE-bench, a gap that measures how much of real enterprise work still sits outside what these models have seen.
TLDR

Real-SWE tests models on code they were never trained on

Specific Labs released Real-SWE, a benchmark built to answer a question public leaderboards cannot: how well do frontier models handle software work at real companies. Instead of drawing tasks from open-source projects that models have almost certainly seen during training, Real-SWE licenses private codebases from actual businesses and asks models to complete genuine engineering tasks in them, including billing logic, tax handling, and customer migrations across services.

The tasks are harder than the public equivalents by design. Changes span an average of 11 files against 6 in comparable benchmarks, and the instructions are deliberately underspecified, forcing an agent to discover how a proprietary system works before it can safely change it. Across 640 scored rollouts on eight model-harness configurations, the top score belonged to Anthropic's Fable 5.1 at 38.8%, according to Specific Labs. OpenAI's GPT-6 Astra followed at 33.8% and Google's Gemini 3.8 Flash at 31.2%.

Horizontal bar chart of Real-SWE resolution rates on private enterprise codebases: Fable 5.1 from Anthropic leads at 38.8 percent, GPT-6 Astra from OpenAI at 33.8 percent, Gemini 3.8 Flash from Google at 31.2 percent, GLM 5.3 from Alibaba at 28.8 percent, Grok 4.6 from xAI and Muse Spark 1.3 from Meta tied at 23.8 percent, Kimi K3 at 18.8 percent, and GPT-5.6 Sol from OpenAI at 16.2 percent
No frontier model cleared 40% on private enterprise code, and the field spread from 38.8% down to 16.2%. Source: Specific Labs, Real-SWE benchmark, September 2026.

Why 38.8% is the number that matters

On public benchmarks such as SWE-bench, the leading models now resolve well over two-thirds of tasks, and that figure has anchored a year of claims that AI can do much of a software engineer's job. Real-SWE cuts that number roughly in half by removing the one advantage those tests quietly give: familiarity. When the code is private and the model has never seen it, the best system in the world resolves fewer than four tasks in ten, and the median rollout that fails does so by missing requirements it was never explicitly given.

99% of tokens in real-world enterprises are hidden away from the frontier models.
Specific Labs, on why private-codebase evaluation matters, September 2026

The cost data sharpens the point. Gemini 3.8 Flash matched far more expensive systems at $2.50 per rollout, and 71.4% of rollouts that ran under ten minutes failed, which says the gap is about reasoning through unfamiliar systems, not about spending more compute. The lesson for enterprises weighing a coding-agent rollout is not that these tools are weak. It is that a benchmark score earned on public code is a poor guide to how a model will perform inside a proprietary system it has never encountered. The frontier of coding ability is now measured less by what a model knows than by how well it can learn a codebase it was never trained on, and by that measure the leaders still have most of the work ahead of them.

In short: Real-SWE, a new Specific Labs benchmark, tests frontier models on private enterprise codebases they were never trained on, and the best of them, Fable 5.1, resolved only 38.8% of tasks, with six of ten tasks below 15%. The scores fall roughly half of what leading models post on public benchmarks, exposing how much real enterprise work still sits outside what these models have seen.

Santage is committed to independent, transparent journalism. This article is produced in accordance with Santage's Editorial Standards and aims to provide accurate and timely information. Readers are encouraged to verify information independently.