- On the third-party Artificial Analysis Intelligence Index, GPT-6 Astra scores 61.2, just 0.2 ahead of Grok 4.6 and GPT-5.6 Sol Max at 61, so the single number most people cite barely moved.
- The gap opens on autonomy. Astra posts 72.6% on OSWorld computer use, reached roughly 47% faster than Sol, and 57.9% on Terminal-Bench 4.0 coding, narrowly ahead of Claude Fable 5.1 at 55.8% and far above Sol's 37.3%.
- Google still has no frontier model in the market, shipping only Gemini Flash updates, which leaves the top tier a three-way contest between OpenAI, Anthropic and SpaceXAI.
The headline index barely moved
The number most readers reach for is the one that says the least this week. On the Artificial Analysis Intelligence Index, the cross-lab aggregate that rolls reasoning, math, coding and knowledge into a single figure, GPT-6 Astra lands at 61.2. Grok 4.6 sits at 61.0. So does GPT-5.6 Sol Max. The most capable model OpenAI has ever shipped clears the previous frontier by two tenths of a point on the measure the industry treats as its scoreboard.
That is not a knock on Astra. It is a warning about the scoreboard. Aggregate indices average across tasks, and once the leaders all score in the high nineties on the classic academic tests, there is little headroom left to average. The index has quietly stopped being able to tell the frontier models apart. To see what actually separates them, you have to break the number back into its parts.
Where Astra actually separates
Broken apart, the picture changes. On OSWorld 2.0, the test of a model driving a real computer through multi-step work, Astra reports 72.6% against Sol's 65.7%, and it gets there in about 40 minutes where Sol took roughly 75. Speed at a given quality is its own capability once these models are running as agents rather than answering prompts, and that is the axis where Astra has moved most.
Coding tells a subtler story. Astra's 57.9% on Terminal-Bench 4.0 is a large jump over Sol's 37.3%, but Claude Fable 5.1 is right behind at 55.8%, and Anthropic shipped Fable 5.1 with a 75% cut to cache-read prices. On the task enterprises are spending the most on, the field is close and the cheaper model is closing, not the pricier one pulling away.
| GPT-6 Astra | Index 61.2, Terminal-Bench 4.0 57.9%, OSWorld 72.6%, Critical cyber tier, $10 / $50, context to 1M |
| Claude Fable 5.1 | Terminal-Bench 4.0 55.8%, $10 / $50 with cache reads cut to $0.25, 1M context, restricted Mythos twin |
| Claude Opus 5 | 96% on SWE-bench Verified at $5 / $25, half Astra's list price, 1M context |
| Grok 4.6 | Index 61.0, frontier parity from post-training rather than a larger base model |
| Google Gemini | no frontier model in market, only Flash updates through Gemini 3.8 Flash |
Price is where the lead gets awkward
Astra arrives at $10 per million input tokens and $50 output, exactly Fable 5.1's rates. That is parity, not a discount, and it sits above Claude Opus 5, which posts 96% on SWE-bench at half the input price. For a buyer who does not need the last few points of computer-use accuracy or the Critical-tier headroom, the frontier now has genuinely cheaper options that were not there a quarter ago. Astra's separation is real, but it is concentrated in autonomy and cyber, the two capabilities most enterprises are least ready to turn on. The gap that shows up in benchmarks is smallest exactly where buying decisions are largest.
The safeguard is now part of the spec
Astra ships at the Critical cyber tier with offensive-security tasks refused by default, the first model to carry that label. Anthropic reached the same design conclusion from the other direction, splitting Fable 5.1 from a restricted Mythos 5.1 twin reserved for vetted cybersecurity and life-sciences work. Two of the three frontier labs now treat the safeguard structure, not just the benchmark, as part of what they ship. For a buyer that means the comparison is no longer only which model scores highest. It is also which capabilities each vendor will let you switch on, and Astra's most distinctive numbers sit behind exactly the gate OpenAI is most cautious about opening.
The Google-shaped hole
The most striking entry in any 2026 comparison is the one that is missing. Google has shipped four Gemini Flash models in roughly a hundred days and still has no frontier flagship in the market, after Gemini 3.5 Pro slipped a third deadline in a pattern that looked structural. The company that was supposed to make this a two-horse race is, for now, watching it from the Flash tier. The top of the market is a contest between OpenAI, Anthropic and, improbably, SpaceXAI's Grok.
The frontier is no longer one number. Astra wins it, but only if you measure the things enterprises are least ready to deploy.
That is the real read on GPT-6 Astra against the field. It is the best model in the market on the capabilities that matter for agents, and it is roughly tied with the previous generation on the capabilities that fit on a chart. Altman's warning that the next generation will be "sobering for everybody" is a promise that the interesting gap, the one in autonomy and cyber, is about to get wider. The scoreboard everyone watches will be the last place it shows up.
Santage is committed to independent, transparent journalism. This article is produced in accordance with Santage's Editorial Standards and aims to provide accurate and timely information. Readers are encouraged to verify information independently.