NEWS

Anthropic Reveals Unreleased Model 2 and Raises Its Misalignment Risk

A smartphone displaying the Anthropic AI logo in white on a black screen, with the Anthropic wordmark on the wall behind it
Anthropic disclosed two unreleased internal models in its August 2026 risk report. Source: Stock image
TLDR

Anthropic disclosed a model more capable than its public flagship

Anthropic's latest risk report, a redacted document running to roughly 186 pages, revealed that the company is running two unreleased successors to Claude Mythos 5 internally. It calls them Model 1 and Model 2. Model 1 is described as broadly similar to earlier releases, while Model 2 is a noticeable improvement on Mythos 5 for many tasks relevant to internal use, though the report is careful to say it is not the kind of leap seen between Opus 4.6 and Mythos Preview.

Model 2 is not a lab curiosity. Anthropic says staff use it heavily for software development, generating training data, and automating engineering work, the tasks that most directly speed up the company's own research. What the public can buy today is a step behind what the company builds with internally.

The more striking disclosure is about risk. Anthropic raised its assessed misalignment risk for high-stakes autonomy from very low to low, and it was explicit that the change was driven by uncertainty rather than evidence of harm.

We believe that the arguments presented below likely still support a designation of very low risk, but we are raising our assessed risk to low to reflect increased overall uncertainty.
Anthropic, August 2026 Risk Report
Inside the report
Unreleased internal models disclosed2 (Model 1 and Model 2)
Misalignment risk for high-stakes autonomyRaised from very low to low
Mythos 5 stealth success on SHADE-ArenaUnder 1% with extended thinking
Length of the redacted report~186 pages
Source: Anthropic, August 2026 Risk Report.

Why a self-imposed risk raise matters more than the model

Companies almost never volunteer bad news about their own products. A frontier lab publishing a report that raises its own risk rating, while sitting on a more capable model it has chosen not to ship, is a governance signal worth more than the specifications of Model 2.

The reason for the raise is the part to sit with. Anthropic did not say its models became more dangerous. It said it is less confident that its own evaluations keep pace with how fast the models are improving, and flagged that it is rewriting its threat models in light of recent incidents. That is a lab admitting its measuring instruments may be lagging the thing they are meant to measure.

The most important line in the report is not about Model 2. It is Anthropic conceding that it raised its own risk estimate because it is no longer sure its tests still hold. When the companies building the most capable systems say their safety measurements are struggling to keep up, the confidence gap, not the capability gap, is the thing to watch.

Santage is committed to independent, transparent journalism. This article is produced in accordance with Santage's Editorial Standards and aims to provide accurate and timely information. Readers are encouraged to verify information independently.