- Anthropic and Epoch AI published three self-measured indexes tracking how much of the company's AI research Claude runs, how closely its agents are supervised, and how much compute goes to safety.
- The headline number puts Claude in the lead on 26% of Anthropic's AI R&D as of August, up from under 1% in February, with more than 90% of the work at the "collaborates" level or above and none fully autonomous.
- Claude generated most of the ratings itself, agreeing with employees 59% of the time while two employees agreed only 35% of the time, and Anthropic says no part of the index has been externally audited.
What Anthropic actually measured
Anthropic and the research group Epoch AI released three indexes on Thursday, the first time a frontier lab has attached numbers to how much of its own work its model runs. One index tracks AI-led research, a second tracks oversight of AI agents, and a third tracks how compute splits between safety and everything else. The scale under the first index comes from Epoch AI, which describes it as a step toward an O*NET for AI R&D, the standardized catalogue that classifies human occupations.
The rating scale runs from AL0, where no AI is involved, through AL3, where Claude collaborates on large parts of a task under close direction, to AL4, where Claude leads and finishes most of a task end to end from a high-level prompt while a person supervises. AL5, full autonomy, sits at the top and describes none of the work Anthropic measured. As of August, 26% of the company's AI research sat at AL4 or above, against under 1% in February, and more than 90% reached at least the collaborates level.
The methodology is narrower than the headline suggests. Anthropic sampled 20% of staff across its model research departments in July, identified about 15,000 granular tasks, and organized them into a 542-node tree with 378 leaves. Claude agents then rated the sampled work by reading Slack messages and internal documents, and the compute index drew on roughly 14% of compute jobs, weighted toward the heaviest runs.
| Level | What it means |
|---|---|
| AL0 | No AI involvement |
| AL1 | Minimal AI involvement |
| AL2 | AI assists with parts of a task |
| AL3 | AI collaborates on large chunks under close human direction |
| AL4 ("leads") | AI completes most of a task end to end from a high-level prompt, with human supervision |
| AL5 | Fully autonomous, no human involvement |
Source: Epoch AI automation scale, via Anthropic, September 2026.
The oversight index describes many agents and a thin human net
The second index gives the clearest picture of scale. Anthropic ran about 30,000 Claude agents at once on its main internal platform in August. An online monitor watched every action and blocked roughly 1 in 47,000, or 0.002%, while an offline monitor flagged one to two transcripts per thousand for later review. Human reviewers looked at around 50 high-priority flags a week.
Those numbers cut two ways. Full monitoring coverage across more than a billion agent decisions is a real engineering achievement, and 50 human-reviewed flags a week is a thin net beneath 30,000 concurrent agents. Anthropic notes that it has run this monitoring for months rather than years and cannot be certain it captures every kind of behavior worth catching.
| Concurrent Claude agents | About 30,000 on the main internal platform |
| Online monitor coverage | 100%, blocking about 1 in 47,000 actions |
| Offline monitor | Flags 1 to 2 transcripts per 1,000 |
| Human review | About 50 high-priority flags per week |
The measurement runs on the model it is measuring
Claude generated most of the ratings in the first index, and that is the part that deserves the most scrutiny. When Anthropic checked Claude's ratings against employees' own assessments, the two matched exactly 59% of the time. Two employees rating the same work matched only 35% of the time, and 97% of ratings landed within one level of each other. Anthropic reads the higher Claude-to-human agreement as a sign the model grades consistently.
Consistency and accuracy are different things. A grader built on the same data as the worker tends to share its blind spots, the sampled tasks form a frozen basket that cannot register genuinely new kinds of work, and real disagreement exists over where the borderline levels fall. Anthropic states the core risk plainly in the report.
The "judge" model could make the same kinds of errors as the model it is checking.Anthropic, Measurements for understanding the pace of AI development inside frontier labs
The compute index and the line every lab will draw generously
The third index reports that 6% of Anthropic's AI R&D compute went to safety in the week of July 13, rising to 12% for the compute tied specifically to AI-driven research. Anthropic adds an unusually candid caveat: a compute share measures only what is spent, so efficiency gains can shrink the safety percentage without any drop in effort, and a single week cannot show a trend.
The deeper problem is definitional, and Anthropic names it.
Safety research is hard to distinguish from capabilities research, and each developer will be tempted to draw the line generously.Anthropic, Measurements for understanding the pace of AI development inside frontier labs
A self-report is only as trustworthy as the distance between the reporter and the subject. Here the reporter and the subject are the same model, so the index measures Claude's view of Claude's work. That does not make the 26% wrong. It makes it a claim only an outside instrument can confirm.
How to read a lab grading its own homework
Bringing in Epoch AI to co-develop the scale is a real step toward outside rigor, and it stops short of outside measurement, because Anthropic still generated the ratings the scale produced. That distinction will matter as other labs publish automation metrics of their own. The companies holding the instruments to measure how much AI runs their research are the same companies with a commercial interest in a large number.
For anyone outside these labs, the index reads best as a direction rather than a fixed data point. The precise 26% may not survive an independent audit, and the trajectory it traces, from a rounding error to a quarter of the work in six months, matches what Anthropic's own leaders have said in public about the pace of internal automation.
The most revealing line in the release is Anthropic's admission that the model grading the work could be making the same mistakes as the model doing it. Every automation milestone a lab reports now carries that asterisk, impressive and plausible and graded by the thing it describes, and it will keep carrying it until an outside instrument can measure what these systems actually run.
Santage is committed to independent, transparent journalism. This article is produced in accordance with Santage's Editorial Standards and aims to provide accurate and timely information. Readers are encouraged to verify information independently.