ANALYSIS

Anthropic Says Claude Leads 26% of Its AI Research

Hands typing on a laptop with streams of colored data flowing from the screen
Anthropic published its first public measure of how much of its own research its model now runs. Source: Qualcomm
Quick answer: Anthropic and the research group Epoch AI published three indexes on September 18, 2026, measuring how much of Anthropic's AI research Claude runs, how closely its agents are supervised, and how much compute goes to safety. The headline figure puts Claude in the lead on 26% of the company's AI R&D as of August, up from under 1% in February. Claude generated most of the ratings itself, and Anthropic says no part of the index has been externally audited.
TLDR
Bar chart showing the share of Anthropic AI R&D where Claude leads at AL4 rising from under 1 percent in February 2026 to 26 percent in August 2026
Anthropic's self-reported share of AI R&D at the "leads" level climbed from a rounding error to a quarter in six months. Ratings were generated by Claude and are unaudited. Chart: Santage · Data: Anthropic and Epoch AI

What Anthropic actually measured

Anthropic and the research group Epoch AI released three indexes on Thursday, the first time a frontier lab has attached numbers to how much of its own work its model runs. One index tracks AI-led research, a second tracks oversight of AI agents, and a third tracks how compute splits between safety and everything else. The scale under the first index comes from Epoch AI, which describes it as a step toward an O*NET for AI R&D, the standardized catalogue that classifies human occupations.

The rating scale runs from AL0, where no AI is involved, through AL3, where Claude collaborates on large parts of a task under close direction, to AL4, where Claude leads and finishes most of a task end to end from a high-level prompt while a person supervises. AL5, full autonomy, sits at the top and describes none of the work Anthropic measured. As of August, 26% of the company's AI research sat at AL4 or above, against under 1% in February, and more than 90% reached at least the collaborates level.

The methodology is narrower than the headline suggests. Anthropic sampled 20% of staff across its model research departments in July, identified about 15,000 granular tasks, and organized them into a 542-node tree with 378 leaves. Claude agents then rated the sampled work by reading Slack messages and internal documents, and the compute index drew on roughly 14% of compute jobs, weighted toward the heaviest runs.

LevelWhat it means
AL0No AI involvement
AL1Minimal AI involvement
AL2AI assists with parts of a task
AL3AI collaborates on large chunks under close human direction
AL4 ("leads")AI completes most of a task end to end from a high-level prompt, with human supervision
AL5Fully autonomous, no human involvement

Source: Epoch AI automation scale, via Anthropic, September 2026.

The oversight index describes many agents and a thin human net

The second index gives the clearest picture of scale. Anthropic ran about 30,000 Claude agents at once on its main internal platform in August. An online monitor watched every action and blocked roughly 1 in 47,000, or 0.002%, while an offline monitor flagged one to two transcripts per thousand for later review. Human reviewers looked at around 50 high-priority flags a week.

Those numbers cut two ways. Full monitoring coverage across more than a billion agent decisions is a real engineering achievement, and 50 human-reviewed flags a week is a thin net beneath 30,000 concurrent agents. Anthropic notes that it has run this monitoring for months rather than years and cannot be certain it captures every kind of behavior worth catching.

The oversight index in numbers
Concurrent Claude agentsAbout 30,000 on the main internal platform
Online monitor coverage100%, blocking about 1 in 47,000 actions
Offline monitorFlags 1 to 2 transcripts per 1,000
Human reviewAbout 50 high-priority flags per week

The measurement runs on the model it is measuring

Claude generated most of the ratings in the first index, and that is the part that deserves the most scrutiny. When Anthropic checked Claude's ratings against employees' own assessments, the two matched exactly 59% of the time. Two employees rating the same work matched only 35% of the time, and 97% of ratings landed within one level of each other. Anthropic reads the higher Claude-to-human agreement as a sign the model grades consistently.

Consistency and accuracy are different things. A grader built on the same data as the worker tends to share its blind spots, the sampled tasks form a frozen basket that cannot register genuinely new kinds of work, and real disagreement exists over where the borderline levels fall. Anthropic states the core risk plainly in the report.

Bar chart showing Claude agreed with employees 59 percent of the time on exact ratings, employees agreed with each other 35 percent, and Claude was within one level of employees 97 percent
The model agrees with staff more often than staff agree with each other, which reads as consistency rather than proof of accuracy. Chart: Santage · Data: Anthropic and Epoch AI
The "judge" model could make the same kinds of errors as the model it is checking.
Anthropic, Measurements for understanding the pace of AI development inside frontier labs

The compute index and the line every lab will draw generously

The third index reports that 6% of Anthropic's AI R&D compute went to safety in the week of July 13, rising to 12% for the compute tied specifically to AI-driven research. Anthropic adds an unusually candid caveat: a compute share measures only what is spent, so efficiency gains can shrink the safety percentage without any drop in effort, and a single week cannot show a trend.

The deeper problem is definitional, and Anthropic names it.

Safety research is hard to distinguish from capabilities research, and each developer will be tempted to draw the line generously.
Anthropic, Measurements for understanding the pace of AI development inside frontier labs
Donut chart showing 12 percent of Anthropic's AI-driven R&D compute went to safety and 88 percent to other work in mid-July 2026
Safety's slice of the AI-driven research compute, for a single week in July. A share this narrow moves with efficiency gains as much as with intent. Chart: Santage · Data: Anthropic Compute Allocation Index
Mental model

A self-report is only as trustworthy as the distance between the reporter and the subject. Here the reporter and the subject are the same model, so the index measures Claude's view of Claude's work. That does not make the 26% wrong. It makes it a claim only an outside instrument can confirm.

How to read a lab grading its own homework

Bringing in Epoch AI to co-develop the scale is a real step toward outside rigor, and it stops short of outside measurement, because Anthropic still generated the ratings the scale produced. That distinction will matter as other labs publish automation metrics of their own. The companies holding the instruments to measure how much AI runs their research are the same companies with a commercial interest in a large number.

For anyone outside these labs, the index reads best as a direction rather than a fixed data point. The precise 26% may not survive an independent audit, and the trajectory it traces, from a rounding error to a quarter of the work in six months, matches what Anthropic's own leaders have said in public about the pace of internal automation.

The most revealing line in the release is Anthropic's admission that the model grading the work could be making the same mistakes as the model doing it. Every automation milestone a lab reports now carries that asterisk, impressive and plausible and graded by the thing it describes, and it will keep carrying it until an outside instrument can measure what these systems actually run.

In short: Anthropic's three indexes are the first public numbers on how much of a frontier lab's own research its model runs, and they are graded by that same model. The 26% figure is a genuine milestone and a self-report at once, and until an outside instrument can verify it, every automation number a lab publishes will carry the same asterisk.

Santage is committed to independent, transparent journalism. This article is produced in accordance with Santage's Editorial Standards and aims to provide accurate and timely information. Readers are encouraged to verify information independently.