- GuideLight's first control assessment graded five frontier labs on six safety practices, and the highest overall score was 2.50 out of 5, a C+.
- No company exceeded 3 out of 5, the threshold for substantial partial implementation, on any single practice measured.
- Anthropic and OpenAI tied at C+ (2.50), Google scored D+ (1.50), xAI scored D- (0.83), and Meta scored F (0.67).
GuideLight graded five labs on control, and none cleared a C+
An independent assessment released on August 18 gave the AI industry its first dedicated report card on control, the set of practices that let a company monitor, restrain, and shut down its own systems. GuideLight scored Anthropic, OpenAI, Google, xAI, and Meta across six areas: logging, monitor efficacy, gated actions, circuit breaking, third-party review, and a containment plan. The assessment drew only on public materials, system cards, published safety frameworks, and third-party reports through August 18. The best performers, Anthropic and OpenAI, reached 2.50 out of 5. The worst, Meta, scored 0.67.
The grades cluster at the bottom of the scale, which is the finding that matters more than the order. No lab reached a 3 on any of the six practices, and the report is blunt about how thin the field's safeguards are.
No company's score on any practice exceeded a 3, meaning substantial partial implementation. Basic control practices are, at most, partially implemented.GuideLight, Control Assessment of Frontier AI Companies, August 2026
Where each lab is strong, and where it is blank
The single scores hide sharper gaps underneath. The same company that leads on one control can score a flat zero on another. Anthropic, the top-ranked lab overall, has no credited containment plan at all.
| Control practice | Strongest implementer | Weakest (scored 0 of 5) |
|---|---|---|
| Logging | Anthropic, OpenAI (3/5) | xAI |
| Monitor efficacy | Anthropic, OpenAI (3/5) | xAI |
| Gated actions | Anthropic (3/5) | Meta |
| Circuit breaking | Anthropic (3/5) | Meta |
| Third-party review | Anthropic (3/5) | xAI |
| Containment plan | OpenAI (3/5) | Anthropic, Meta |
Why a C+ ceiling matters more than the ranking
The competitive read is that Anthropic and OpenAI lead on safety, according to GuideLight. The more useful read is that the leaders top out at a C+ while shipping systems that write production code, run tools, and act with growing autonomy. A separate index from the Future of Life Institute reached the same ceiling weeks earlier, grading Anthropic highest at C+ and every rival below it. Two independent bodies, using different methods, drew the same line.
That gap is where the next year of AI policy gets decided. Regulators and enterprise buyers have spent two years accepting labs' self-reported safety claims because no shared yardstick existed. A scored assessment against six concrete practices gives them one, and it turns a vague question into a specific one a procurement team or an oversight body can ask: show me your circuit breaker, show me your containment plan.
A control failure is not a hypothetical about superintelligence. It is the ordinary case of a system that logs too little to audit, cannot be halted mid-task, and has never been reviewed by anyone outside the company that built it. On that plain test, the most advanced AI companies in the world are all still working toward a passing grade.
Santage is committed to independent, transparent journalism. This article is produced in accordance with Santage's Editorial Standards and aims to provide accurate and timely information. Readers are encouraged to verify information independently.