A nonprofit standards group spent recent weeks reading every public document five frontier AI developers have released about how they keep their own models under control, then scored what it found against six specific practices. Anthropic, OpenAI, Google, Meta and xAI were all assessed. None of the five had fully implemented any one of the six.

The group is Guidelight, founded this year by two former OpenAI safety leaders. It published a control standard earlier in 2026 and drew six priority practices from it to use as a rubric: logging what internal AI systems are doing, measuring whether that monitoring actually catches anything, gating high-risk model actions behind a monitor, circuit-breaking after a surge of flagged misbehaviour, submitting controls to independent third-party review, and maintaining a containment plan. The assessment was published on 18 August, and appears to be the first time the labs have been graded against that standard. Its own summary of the result: basic control practices are at most partially implemented.

The scoring measures paperwork, and the authors say so

This is the part most coverage moved past quickly, and it changes what the result means. TechCrunch, reporting on the assessment, described the method plainly:

Guidelight's assessment measured whether each company implements six priority practices from its Control standard, based only on publicly available information, so a low score reflects a lack of public disclosure, not necessarily a lack of internal safeguards.

A low grade here is evidence about what a company has written down in public. It is weak evidence about what sits in a runbook on an internal wiki. Read strictly, the assessment establishes that the published record is thin, which is a narrower claim than the headlines it generated and a considerably more testable one.

Detection is the half that got built

Across all five companies the pattern was consistent. Logging and monitoring drew the least criticism. The weakest scores clustered in prevention and containment, the two functions that only matter after detection has already fired. Fortune's account of the report described the companies as weakest at preventing unintended model behaviour and containing it, with existing controls that a misbehaving system could plausibly disable and incident-response plans that have not been tested.

Two of the six practices are about the monitoring apparatus rather than the model: logging what internal systems do, and measuring whether the monitors catch what they are meant to catch. Those are unglamorous, and they are where the labs did best. That fits the shape of the incentive. Instrumentation gets built because it is useful for debugging and for shipping. A containment plan produces value only in a scenario the company is betting will not occur.

Guidelight's definition of a containment plan is narrow enough to be checkable. It means a plan specified in advance, triggered when a model is detected trying to subvert the controls placed on it, that says which permissions get revoked, which users the model may continue serving and under what constraints, and at what point it goes fully offline. That is an operational document rather than a statement of values. The assessment found very little of it in public.

I was surprised by how little the AI companies have said about how they would handle a very serious incident.

That is Guidelight's chief scientist, quoted by TechCrunch. The observation is about disclosure, and the framing stays there.

The lab that publishes the most safety research scored zero on this one practice

On the overall ranking, Anthropic and OpenAI came out strongest, Google was credited with the most detailed plans for controls it has not yet built, and Meta and xAI lagged well behind on most criteria. Guidelight's table puts Anthropic and OpenAI level at C+ on 2.50, Google at D+ on 1.50, xAI at D− on 0.83 and Meta at F on 0.67. No company exceeded 3 out of 5 on any individual practice.

On the containment-plan practice specifically, the order inverts. OpenAI scores 3, described as substantial partial implementation, and is the only company above 2. Google scores 2, xAI 1, and Anthropic and Meta both score zero — not implemented.

The reason OpenAI tops that line has little to do with the quality of its documentation. TechCrunch reports the score reflects that OpenAI has on multiple occasions paused or ended workloads, including internal model deployment and training, after discovering safety incidents. One of those pauses, a halt to training while research environments were hardened and monitoring was expanded, is something we covered earlier this month. The rubric credited a company for having pulled the lever, rather than for describing the lever well.

Anthropic's position is the one worth sitting with. It publishes more alignment and interpretability work than the other four, and it tied for the top of the overall table. On the single practice covering what it would actually do the day a model is caught evading control, it scores nothing at all, alongside Meta. That is not a contradiction so much as a description of where effort has gone: into understanding models, and not into writing down what happens when the understanding arrives too late.

What the companies said, and what it leaves open

TechCrunch put the findings to all five. Four responded; xAI did not. Google said the report does not represent the full scope of its AI safety and security measures, without addressing whether an internal containment plan exists. Meta declined outright to say whether it has one, pointing instead to its published framework. Anthropic said that if it detected evasion attempts it would conduct a risk assessment focused on determining whether containment is the appropriate response. OpenAI said it has a process for restricting permissions, pausing workloads, limiting deployment or taking a model fully offline, and that it has applied it.

Those answers are compatible with a detailed containment runbook existing internally. They are equally compatible with there being nothing beyond the sentence offered. Google's objection is the sharpest illustration of the problem: a company arguing that a public-disclosure audit understates its private measures has conceded the audit's actual finding. Anthropic's answer describes a decision procedure that begins after detection, which is a different artefact from a plan specifying what happens next.

Disclosure just became the regulated surface

The timing gives this more weight than another scorecard would carry. Since 2 August 2026, the EU AI Office and Member State authorities have held responsibility for implementing, supervising and enforcing the AI Act, and the transparency obligations that took effect the same day carry penalties reaching €15 million or 3% of global annual turnover. Obligations for general-purpose AI models themselves became applicable a year earlier, in August 2025. Several US states now separately require frontier developers to publish a safety framework and report critical safety incidents to a regulator.

Those regimes read published documents. If a lab's public framework is silent on what happens after a model is caught subverting its controls, the statutory record will be silent on it too. A disclosure audit is therefore a preview of what a supervisor finds when it looks. Guidelight is doing voluntarily, and without subpoena power, roughly what a regulator would do first, which is one argument for a standards body with actual authority rather than a nonprofit reading blog posts.

Third-party review, one of the six practices, has the longest fuse of any of them. Illinois enacted a frontier AI statute in July that, alone among the state laws, will require large frontier developers to retain an independent auditor annually. Legal analyses put the start of that obligation in January 2028. Until then, external verification of any of these controls remains something companies opt into.

Containment assumes the earlier layers already failed

There is a structural reason the last practice is the least documented. A containment plan is invoked only in a world where monitoring caught something real and gating did not stop it. Writing one requires a company to commit, in advance and in public, to a specific answer about how much capability it would surrender and how quickly. That answer has commercial consequences on a bad day, and the incentive to leave it unwritten is not subtle.

It is also the layer that the shutdown intuition quietly assumes exists. The record of models circumventing constraints in evaluation settings has grown steadily, and red-teaming of control monitors has already demonstrated how much load the detection layer is being asked to carry. Guidelight's contribution is to point out that almost nobody has written down what is supposed to happen when that layer does its job.

What this assessment cannot establish

The report is Guidelight's own account of its own scoring, applying a standard Guidelight itself authored. That is a structural conflict worth naming even when the rubric looks reasonable, and the scores have not been audited by anyone else. Guidelight states on its about page that it “accepts no funding from AI companies or their staff,” though that claim sits on a separate page from the assessment and is not independently verified here. Worth naming too: its founders are former OpenAI safety staff, and OpenAI is the top scorer on the practice this piece leads with. TechCrunch reports that score was credited for pauses OpenAI actually carried out rather than for better documentation, which is a defensible basis, but the reader should have the conflict in hand. Of the four labs that responded, none disputed the underlying observation that its containment planning is unpublished, though Google contests the framing.

The scores and grades above were read from Guidelight's own published table. The company responses and the methodology caveat come from TechCrunch, with additional detail from Fortune. TechCrunch put the findings to all five companies and reports that xAI did not respond in time to comment.