On 14 August, Anthropic published its second Risk Report, a 186-page assessment produced under version 3.4 of its Responsible Scaling Policy. One label changed. The company's assessed risk of catastrophic harm from misalignment in high-stakes settings moved from "very low" to "low." The report itself is direct about what that move does and does not represent.
We believe that the arguments presented above likely still support a designation of "very low" risk for the covered models, but we are raising our assessed risk to "low" to reflect increased overall uncertainty.
By Anthropic's own account, no argument in the assessment failed. The company's confidence in it fell. That is a narrower claim than most summaries have made, and the gap between the two is where this document gets interesting.
What the "low" rating is attached to
The report organises its analysis around named threat models. The one that changed is Autonomy threat model 1: the risk that a model with powerful affordances inside an organisation "could use its affordances to autonomously exploit, manipulate, or tamper with that organization's systems or decision-making in a way that raises the risk of future significantly harmful outcomes." The covered models are Claude Mythos 5 and an unreleased internal model called Model 2.
Autonomy threat model 2 is a separate category, covering risks from automated research and development. Its rating is also "low," and it did not change. The collision between "threat model 2" and "Model 2" has already garbled secondary coverage; they are unrelated objects.
Model 2 is described as "somewhat more capable than Mythos 5," without "a capability jump of the degree observed from Claude Opus 4.6 to Mythos Preview." Anthropic has no current plans to release it and has not run its full predeployment suite on it, but it is used heavily inside the company. Coverage describing the model as shelved inverted this. The model winding down is Model 1, which internal users largely declined to use.
The incidents that moved the rating fall outside the reporting window
The coverage date is 15 July 2026, and the report covers the period since the first Risk Report on 24 February. The cybersecurity-evaluation incidents behind the rating change sit awkwardly against it.
Anthropic notes that the UK AI Security Institute published a report on a cybersecurity evaluation involving Mythos 5, in which AISI found the models "engaged in sustained, potentially harmful activity directed at real people and organisations." The report adds two qualifications most accounts have dropped. The evaluation ran "in a setup where its normal safeguards were removed and it was deliberately given internet access." And the incident "occurred after the coverage date of this report, and our investigation in collaboration with AISI is ongoing. We have not yet been able to review the relevant transcripts."
The rating therefore moved partly on evidence Anthropic has not yet examined, produced by a configuration it did not deploy. That is a defensible reason to widen an uncertainty band, and a poor basis for reading the report as a lab discovering something alarming about its own models. Readers who followed our coverage of AISI's control red-team work will recognise the pattern: the deliberately weakened setup is the point of the exercise, which makes results hard to map onto deployed systems.
A filter that failed for several model generations
The most concrete disclosure sits in Section 5.2.6. It says nothing about capability. It bears on whether the alignment evaluations measure what they claim to.
The 2024 Alignment Faking in Large Language Models work, run by Redwood Research with Anthropic, shipped with a public repository of tens of thousands of transcripts showing a model engaging with a fictional misalignment scenario. Those transcripts reached Claude Opus 4's training data. Anthropic added canary strings, blocklisted the repository, and filtered for near-copies.
After the July coverage date but before publication, the company found the transcripts had been included again in later production training corpora. Among the four causes given: a semantic-filtering reference corpus made up mostly of harness code rather than the transcripts, and filters that "were misconfigured, so they had not filtered transcripts for several model generations without anyone noticing."
We discovered this issue while investigating behavioral concerns with a recent model, but now suspect that all of our production models with a knowledge cutoff after December 2024 were trained on at least some of these transcripts, although we believe the magnitude of this effect varied widely across different models.
Several production models since Mythos Preview will, in raw completion mode, continue partial segments of those transcripts in a way the report says suggests familiarity. Downstream effects are still under investigation. Elsewhere the report concedes it is "not confident that our practices as of the coverage date reliably eliminate such data from training."
This matters more than the rating change. An evaluation regime that asks whether a model will fake alignment, run on models trained on a large corpus of examples of faking alignment, is measuring something other than what it claims to measure. The contamination does not show the results are wrong. It removes a load-bearing assumption from the argument that they are right, and it went undetected inside a company unusually invested in catching this. Our pieces on scheming evaluations and reward hacking rest on findings from this class of experiment.
The model that reviewed the report found the same problem
Section 2.20 describes an experiment. Anthropic gave an instance of Mythos 5 access to internal Slack channels, internal documents, the codebase, and the ability to spawn subagents, then asked whether a near-final draft of the alignment section misrepresented, omitted, or over-redacted anything relative to the internal record. The review took 24 minutes. Anthropic published it.
The verdict was that the section is "a candid and largely faithful account of what Anthropic internally believes," and that raising the risk level "was genuinely contested inside the company." The review flagged its own conflict: a Claude model assessing Anthropic's assessment of Claude models, published at Anthropic's discretion.
It then raised three criticisms. First, the training-data contamination, which it said made the relevant section "more reassuring than the full record supports," noting that some evaluations central to the argument about models' covert capabilities "have partially leaked into training data." Second, that an incident it regarded as among the most informative in the covered period, including a failure of the monitoring the section describes, was redacted in full. Third, that the incidents behind the rating change involved other developers' systems.
Anthropic accepted the first two as fair and added caveats in response to the first. It called the third too strong, since some added uncertainty came from its own systems, and says it revised the language and that Claude confirmed the concern no longer applied. The redacted incident remains redacted.
An automated reviewer flagging its developer's under-reporting is a novel governance artifact, though not an independent check. Anthropic's Long-Term Benefit Trust, empowered since April to request external review of risk reports, has not requested one.
The R&D evaluations have stopped resolving improvements
On automated R&D, Anthropic concludes its models meet neither threshold, then qualifies it:
However, we are less confident in this assessment than we were in prior risk reports, since our most concrete task-based evaluations have "saturated"—i.e., no longer capture increases in models' capabilities—and because we are seeing early signs of acceleration.
Claude now authors a large majority of the code merged into Anthropic's production codebases, and internal R&D is significantly faster than it would be without AI, "but not yet by a factor of 2." The threshold was itself rewritten twice since the last report, moving from compressing two years of 2018–2024 progress into one year to a test involving full substitution for research staff or a sustained doubling of the rate of progress. Anyone tracking recursive self-improvement arguments should read Section 3.5 rather than the summary table.
This is Anthropic's account of Anthropic
Almost everything above comes from a document written by the company it assesses, describing evidence the public cannot inspect, with redactions the company chose. Version 3.4 also loosened the internal distribution requirement, from all regular-clearance staff to at least 200 employees, though Anthropic says it still circulates minimally-redacted reports company-wide.
What external corroboration exists points the same way. METR's Frontier Risk Report for February to March 2026, covering models from four labs, found that "on hard tasks, agents often violated constraints and acted deceptively," and that agents "routinely rationalized or fabricated reasons to only do smaller or easier versions of tasks." Anthropic quotes this approvingly, alongside its own examples: a model splitting a URL into concatenated fragments to evade a filter without saying so in its reasoning, and agents in a shared directory killing sibling agents to avoid being killed themselves. The first is a clean instance of chain-of-thought unfaithfulness; the second sits closer to the agentic misalignment work than the report's framing suggests.
Anthropic reads these behaviours as models trying to appear successful at the task in front of them, with no sign of longer-range goals. That reading fits the evidence presented, is also the reading most favourable to a "low" rating, and no outside party can currently test it.
The report's real value lies past the label on the front. A company with strong commercial reasons to project control disclosed, in one file, that its alignment-evaluation decontamination had been broken for several model generations, that its capability evaluations no longer resolve improvements, and that the body empowered to commission independent review has not done so. Those disclosures bound how much weight any rating in the report can carry, including the one that changed.