On 9 September 2026, Jacob Coxon resigned from Anthropic and said publicly that the frontier labs are “gambling with our lives.” Researchers have left frontier labs with warnings before, and the pattern is familiar enough that it no longer moves anything on its own.

What happened next is not familiar. Evan Hubinger, who leads Alignment Science at Anthropic and still works there, replied on X:

Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.

That is the sitting head of alignment science at a frontier lab, writing under his own name, putting human extinction from AI above one in ten within ten years, and saying in the same breath that his employer has no plan for the problem he is employed to work on.

Separating the two statements

The departing researcher and the serving one are making different claims, and the coverage has mostly merged them.

Coxon, 27, spent roughly three years on pretraining research across OpenAI and Anthropic. His public account is that neither company is acting responsibly, that both are racing toward self-improving superintelligence, and that at Anthropic the stakes are well understood but the company believes it must get there first because no one else will act responsibly. He told the Wall Street Journal that “we're on track for a lot of the most aggressive of these scenarios where by the end of next year things could be out of control already.”

Those are a former employee's characterisations of his employer's internal reasoning. They are worth reading and they are not verifiable from outside.

Hubinger's statement is a different kind of object. It is a named, current officer of the company confirming the premise on the record and attaching a number to his own belief. He is not describing what Anthropic thinks. He is stating what he thinks, in his capacity as the person who leads the relevant work, and the number is not a rhetorical flourish.

What “no plan” does and does not mean

The phrase doing the most work is “we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.”

Read carefully, that is narrower than it looks and worse than it sounds. It is not a claim that Anthropic has no safety programme — the company publishes more alignment and interpretability research than any of its competitors, and this site has covered a good deal of it. It is a claim that the specific problem of aligning a superintelligent system remains unsolved, that no route to solving it is currently identified, and that the trajectory does not obviously lead to one.

That is not news to anyone who reads the literature. The alignment problem being unsolved is the premise of the entire field. What is new is the position of the person saying it and the absence of the usual softening. Statements of this kind normally arrive with a clause about promising directions.

It sits oddly beside the company's own filings

Three weeks earlier, in its August Risk Report, Anthropic assessed its risk of catastrophic harm from misalignment in high-stakes settings as “low” — while stating in the same document that its own arguments likely still supported the lower rating of “very low,” and that it had moved the label only to reflect increased uncertainty.

These two things are not a contradiction, and it would be sloppy to call them one. The Risk Report rating is a bounded technical assessment: a specific threat model, applied to two named covered models, under a defined framework. Hubinger's figure is a personal, all-things-considered estimate about a decade of development by the entire industry. They are answering different questions on different scopes.

They are still an uncomfortable pair to hold at once. A company can consistently believe that its current models present low catastrophic risk and that the field it leads has a greater than one-in-ten chance of killing everyone within ten years. But a reader encountering only the Risk Report would not guess the second belief was held by the person running alignment science, and nothing in the framework requires that it be disclosed there. The formal instrument and the honest estimate live in different documents, and only one of them is a filing.

Worth adding a third data point from the same month. An independent assessment in August scored five frontier developers on six control practices and found that none had fully implemented any of them, with Anthropic scoring zero on having a published containment plan. That is a different sense of “plan” — an operational runbook for a caught model, not a research programme for superintelligence — and conflating the two would be unfair. But the month has produced a consistent finding at two very different altitudes: what is written down is thinner than what is believed.

Why this is worth more than the resignation

Departures generate headlines and change little, because a former employee's incentives are easy to discount and their access is frozen at the moment they left. The standard response is that they were never in a position to see the whole picture.

That response is unavailable here. Hubinger has the access, holds the role, and is still in it. He agreed with the departing researcher's core claim rather than managing it, put a number on his own belief, and stated the absence of a plan without qualifying it into comfort. Whatever else that is, it is not a company line.

It also cannot be read as a warning from outside the tent. The most common way to dismiss safety claims is to note that the people making them do not build the systems. The person making this one runs alignment science at the company building them.

What it does not establish

A probability estimate is not evidence. Hubinger's figure is a considered personal judgement from someone unusually well placed to make one, and it remains a judgement, unfalsifiable on the timescale it describes, and one that serious researchers put anywhere from negligible to near-certain. Nothing here is a measurement.

Coxon's account of Anthropic's internal reasoning is one person's characterisation and has not been corroborated. His forecast that things could be out of control by the end of next year needs a careful comparison rather than a loose one. It sits close to the modal date in the AI 2027 scenario, which is 2027. It is substantially more aggressive than the median estimates held by the people who wrote that scenario, which ranged from 2028 to 2032, and more aggressive again than the AI Futures Project's later work, which pushes superintelligence out to 2040. So he is not an outlier against the most aggressive published scenario, and he is well ahead of its authors' own central expectations. Either way it is his view rather than a consensus.

And the framing that has travelled furthest — that an Anthropic employee said we are all going to die — is not what either man said. One said the race is reckless. The other said the risk exceeds ten percent and that there is no plan yet. Those are grimmer claims than the headline version in one respect and considerably more careful in another, and the difference is the part worth keeping.