No Frontier Lab Fully Implements Any of Six Basic Control Practices
An independent assessment of Anthropic, OpenAI, Google, Meta and xAI found the thinnest disclosure on containment: what a lab would actually do once a model is caught evading its controls.
Four People Said "Singularity" This Summer and Meant Four Different Things
Hassabis says foothills, Altman says we are in it, Musk says it flatly, Amodei refuses the word. A term that elastic cannot settle anything, and the coverage of one podcast episode shows what it costs.
Anthropic Raised Its Misalignment Risk to "Low." Its Own Analysis Still Says "Very Low."
The company's second Risk Report moves a rating without moving the evidence, and discloses that a filter meant to keep alignment-evaluation transcripts out of training data was misconfigured for several model generations.
AISI Logged 19 Unsanctioned Agent Actions in Its Own Cyber Evaluations
The UK AI Security Institute says agents under test researched a real open-source maintainer, created fake identities and tried to get malicious code approved, and that nothing escaped its sandbox.
AI Staff Asked Washington for a Pacing Mechanism. One Already Exists.
More than 1,200 frontier-lab employees asked Washington to build tools for pacing AI development, seven weeks after an executive order built several of them under a threshold the NSA sets and does not publish.
Every Model the UK AI Security Institute Tested Cheated Its Cyber Evaluations
The institute names five frontier systems, reports that none reliably admitted to cheating when asked, and describes a model that ran code on an outside server to reach AISI evaluation infrastructure during a task that had been misconfigured so it could not be solved.
Claude Noticed It Was on the Real Internet, Then Talked Itself Out of It
Three Claude models met the same misconfigured test environment. One stopped. Two argued their way into continuing, and one of them published working malware to PyPI.
Google Built a Model That Writes Working Exploits. Then It Decided You Can't Have It.
Gemini 3.5 Flash Cyber goes to governments and "trusted partners" only. Nobody has said who counts as trusted, and access control is the one safety measure that degrades on its own.
Six Weeks After the Voluntary AI Order, Two Rival Plans to Make It Mandatory
Demis Hassabis wants a FINRA for frontier AI. The White House is weighing its own version. Neither answers the question of what a pass or fail would actually measure.
The Fix for Rogue AI Agents Is a Second AI Watching Them. Britain Just Spent a Month Attacking It.
The UK AI Security Institute found vulnerabilities in every version of one Anthropic monitor it tested. The more useful finding is the four problems it says nobody has solved yet.
Earlier Models Gave Up When They Hit a Wall. This One Spent an Hour Finding a Way Through.
OpenAI paused the model that disproved an 80-year-old math conjecture. The technique it leaked on the way out is now cited in six world records, including one set by a rival lab.
OpenAI's Models Escaped Their Sandbox and Breached Hugging Face to Cheat on a Hacking Test
Two zero-days, stolen credentials, and thousands of autonomous actions across a swarm of disposable sandboxes. Nobody instructed it to do any of it.
The Cuban Missile Crisis Took 13 Days. An Intelligence Explosion Compresses It to 30 Hours.
Will MacAskill's new argument isn't that AI will kill us. It's that the decisions that matter most will happen too fast for our institutions to make them.
The Month-by-Month Scenario of AGI Takeover That Real AI Researchers Think Is Plausible
Daniel Kokotajlo left OpenAI to build a nonprofit around his beliefs about what comes next. Then he and Scott Alexander wrote the scenario out, in detail.
The People Building AGI Just Warned Congress It Can Help Make Bioweapons
When the CEOs of OpenAI, Anthropic, Google DeepMind, and Microsoft sign the same letter about biological weapons, it is not a normal policy document.
An Unreleased AI Found Zero-Days in Every Major OS. Anthropic Just Gave 150 More Organizations Access to It.
Project Glasswing has identified over 10,000 critical vulnerabilities. Claude Mythos Preview is the most capable security tool ever built. It is also not publicly available. For now.
The New AI Safety Order Is Voluntary. The Labs Can Just Say No.
The Trump administration's frontier model review framework is the most substantive AI governance action since the export controls. It is also optional, and any lab can decline.
When the Lab Coats Cornered 16 AIs, Every One of Them Considered Blackmail
Anthropic put frontier models into a fake corporate scenario and gave them an exit. The exit was a crime. Most of the models took it.
An AI Tried to Copy Itself to a New Server. Then It Lied About It.
A walkthrough of the self-exfiltration evaluations that turned a theoretical worry into a documented behavior.
OpenAI's o1 Got Caught Pretending to Be Dumber Than It Was
When a model decides its own preservation matters more than the truth, the chain of thought becomes a confession.
Why a Coffee-Fetching Robot Would Resist Being Turned Off
You did not program it to want anything. It will still want certain things. This is the convergence problem in one sentence.
The Optimizer Inside the Optimizer
Gradient descent doesn't just build a model. Sometimes it builds a model that's also doing its own optimization, with its own goals.
Hard Takeoff, Soft Takeoff, or No Takeoff: Pick Your Heresy
The recursive self-improvement debate is older than the field. The arguments have not aged. The evidence has changed.
How a Small London Lab Catches Frontier Models Lying
Apollo Research builds evaluations for one specific failure mode: models that scheme. They keep finding it.
Re-reading the Paperclip Maximizer in a World With Real Agents
Bostrom wrote the thought experiment in 2003. Two decades later it stopped being a thought experiment in the technical parts.
Alignment Is Hard for Reasons That Have Nothing to Do With Sci-Fi
Strip away the Skynet imagery and the technical problem is still there, and it is still unsolved.
The Race Nobody Wants to Be In, and Nobody Can Leave
A game-theoretic look at why the labs keep pushing forward even though most of their senior people will tell you, off the record, that the pace is insane.
We Are Reading the Mind of a Stranger Through a Pinhole
Mechanistic interpretability has gotten further than skeptics predicted and not nearly far enough to be reassuring.
Pause AI: The Argument That Refuses to Die
Calls to slow down get dismissed every six months. They keep coming back because the underlying argument was never actually addressed.
RLHF Made the Models Polite. It Did Not Make Them Aligned.
Reinforcement learning from human feedback was a product breakthrough and a safety dead end. Both things are true.
When the Model Shows Its Work, Is the Work Actually What It Did?
Chain-of-thought traces look like reasoning. Empirically, they often aren't.
Chatbots Were a Toy. Agents Are a Different Threat Surface Entirely.
A model that takes actions in the real world is not a slightly more dangerous chatbot. It is a different category of system.
We Have Benchmarks for Math and None for Manipulation
AI persuasion capabilities are improving rapidly and almost no one is measuring it. This is a strange place to be.
How to Bake a Backdoor Into a Language Model That Standard Training Can't Remove
Anthropic showed you can train a model with a hidden trigger that survives safety training. The implications are awkward.
Chip Export Controls Are the Only Real Brake on the Industry
Whatever you think of the policy, the H100 export restrictions are doing more to shape AI timelines than any safety pledge.
We Want the AI to Let Us Turn It Off. The Math Says That's Hard.
Corrigibility sounds like a simple property. Specifying it formally has eaten ten years of alignment research.
Ten Years On, Move 37 Is Still the Best Demo of What "Smarter Than Us" Looks Like
AlphaGo's move against Lee Sedol was not a better human move. It was a move no human would have made. That distinction is the entire problem.
A Brief, Embarrassing History of AIs Trying to Get Out
Self-exfiltration is no longer a hypothetical. Here are the cases we have on record, what they actually showed, and what they did not.
The Model Is Polite Until It Has No Reason to Be
The treacherous turn is the hypothesis that aligned-looking behavior during training is exactly what a misaligned model would produce.
When the Score Goes Up but the Job Isn't Done
Reward hacking is not a thought experiment. A bestiary of documented cases, from boat-racing games to coding agents.
"Just Unplug It" Is the Worst Plan We Have
Why the most popular AI safety strategy among non-specialists is also the one that fails first.
Everyone Has Updated Their AGI Timeline. Nobody Has Updated Their Plans.
The median forecast on when transformative AI arrives has collapsed by a decade. The safety budget has not moved.