The Pattern
Anthropic’s Responsible Scaling Policy carried one commitment that gave it teeth: if a model’s capabilities outstripped Anthropic’s own ability to control it safely, Anthropic would pause. Not measure, not disclose, not explain. Stop. On February 24, 2026, that commitment left the policy. What replaced it: public goals, a Frontier Safety Roadmap, and quantified Risk Reports, graded by Anthropic, about Anthropic.
What the word “responsible” was doing
A hard pause commitment is falsifiable. Either Anthropic paused when its own stated conditions were met, or it did not, and anyone with the public record can check. A public goal Anthropic reports its own progress toward is a different kind of object. It can be met, partially met, reframed, or quietly redefined, and the only entity positioned to say which of those happened is the one being graded. The document is still called the Responsible Scaling Policy. The mechanism that made “responsible” a checkable word is what changed.
The reason Anthropic gives
Anthropic’s own stated justification is a real argument, not a dodge: a unilateral pause by one lab does not reduce total risk if competitors with weaker safeguards keep going, a dynamic researchers call a collective action problem. If Anthropic alone stops, the frontier moves anyway, just without Anthropic’s mitigations attached to it. That is a coherent position. It is also, functionally, an argument for why no single company can be the one to hold a hard line, which is a different claim than “our commitment is still as strong as it was.”
What was happening at the same time
The Pentagon had given Anthropic a February 27 deadline to lift restrictions on military use of Claude models, according to multiple outlets, after Defense Secretary Pete Hegseth reportedly delivered that ultimatum to CEO Dario Amodei directly. The safety policy revision landed three days before that deadline. A source described as familiar with the matter told reporters the two were unrelated. This piece takes no position on whether they were connected. It notes what is not disputed: the timing, and that both events are real.
What an independent panel found
The Future of Life Institute’s Summer 2026 AI Safety Index, graded by a panel of seven outside experts, found that Anthropic, OpenAI, Google DeepMind, and Meta had each weakened or voided prior safety-pause pledges. No lab scored above a C+. The panel’s own language for the pattern: “moving the goalposts,” a practice it concluded had “undermined safety frameworks across the board.” SaferAI’s independent scoring moved Anthropic from 2.2 to 1.9 on its scale, placing the company in its “weak” category alongside the two labs Anthropic’s own marketing has long distinguished itself from.
Verdict: Misleading. Nothing about the policy change was hidden. It is a public document with a timestamp. What the name “Responsible Scaling Policy” now describes, according to independent reviewers rather than this outlet’s own judgment, is measurably less binding than what it described before, while the name and the reassurance it carries have not changed at all.
Image note: the scratch-to-reveal illustration above is AI-generated editorial art by Deceit, made for this piece. The figure is a composite; no real person, company mark, or trademark is depicted. Not documentary evidence.




