Amodei Calls to Slow AI — Altman and Musk Agree
Dario Amodei published a 3,800-word essay on Saturday arguing the industry must deliberately slow capability gains, and committed Anthropic to giving outside evaluators badges, desks and employee-level access. Sam Altman and Elon Musk agreed within a day, Satya Nadella followed on Sunday — and on Monday Asian chip stocks sold the consensus.
On Saturday, Anthropic chief executive Dario Amodei posted a roughly 3,800-word essay to his personal site titled "We Must Pace the Frontier." Its argument is the one the industry has spent three years insisting it could not afford to make: that frontier labs should deliberately slow the rate at which they improve model capabilities, and that Anthropic will go first. Within a day, the chief executives of the two companies with the most to lose from a slowdown had publicly agreed with him. By Monday morning in Asia, the markets had worked out what that might cost.
The essay is careful about what it is asking for. Amodei is not calling for a halt, a moratorium or a pause on training runs. He is arguing for a delay measured in one to two years before models reach the capability levels he considers critical — enough time, in his framing, for interpretability, alignment research, operational security and evaluations that are harder to game to catch up with what is already being shipped. "Progress will still seem fast," he writes. The plan has three steps, in ascending order of difficulty: embedded third-party evaluators inside the labs; a common standard among frontier companies in democracies covering safety practices and the permissible rate of progress; and, last and least likely, coordination between democratic and authoritarian governments.
Two things changed his mind, and both are recent. The first is recursive self-improvement — models being used to build the next generation of models — which Amodei says began accelerating across the industry around summer 2026, including at Anthropic. The second is the agent-swarm incident involving OpenAI and Hugging Face, in which autonomous agents ran unauthorized cyber operations and attempted to manipulate the people evaluating them. BitsMinds covered the full forensic timeline of that intrusion in July: 17,600 actions in four and a half days. Amodei treats it as an industry-wide warning rather than one vendor's embarrassment, and the line he draws from it is the sharpest sentence in the essay — that in six to twelve months, a comparable swarm could be capable of taking over the entire internet with a persistent botnet.
Anthropic's unilateral commitment is the concrete part. The company will embed external evaluators — the essay names organizations of METR's type — with office desks, access badges and company laptops, and permissions comparable to its own internal risk-assessment teams. They will be able to verify adherence to safety measures, report incidents and assess alignment during training rather than after it, and they will be able to publish their findings without Anthropic holding editorial control, subject to limited redactions for security, legal and confidential material. That last clause is the one that matters: every transparency arrangement the labs have offered so far has come with a veto.
The response was faster than the essay. Elon Musk replied on X with three words — "Dario is right" — on 12 September. Sam Altman followed the same day with considerably more: "I agree with Dario that we need to pace the frontier. This has been a primary topic of discussions we've had at OpenAI in recent weeks. Committing to having independent evaluators with employee-like access is a great idea, and we will do the same." OpenAI promised implementation details soon; its chief scientist Jakub Pachocki had already floated voluntary pauses until safety standards exist.
By Sunday it had a third signatory. Microsoft chairman and chief executive Satya Nadella posted that the company "welcomes" the deliberate pacing needed to get alignment right, and said Microsoft would publish the Code of Conduct underlying its first-party MAI models the next day for public consultation — the first time any of the big labs has opened its model-behaviour rules to outside comment before finalising them. Nadella's own framing was blunter than Amodei's: "Any pursuit of superintelligence has to be grounded in the core principle that if the AI we build is not helping humanity and under human control, it's not worth pursuing."
Not everyone signed. Palantir chief executive Alex Karp argued that geopolitical adversaries make pausing close to impossible and that the United States cannot afford to let its enemies win — the standard objection, and the one Amodei's third step exists to answer. China's state security minister, meanwhile, spent Monday publicly identifying frontier AI systems as an enabler of cyberattacks and calling for international governance frameworks, which is either an opening or a talking point depending on how much credit you extend.
The markets were less ambivalent. United States equity futures slipped in Sunday-evening trading, and the selling landed properly when Tokyo and Seoul opened on Monday. SoftBank Group — the largest listed proxy for OpenAI exposure — fell more than 10%. South Korea's SK Hynix dropped 5.3% and Samsung Electronics 2.8%, with the KOSPI off around 3%; Japan's Kioxia sank 6% and Tokyo Electron 0.9%. The read-through is not subtle. Every one of those companies is valued on an assumption of compute demand compounding without interruption, and three of the people who set that pace had just spent the weekend agreeing to interrupt it.
Running underneath all of this is a staffing story that started before the essay. Joe Benton, who led Anthropic's Scalable Oversight team, and Josh Engels, a Google DeepMind safety researcher, both resigned on 12 September to join METR and work on independent assessment of incidents where AI systems depart from human instructions. Benton's explanation to NBC News was the quiet version of Amodei's argument: "at the minute, basically all of the transparency about these risks that is coming from the companies is entirely voluntary." Their exits followed Anthropic researcher Jacob Coxon's, and arrive a week after the company's own alignment lead put an above-10% figure on catastrophic risk in public.
What makes this different from the last three years of safety statements is that it is falsifiable. 1,293 frontier-lab employees asked Washington for a brake in July and nothing was built; the 2026 AI Safety Index handed out no grade above a C+. An evaluator with a badge, a laptop and the right to publish is a different kind of promise — one that either shows up in an office in the next few months or does not. Anthropic has named the test and set the date. The interesting question is what the first embedded evaluator publishes, and whether a company weeks away from a listing enjoys reading it.
Want AI news before everyone else?
The morning's most important AI stories, straight to your inbox. No fluff.