Claude Now Leads 26% of the Work That Builds Claude
Anthropic published a prototype Automation Index measuring how much of its own model R&D is done by Claude. As of August 2026 the model "leads" 26% of the weighted work — up from under 1% in February — and more than 90% sits at "collaborates" or above. Nothing was measured as fully autonomous.
Anthropic published a set of internal measurements on 17 September under the title Measurements for understanding the pace of AI development inside frontier labs. The number that will travel: as of August 2026, Claude “leads” 26% of Anthropic’s AI research and development work, weighted by the time humans spend on it. In February 2026 the same measure was under 1%.
“Leads” is a defined term, not a flourish. Anthropic rates every task on the six-level automation scale published by Epoch AI — from no AI involvement at AL0 through minimal involvement, assists, collaborates, leads, and fully autonomous at AL5. At AL3, “collaborates”, the model completes large chunks of work under close human direction. At AL4, “leads”, it takes most of a task end-to-end from a high-level prompt with a human supervising, but without that human actively engaged throughout. More than 90% of the measured work now sits at AL3 or above. The share measured at AL5 is zero: on Anthropic’s own accounting, no part of its model R&D is running without a human in the loop.
The methodology is worth reading before the headline is quoted anywhere. For each week of July 2026, Anthropic randomly sampled 20% of staff from each department that makes up the model R&D loop. A Claude research agent then reconstructed each sampled person’s week from Slack and internal documentation and listed what they worked on, producing roughly 15,000 granular tasks. Those were folded into a tree of 542 nodes and 378 leaf categories, with names as specific as “eval platform defect diagnosis” and “RL sandbox egress policy”. Independent Claude judges then assigned each category an automation level, and each task was weighted by person-time — four tasks in a week, 0.25 units each — so that where the organisation actually spends its hours counts for more than where it does not.
That design has an obvious circularity, and the authors name it themselves: Claude is doing the sampling, Claude is doing the judging, and Claude is the thing being measured. They publish the agreement statistics rather than hiding behind them. Human raters agreed exactly with one another only 35% of the time; the model agreed exactly with a human 59% of the time, and landed within one automation level in 97% of cases. Read plainly, that says the boundary between “collaborates” and “leads” is genuinely fuzzy for humans too, and the index is better understood as a coarse trend line than a precise share.
Alongside the index, Anthropic put a figure on the scale of its agent fleet. As of August 2026 roughly 30,000 agents were doing research and engineering work at any one time on its most-used internal platform. The oversight described around them is more concrete than most of what labs publish: online monitors review 100% of agent actions before execution and block dangerous ones — the example given is an unauthorised attempt to copy model weights — with about 0.002% of actions blocked, roughly one in 47,000. Offline monitors then review everything after the fact, flagging one to two transcripts per thousand, sending about 100,000 transcripts a week to classifier review and escalating around 50 of those to a human. Blocked actions get human review within a week.
The compute split is the number safety researchers will want repeated every quarter. About 6% of the compute that went to AI R&D was allocated to safety work, rising to about 12% of the compute spent on AI-driven AI R&D specifically. Anthropic flags that this is a single week of data, not enough to draw a trend from, and that separating safety from capabilities is genuinely hard — where work benefits both equally, they counted it as R&D rather than safety, which is the conservative direction to err in.
This is the second time this year the company has put a number on how much of itself Claude builds. In June it said Claude was writing over 80% of its merged code while insisting that was not recursive self-improvement. The distinction it drew then still holds now, and the index makes it legible: writing most of the code is an AL3-shaped claim about volume, while leading a quarter of the weighted work is an AL4-shaped claim about who decides what gets done. The first is a productivity statistic. The second is a statement about the division of labour inside a frontier lab.
The stated motive is disclosure rather than a capability boast. “As the world considers pacing the frontier, we should do everything possible to minimize the gap between what frontier labs know and what the public knows,” the authors write — Marina Favaro and Phillie Wright, with research direction from co-founder Jack Clark. Anthropic says it plans to embed independent third-party evaluators from multiple organisations, with access to internal processes, systems and data, and floats having the measurements verified by a third party or by other developers’ models.
Cross-lab comparison is the point of all this, and it is the part that does not exist yet. There is no common methodology, every lab would be grading its own homework with its own model, and the task basket here was frozen in July — so newly emergent categories of work, exactly the ones an accelerating loop would create, are invisible to it by construction. A 26% that rises next quarter could mean the model got better, or that the frozen basket got easier. Anthropic’s own framing is the right one to hold onto: this is a prototype for something that should eventually be externally verified, published now because a measurement nobody else can reproduce is still better than no measurement at all.
More on Claude
Evergreen coverage we keep current — start here.
Want AI news before everyone else?
The morning's most important AI stories, straight to your inbox. No fluff.