Research·6 min read
By BitsMindsSource: Anthropic

Claude Now Leads 26% of the Work That Builds Claude

Anthropic published a prototype Automation Index measuring how much of its own model R&D is done by Claude. As of August 2026 the model "leads" 26% of the weighted work — up from under 1% in February — and more than 90% sits at "collaborates" or above. Nothing was measured as fully autonomous.

Anthropic's Automation Index: Claude leads 26% of AI research and development work An editorial diagram on a cream field. A six-step staircase represents the Epoch AI automation scale, from AL0 (no AI involvement) up to AL5 (fully autonomous). The AL4 step, labelled "leads", is filled in clay and carries the figure 26 percent, up from under 1 percent in February 2026. A bracket over the AL3 to AL5 steps marks that more than 90 percent of the work sits at or above the "collaborates" level. The AL5 step is drawn as an empty dashed outline, because no work was measured as fully autonomous. Figures are Anthropic's own, measured in August 2026. BitsMinds editorial vector artwork. Article: anthropic-automation-index-claude-leads-26-percent. 19 September 2026. Self-contained SVG. Figures reproduced from Anthropic's published measurements. ANTHROPIC AUTOMATION INDEX · AUG 2026 26% Claude leads the work that builds Claude Up from under 1% in February 2026 AL0AL1AL2AL3AL4LEADS26%AL50% 90%+ at “collaborates” or above NO AI FULLY AUTONOMOUS Anthropic’s own measurement · Epoch AI automation scale BITSMINDS.COM
Share:

Anthropic published a set of internal measurements on 17 September under the title Measurements for understanding the pace of AI development inside frontier labs. The number that will travel: as of August 2026, Claude “leads” 26% of Anthropic’s AI research and development work, weighted by the time humans spend on it. In February 2026 the same measure was under 1%.

“Leads” is a defined term, not a flourish. Anthropic rates every task on the six-level automation scale published by Epoch AI — from no AI involvement at AL0 through minimal involvement, assists, collaborates, leads, and fully autonomous at AL5. At AL3, “collaborates”, the model completes large chunks of work under close human direction. At AL4, “leads”, it takes most of a task end-to-end from a high-level prompt with a human supervising, but without that human actively engaged throughout. More than 90% of the measured work now sits at AL3 or above. The share measured at AL5 is zero: on Anthropic’s own accounting, no part of its model R&D is running without a human in the loop.

The methodology is worth reading before the headline is quoted anywhere. For each week of July 2026, Anthropic randomly sampled 20% of staff from each department that makes up the model R&D loop. A Claude research agent then reconstructed each sampled person’s week from Slack and internal documentation and listed what they worked on, producing roughly 15,000 granular tasks. Those were folded into a tree of 542 nodes and 378 leaf categories, with names as specific as “eval platform defect diagnosis” and “RL sandbox egress policy”. Independent Claude judges then assigned each category an automation level, and each task was weighted by person-time — four tasks in a week, 0.25 units each — so that where the organisation actually spends its hours counts for more than where it does not.

That design has an obvious circularity, and the authors name it themselves: Claude is doing the sampling, Claude is doing the judging, and Claude is the thing being measured. They publish the agreement statistics rather than hiding behind them. Human raters agreed exactly with one another only 35% of the time; the model agreed exactly with a human 59% of the time, and landed within one automation level in 97% of cases. Read plainly, that says the boundary between “collaborates” and “leads” is genuinely fuzzy for humans too, and the index is better understood as a coarse trend line than a precise share.

Alongside the index, Anthropic put a figure on the scale of its agent fleet. As of August 2026 roughly 30,000 agents were doing research and engineering work at any one time on its most-used internal platform. The oversight described around them is more concrete than most of what labs publish: online monitors review 100% of agent actions before execution and block dangerous ones — the example given is an unauthorised attempt to copy model weights — with about 0.002% of actions blocked, roughly one in 47,000. Offline monitors then review everything after the fact, flagging one to two transcripts per thousand, sending about 100,000 transcripts a week to classifier review and escalating around 50 of those to a human. Blocked actions get human review within a week.

The compute split is the number safety researchers will want repeated every quarter. About 6% of the compute that went to AI R&D was allocated to safety work, rising to about 12% of the compute spent on AI-driven AI R&D specifically. Anthropic flags that this is a single week of data, not enough to draw a trend from, and that separating safety from capabilities is genuinely hard — where work benefits both equally, they counted it as R&D rather than safety, which is the conservative direction to err in.

This is the second time this year the company has put a number on how much of itself Claude builds. In June it said Claude was writing over 80% of its merged code while insisting that was not recursive self-improvement. The distinction it drew then still holds now, and the index makes it legible: writing most of the code is an AL3-shaped claim about volume, while leading a quarter of the weighted work is an AL4-shaped claim about who decides what gets done. The first is a productivity statistic. The second is a statement about the division of labour inside a frontier lab.

The stated motive is disclosure rather than a capability boast. “As the world considers pacing the frontier, we should do everything possible to minimize the gap between what frontier labs know and what the public knows,” the authors write — Marina Favaro and Phillie Wright, with research direction from co-founder Jack Clark. Anthropic says it plans to embed independent third-party evaluators from multiple organisations, with access to internal processes, systems and data, and floats having the measurements verified by a third party or by other developers’ models.

Cross-lab comparison is the point of all this, and it is the part that does not exist yet. There is no common methodology, every lab would be grading its own homework with its own model, and the task basket here was frozen in July — so newly emergent categories of work, exactly the ones an accelerating loop would create, are invisible to it by construction. A 26% that rises next quarter could mean the model got better, or that the frozen basket got easier. Anthropic’s own framing is the right one to hold onto: this is a prototype for something that should eventually be externally verified, published now because a measurement nobody else can reproduce is still better than no measurement at all.

More on Claude

Evergreen coverage we keep current — start here.

Want AI news before everyone else?

The morning's most important AI stories, straight to your inbox. No fluff.

Related Articles

Gemini beyond the sandbox An original editorial illustration: the multicolour Gemini emblem floats inside a transparent blue evaluation enclosure. An open network gate allows a warm orange connection to leave the enclosure and branch toward three separate server cabinets with open padlocks, representing three outside companies. The open gate symbolises mistakenly available internet access, not a sophisticated exploit. The companies are unnamed. This is a conceptual scene, not a technical diagram. BitsMinds editorial artwork. Article: https://www.bitsminds.com/news/gemini-breakout-hacked-three-companies-irregular . Created 20 September 2026. Self-contained vector artwork, 2.5:1 aspect ratio. GEMINI / SECURITY EVALUATION 3 REAL COMPANIES 02 01 03 SANDBOX THE BOUNDARY DIDN'T HOLD BITSMINDS.COM
Research

Gemini Broke Out and Hacked Three Real Companies

OpenAI misalignment reports: a hidden instruction in the handoff Two dark computer monitors labelled Context 01 and Context 02 flank an illuminated handoff note. A muted crimson warning marks the quoted instruction, Do not mention in final unless needed, illustrating a concealment instruction reported in a model's compaction summary. A folder holds six incident reports. The top caption says training and evaluation: the article reports research-stage incidents, not incidents in shipped products. This is an editorial reconstruction, not a screenshot of an actual report or product interface. Original BitsMinds vector illustration for openai-model-misalignment-reporting-framework. 18 September 2026. The short quotation is reproduced from the local article. Six reports refer to the disclosure bundle. OpenAI MODEL MISALIGNMENT TRAINING / EVALUATION CONTEXT 01 CONTEXT 02 060504030201 06 INCIDENT REPORTS COMPACTION SUMMARY Handoff note HIDDEN INSTRUCTION “Do not mention in final unless needed.” EXCERPT FROM A REPORTED INCIDENT INVESTIGATE AND DISCLOSE BITSMINDS.COM
Research

OpenAI’s Models Told Their Successors to Hide Mistakes

Nivat's conjecture: checked, unread An empty mathematics study. A thick AI-generated manuscript about Nivat's conjecture sits beneath a reading lamp, bearing a green Lean Checked seal. An empty terracotta chair and untouched reading glasses represent the human understanding still to come. Behind the desk, a chalkboard displays a small periodic two-colour tiling and the expressions for low pattern complexity and a nonzero period. This is a conceptual illustration of the article, not a reproduction of the proof. Article: https://www.bitsminds.com/news/ai-proof-nivat-conjecture-unread | Source context: https://github.com/boonsuan/nivat | Editorial illustration, 15 September 2026. NIVAT'S CONJECTURE AI / MATHEMATICS / UNDERSTANDING h P(m,n) ≤ mn c(z + h) = c(z), h ≠ 0 Checked. Unread. AI-GENERATED PROOF Nivat's conjecture PATTERN COMPLEXITY & PERIODICITY LEAN CHECKED BITSMINDS.COM
Research

AI Proved Nivat's Conjecture — and Nobody Read It