All News
anthropicclauderecursive-self-improvementai-safetyagentsmetrics

Claude leads 26% of Anthropic's own AI research. The index that measures it was built by Claude

Anthropic's R&D Automation Index says Claude leads 26% of its AI research, up from under 1% in February. The tasks, the ratings and the judge are all Claude.

Vlad MakarovVlad Makarovreviewed and published
7 min read
Claude leads 26% of Anthropic's own AI research. The index that measures it was built by Claude

Anthropic published a number on Thursday that its chief executive has spent the past week arguing should inform how the world handles frontier AI: how much of the company's own research is now done by AI. The answer, measured against weeks of internal work records, is that Claude "leads" 26% of Anthropic's AI research and development, meaning the model completes most of a task end-to-end from a high-level prompt while a human supervises. In February, the same chart shows, that share was under 1%. The post also discloses roughly 30,000 agents running on the company's systems at any one time, and a safety share of research compute in the single digits.

What the number actually says

The figure comes from a prototype the company calls the Anthropic R&D Automation Index, one of three measurement tools laid out in "Measurements for understanding the pace of AI development inside frontier labs," published September 17 on Anthropic's Institute site. Marina Favaro and Phillie Wright co-authored it, with editorial support from Santi Ruiz, Adam Farina and Sarah Pollack; Jack Clark provided research direction.

The index borrows its scale from the Automation Level taxonomy developed by Epoch AI, which runs from AL0 (no AI involvement) to AL5 (the AI operates fully autonomously, with no human in the loop). Two rungs matter for reading the headline. At AL3 the model "collaborates": it does large chunks of work under close human direction. At AL4 it "leads": it handles the whole task from a prompt and the human only reviews and approves the result. Three findings sit underneath the 26%. Claude is not operating fully autonomously on any subset Anthropic measured. The share of work at or above "collaborates" is over 90%. And the "leads" line climbs from effectively nothing at the start of the year, which is how Bloomberg's Shirin Ghaffary reported it.

How the index is built

No one can hand-list every research task at a frontier lab, so Anthropic assembled the catalogue bottom-up from work records, including Slack and internal documentation. For each week in July 2026 it randomly sampled 20% of staff in every department that makes up the model R&D loop. A Claude research agent read each sampled person's week and listed what they worked on, producing a flat list of about 15,000 granular tasks. A second Claude pass organised those into a hierarchy of 542 nodes, of which 378 are leaves with names like "eval platform defect diagnosis and fixes," "RL sandbox egress and network policy" and "serving incident postmortems."

That tree is then frozen, so every later measurement runs against the same basket of work. Each node is weighted by person-time: every person gets one unit of weight per week, split evenly across the tasks they touched during that week, the crude approximation that decides which categories count for most.

Who grades the grader

The ratings are not human assessments that Claude summarised. A Claude agent researches how each kind of work is done, and an independent Claude judge assigns an automation level from the evidence. Anthropic checked it by asking staff who own the relevant work areas to rate automation there without seeing what the models had gathered. Model-versus-human exact agreement was 59%; human-versus-human agreement was 35%; ratings landed within one level of each other 97% of the time. The company concedes there is "real room for disagreement on borderline cases, such as where exactly 'AI collaborates' ends and 'AI leads' begins."

It also names the obvious objection directly:

"we're using our own models to evaluate our systems, which could mean that the 'judge' model could make the same kinds of errors as the model it is checking."

The proposed fix is verification by a third party or by other developers' models. Two limits are self-reported. The frozen basket describes the work humans were doing in July, not whether new kinds of work are appearing; comparing an alternate January 2026 basket found no rise in novel tasks; Anthropic plans to rebuild and re-version the index periodically.

Thirty thousand agents, fifty escalations a week

The oversight numbers cover one system: the most-used internal agent platform, where roughly 30,000 agents were doing research and engineering work at any one time in August. Anthropic splits its monitors into two kinds, and both report total coverage of agent activity.

MetricOnline monitorsOffline monitors
Coverage100% of actions screened before execution100% of actions ingested after the fact
TimingUsually seconds; humans review blocked actions within a weekFlags about 100,000 transcripts a week, then classifiers
Escalation0.002% blocked (about 1 in 47,000)About 50 top-priority flags a week reach humans

The blocking rate comes from more than a billion decisions by research and engineering agents during August; the offline side flags one to two transcripts in every thousand. Two scaffold choices are meant to make this auditable. Agents get persistent individual identities that survive model upgrades, so one agent treats another's output as a claim to check and any action can be traced per agent over time. They also talk over a shared open channel, so messages can be cross-referenced with transcripts and agents do not end up "playing telephone."

Six percent

For compute, Anthropic took a single week, July 13 to 20, and sorted every workload into buckets. About 6% of the compute going to AI R&D was allocated to safety work, and about 12% of the compute going to AI-driven AI R&D was. The company calls these deliberately conservative: a token that advanced capabilities and safety equally was counted as R&D, and the safeguards classifiers that filter model outputs are excluded from the numerator entirely. Classification was itself done by a prompted Claude reading run metadata and code across roughly 10,000 runs a week, sampled at about 14% and weighted toward the largest runs.

Two caveats matter more than the number. Compute share measures what was spent, so a more efficient safety classifier lowers the safety fraction without anyone doing less safety work, and the boundary between safety and capability research is drawn by the company being measured. Anthropic's own framing here is unusually blunt: "each developer will be tempted to draw the line generously. The burden of proof should sit with the developer."

A measurement and an ask shipped together

The post is not only a disclosure. It is also an exhibit for the argument Amodei made a week earlier in his pacing essay, which this site covered when rivals and Congress reacted to it. Anthropic says plainly that these numbers would be expected to shift if there were coordination on pacing the frontier, and that it plans to embed independent third-party evaluators with access comparable to its internal risk assessment teams — the verifiability step at the centre of that proposal. Read the release next to the safety-coordination talks between labs and the two documents pull in the same direction: publish the internal measures, then hand someone outside the company the keys to check them.

What would make the numbers comparable

Nothing here is yet a cross-lab measurement. No other developer publishes an automation index, an agent escalation rate or a safety compute share, and the methodology that produces Anthropic's figures was designed inside Anthropic. Cross-lab comparison needs a shared methodology and judges that are not the vendor's own models, which the post acknowledges. Until that exists, the honest reading of a single upward line is narrow: the work that existed in July is being handed off more often, and one company scored the handoffs itself. It is worth remembering that a rising automation index is not by itself recursive self-improvement — the recent Reddit misreading of a Google preprint is a useful reminder of how fast that leap gets made. What would settle it is the thing Anthropic has promised but not yet delivered: an outside evaluator with the same access as an insider, publishing what it finds, including when the number falls.

Related Articles

Scroll down

to load the next article