Reading elevation — this note’s pacing, drawn from its own paragraphs

Anthropic Put Numbers on AI Building Itself: What the R&D Automation Index Actually Shows

At a glance

Anthropic self-published new metrics on how much of its own AI R&D is now done by Claude. Here's what the numbers show — and the self-measurement caveat that matters.

Anthropic Put Numbers on AI Building Itself: What the R&D Automation Index Actually Shows

Defensibility check: these are Anthropic’s own self-published, self-measured metrics — not independently audited. I’m presenting the numbers with that caveat built in throughout, not as verified external fact.

Anthropic just published something genuinely unusual: actual numbers on how much AI is building AI inside one of the labs racing to build it. This follows directly from CEO Dario Amodei’s “Pace the Frontier” essay calling for industry coordination on AI safety — and it’s the concrete follow-through, not just more talk. Here’s what the numbers actually show, and what they don’t.

What Anthropic built

Anthropic created something it calls the R&D Automation Index — a prototype measurement of how much of its own AI research and development work is performed by Claude versus humans. It’s built on an automation scale developed by Epoch AI, running from AL0 (no AI involvement) to AL5 (AI operates fully autonomously, no human in the loop).

The methodology, in brief: Anthropic sampled 20% of staff across R&D departments each week in July 2026, had a Claude research agent catalog what tasks they worked on, organized roughly 15,000 granular tasks into a 542-node hierarchy, then had another Claude model rate the automation level of each task category based on gathered evidence.

What the numbers say (as of August 2026)

  • Claude “leads” 26% of Anthropic’s AI R&D work — meaning it can complete most of a task end-to-end from a high-level prompt while a human supervises. This is up from under 1% in February 2026.
  • Over 90% of R&D work is at or above “AI collaborates” — AI doing large chunks of work under close human direction.
  • Claude is not operating fully autonomously (AL5) for any measured subset of work. No fully unsupervised AI-building-AI yet, by Anthropic’s own accounting.

Alongside this, Anthropic reported two other measurements: oversight of its ~30,000 active R&D agents (100% of actions pass through automated monitoring, with roughly 1 in 47,000 decisions blocked), and compute allocation (about 6% of AI R&D compute went to safety work generally, rising to 12% for AI-driven R&D specifically).

The part that matters most: who’s grading this exam

Here’s the detail that should sit at the front of anyone’s read on this: Anthropic measured itself, using its own models, and published the results itself. The company acknowledges this directly — a Claude agent gathered the evidence, and a separate Claude model acted as judge. Anthropic reports that its model-judge agreed with human staff raters about as often as human raters agreed with each other (59% exact agreement vs. 35% human-to-human), which is a genuinely useful data point — but it doesn’t change the fundamental structure: the company building the AI is also the company measuring the AI, with no external verification yet in place.

Anthropic itself flags this as a gap: it says it plans to embed independent third-party evaluators with access comparable to internal risk teams, and that METR has previously red-teamed its offline monitoring platform. Both are real steps toward external verification — but as of this publication, they describe a plan and one prior narrow review, not an audited version of the numbers in this specific report.

Why this is still worth taking seriously

Self-measured isn’t the same as meaningless. Anthropic is one of the only frontier labs voluntarily publishing this kind of internal operational data at all, and doing so with enough methodological detail (542-node task tree, explicit automation-level definitions, published disagreement rates) that outside researchers can actually critique the approach rather than just the headline number. That’s a meaningfully higher bar than a vague “AI is helping us build AI” claim with no numbers behind it.

The company also explicitly invites comparison: it says any frontier developer could publish the same three measures using a public methodology, enabling cross-lab comparison over time. Whether competitors take that offer is the real test of whether this becomes an industry norm or stays a one-lab exercise.

The plain-English bottom line

Anthropic says Claude now does the bulk of the work on roughly a quarter of its own AI R&D tasks — up sharply from earlier this year — while remaining under human supervision throughout, with no fully autonomous AI-building-AI happening yet by its own definition. That’s a real, specific, and useful data point. It’s also a number the company that benefits from looking responsible chose to measure, define, and publish about itself, with independent verification still a stated plan rather than a completed fact. Both things are true at once, and neither cancels the other out.

Sources: Anthropic, “Measurements for understanding the pace of AI development inside frontier labs” (September 2026, primary); Dario Amodei, “We must pace the frontier” (Anthropic CEO essay); independent coverage via FourWeekMBA and Korea Times summarizing the release.