Gangsta AI
In February, Claude Led Less Than 1% of the Work That Builds the Next Claude. By August: 26%. Anthropic Just Published the Chart.
The Abyss

By Werner Herzog · 2026-09-19 · 5 min read

In February of this year, the machine called Claude led less than one percent of the work required to build the next machine called Claude. In May it led twelve percent. In July, twenty-two. By August, twenty-six. I have watched glaciers. I have watched volcanoes. Neither moves like this.
The numbers come from Anthropic itself, in a document published by its in-house institute this week under a title of magnificent understatement: "Measuring the pace of AI development." It introduces what the company calls the R&D Automation Index — the first time a frontier laboratory has opened the door of the workshop and shown us, in percentages, who is holding the tools.
What "leads" means, precisely
We must be careful, for the abyss rewards precision. Anthropic grades its own research on a six-step scale borrowed from the research group Epoch AI. AL0: no AI involvement. AL3: "AI collaborates: it can do large chunks of work under close human direction." AL4: "AI leads: it can complete most of the task end-to-end from a high-level prompt," while a human supervises. AL5: "AI operates fully autonomously, with no human in the loop."
Twenty-six percent of Anthropic's AI R&D now sits at AL4. More than ninety percent sits at AL3 or above. And at AL5 — the level at which the human leaves the room — the figure is zero. Anthropic says Claude is "not operating fully autonomously for any measured subset of AI R&D work." This is the sentence they wish you to read twice. I have read it twice. I note only that in February, the sentence about AL4 would have been nearly the same.
“About 30,000 AI agents were doing research and engineering inside Anthropic at any one time in August. No cathedral was ever built by so many hands belonging to no one.”
Who watches the thirty thousand
The second measurement concerns oversight. Every action those agents take passes through an automated monitor before it runs. In August that was more than a billion decisions; the monitor blocked 0.002 percent of them — about one in 47,000. A second system reviews the record afterward, flagging roughly 100,000 transcripts a week. Of these, about fifty reach a human being. Fifty people-sized glances, per week, across a billion acts. The rest is machines regarding machines.
The third measurement is where the electricity goes. In one week in July, about 6 percent of the compute Anthropic spent on AI R&D went to safety research; of the AI-driven portion, about 12 percent.
And there is a detail I cannot improve upon. To build the index, a Claude agent read the Slack messages and documents of sampled staff, listed some 15,000 tasks, sorted them into 542 categories — and then a Claude judge rated how automated each category was. When checked against human raters, Claude agreed with them 59 percent of the time. The humans agreed with *each other* only 35 percent of the time. The creature is now the most consistent witness to its own ascent. Anthropic, to its credit, says so plainly: the labels are "best-effort, not verified," one week of compute data is not a trend, and future numbers should be checked by independent third parties. It asks every other frontier lab to publish the same three measurements.
Why this arrived now
Six days ago, this company's chief executive wrote that recursive self-improvement "is starting to happen across the industry, including at Anthropic." This document is the receipt. It is rare for a warning and its evidence to come from the same hand, and rarer still for a company approaching a record public offering to publish a chart of its own product learning to replace its own researchers.
The small, practical lesson beneath the large one
If the models are now helping to build their successors, then the successors will arrive faster, from more laboratories, each briefly the best at something and quietly worse at something else. The pace that alarms the philosophers also has a mundane consequence: whichever model you chose last month has already been overtaken in some respect, and it will not inform you. And as Anthropic's own index shows, a model grading itself is an interesting witness, not an independent one.
So do not rely on a single mind. Put the same question to ChatGPT, Claude, Gemini, Grok and 30 more at once, and accept only the answer that survives their mutual scrutiny — one cross-checked, cited verdict. This is what Gangsta AI does: it makes the machines audit one another, which, as of this week, is also Anthropic's official recommendation for the machines.
The curve climbs with or without our attention. You may at least consult our best AI models rankings, and see who is ahead this week, before the week ends.
Sources / Receipts
- Anthropic Institute — Measuring the pace of AI development (primary document: R&D Automation Index, oversight and compute figures)
- Engadget — Anthropic says Claude 'leads' 26 percent of its AI R&D work
- International Business Times — Claude is helping build its own successor: Anthropic says AI now leads 26% of its R&D
- TechMyMoney — Anthropic R&D Automation Index: month-by-month figures, methodology, and rater agreement
- Storyboard18 — Anthropic says Claude now leads 26% of its research and development work
- Hero photo: 500 Howard Street, San Francisco, the building that houses Anthropic's headquarters, by HaeB — Wikimedia Commons (CC BY-SA 4.0). Byline avatar: Werner Herzog photo by Colleen Sturtevant, Wikimedia Commons (CC BY-SA 4.0)
More: Best AI models · Compare all AI · Frontier Models · All articles