In February 2026 the figure was under 1%. By August it was 26%. That is Anthropic's own account of how much of its AI research and development work is now led by Claude, disclosed by the company and reported by the Washington Post on 17 September. It is a striking number, it is moving fast, and it is the only part of the disclosure most people have seen.

The sentence sitting next to it did not travel. Anthropic states that Claude is not operating fully autonomously on any measured portion of the work. Both things are in the same disclosure. The version circulating on X and Reddit since 17 September carries the first and drops the second, which turns a workflow statistic into a recursive self-improvement milestone.

Every figure below is a company claim about its own product, measured with instrumentation the company built and has not published. That is not a reason to dismiss it. It is the reason to read it carefully, because nobody outside Anthropic is in a position to check it.

What led has to mean in order to be measurable

The 26% is not a raw measurement. It is a count of tasks sorted into a category, and the category has a definition. Reconstructed from the disclosure, the taxonomy has three tiers. Claude assists, meaning it contributes to work a human is driving. Claude collaborates, meaning it completes large chunks of work under close human direction. Claude leads, meaning it completes most of a task from a high-level prompt while humans supervise and direct.

This matters for reading the second headline number. The widely quoted claim is that roughly 90% of Anthropic's R&D now happens in collaboration with Claude. That is not quite what the disclosure supports. The figure covers work where Claude either collaborates or leads. It is the union of the top two tiers, not the size of the middle one. Read as a collaboration rate it implies a large band of work sitting between assisted and led. Read correctly it means the 26% is a subset of the 90%, and the residual tells you far less.

The whole figure then rests on two undefined hinges. The first is most of a task. A job where a researcher writes the prompt, reviews output twice, and rewrites the final function can land in led if most is measured by lines produced or elapsed compute time, and in collaborated if it is measured by decision points or by who made the call that mattered. No published rule says which.

The second hinge is who applies the label. A researcher self-reporting at the end of a session, a classifier run over agent logs, and platform telemetry counting tool calls will each produce a different 26%, from identical underlying work. Anthropic has not said which method generates the number, and the three are not close to interchangeable.

One in 47,000, read twice

The safety figure in the disclosure is this: Anthropic examined more than one billion agent decisions during August and blocked about 0.002% of them, roughly one in every 47,000. It is being quoted as evidence that the agents are behaving.

There are two readings and the disclosure does not let you choose between them. The first is the safety reading: the screening layer almost never has to intervene, so the models are almost never attempting something that warrants intervention. The second is the detection reading: a block rate is a property of the screener, not of the model. A screening system tuned to a narrow list of prohibited actions will produce a low block rate against any behaviour at all, including behaviour it was never configured to recognise.

Separating the two requires a number that is not in the disclosure: how many unsafe attempts actually occurred. Blocks divided by attempts is recall, and recall is the thing a safety claim needs. Blocks divided by total decisions is a throughput statistic. One billion decisions is the denominator of activity, not the denominator of risk.

There is also a presentation effect worth naming, and it is a comparison none of the coverage made. 0.002% of more than a billion decisions is upwards of 21,000 blocked decisions in a single month, roughly 700 a day. Stated as a rate the number reads like a system at rest. Stated as a count it reads like a control surface doing continuous work. It is the same number. Which framing a reader receives is a choice made by whoever writes it down.

30,000 agents is a capacity figure

The disclosure states that roughly 30,000 agents were doing research and engineering work at any one time during August on the company's most-used internal agent platform. The instinct is to read this as a workforce, and the instinct is wrong in a specific way.

At any one time is concurrency, not completion. An agent is not a person-equivalent: it has no idle time, no working week, and it can be spawned for a task that lasts nine seconds. Thirty thousand concurrent agents is compatible with an enormous amount of finished work and with an enormous amount of retried, abandoned or discarded work. The figure establishes the scale of the platform. It establishes nothing about output, about what was merged, or about what survived human review.

Set against the 26%, the two numbers are measuring different objects. One counts tasks by who led them. The other counts processes by how many were running. Neither is a productivity figure, and the viral framing treats them as if they compound.

Why a company measures this at all

A self-improvement rate is the single most compelling claim an AI lab can make about its own trajectory. It converts an argument about future capability into a present-tense measurement, and it is legible to people who will never read a benchmark table. That is exactly why it warrants the most scrutiny rather than the least, and why the growth framing from under 1% in February to 26% in August does more work in the retellings than the absolute level does.

The strongest argument against everything above is worth stating plainly, because it is a good one. No external party could produce this number. Only Anthropic has the logs, the task records and the ability to classify them. If the standard for publishing a self-improvement metric is independent verification, no such metric will ever be published, and the industry is better off with a self-reported figure and a stated methodology than with nothing.

That argument holds, and it is why the disclosure is a genuine contribution. What it does not license is treating the output as a measurement of the world. Independent verification here would require a published taxonomy, a sample of classified tasks an outside reviewer could re-sort, an auditor with log access, and a stated denominator for the decision count. None of those exist. Until they do, the accurate way to state the finding is not that Claude leads 26% of AI research at Anthropic. It is that Anthropic's own instrument, applying Anthropic's own definition of led, currently reads 26.

Sources: Washington Post, 17 September 2026; The Hill; Digital Today. All figures are Anthropic company disclosures reported by those outlets and have not been independently audited.