Anthropic says Claude now “leads” 26% of its measured AI research and development work, under a new internal framework intended to quantify how much of frontier-model development is being performed by AI.
The company’s definition of “leads” is important. It does not mean fully autonomous research. Anthropic’s scale runs from no AI involvement through assistance and collaboration to “AI leads,” where a model can complete most of a task end to end from a high-level prompt while a person supervises. Anthropic says Claude is not operating fully autonomously for any measured subset of its AI R&D work, while more than 90% of the measured work is at or above the collaborative level.
The figure matters because it gives enterprises a concrete way to think about a broader shift: the human role is moving from using a tool to supervising a machine-led workstream. That is a very different operating model from a conventional productivity assistant.

Anthropic’s methodology also deserves scrutiny. The company built a map of roughly 15,000 granular R&D tasks using internal work records, then used Claude agents to organise and research those tasks and another Claude-based judge to assign automation levels. Anthropic acknowledges the obvious limitation: using its own models to evaluate its own automation can introduce correlated errors. It says its model’s assessments were checked against ratings from human staff, while also arguing that third-party verification would improve comparability.
For enterprise governance, the lesson is not that every company will soon have AI doing a quarter of its research. Anthropic is a frontier AI lab and therefore an unusual environment. The more transferable lesson is that organisations will need controls for supervision at scale.
A policy that assumes a person directly performs every significant action becomes less useful when a single employee may oversee multiple AI-driven tasks, each of which can create code, analyse data or interact with internal systems. Auditability, exception handling and escalation become core design requirements rather than post-deployment compliance work.
Anthropic’s numbers are self-reported and method-dependent, not an independent industry benchmark. But the company has at least put a measurable concept on the table. For governance teams, that makes the question more practical: not simply how capable is the model, but how much of a real workflow is it trusted to lead, and what evidence shows that human supervision still works?
