Frontier Lab Transparency: Anthropic Discloses Automation, Agent Oversight, And Safety Compute Metrics

AI analysis firm Anthropic launched a set of inside measurements supposed to provide the general public, regulators, and third-party observers clearer visibility into how shortly frontier AI growth is continuing — together with how a lot of the work of constructing new fashions is now being carried out by AI itself.
The centerpiece of the discharge is the Anthropic R&D Automation Index, a prototype metric that catalogues each sort of AI analysis and growth process on the firm and charges every on an automation scale developed by Epoch AI. The scale runs from AL0, which means no AI involvement, to AL5, which means totally autonomous operation with no human within the loop. AL3 denotes duties the place AI collaborates underneath shut human route, whereas AL4 covers duties AI completes end-to-end from a high-level immediate with human supervision.
As of August 2026, Claude isn’t working totally autonomously in any measured space of AI R&D. However, AI leads roughly 26% of Anthropic’s R&D work, and greater than 90% of duties are at or above the extent the place AI collaborates with people. Anthropic frames these figures as a technique to observe how shut the {industry} is to recursive self-improvement — a state of affairs wherein a mannequin autonomously builds its personal successor — and notes that the numbers would probably shift underneath any coordinated effort to gradual frontier growth.
The firm pairs this with knowledge on the oversight of AI brokers. Roughly 30,000 brokers had been conducting analysis and engineering work on Anthropic’s most-used inside platform, and each certainly one of their actions passes via real-time on-line displays earlier than execution, in addition to offline displays that assessment exercise afterward. Over greater than a billion agent choices analyzed throughout August, about 0.002% — roughly one in 47,000 — had been blocked. Offline displays flag one to 2 transcripts per thousand, with round 100,000 flagged weekly and solely about 50 escalated to human assessment.
Compute Allocation Emerges as a Verifiable Pacing Lever
The third measurement addresses compute, some of the verifiable inputs to AI growth and subsequently a possible lever for any future industry-wide pacing regime. In a one-week snapshot of Anthropic’s complete compute utilization, roughly 6% of compute dedicated to AI R&D went to security work; inside AI-driven R&D particularly, the determine was about 12%. The firm describes these as intentionally conservative estimates, noting that compute is an imperfect proxy for security funding since a lot security analysis is design-intensive fairly than compute-intensive.
Anthropic positions the discharge as complementary to its Responsible Scaling Policy threat stories and its Advanced AI Framework coverage proposal, which outlines transparency obligations that governments may impose on labs. The firm additionally introduced plans to embed unbiased third-party evaluators from a number of organizations inside Anthropic, granting them entry to inside techniques and knowledge akin to what its personal threat evaluation groups obtain.
Cross-lab comparability stays a problem. Anthropic acknowledges there is no such thing as a shared methodology for automation measurement, and that utilizing its personal fashions to guage its techniques creates potential blind spots. It suggests third-party verification or analysis by different builders’ fashions, with safeguards for competitively delicate data, as a path towards standardized reporting that governments and the general public may depend on.
The submit Frontier Lab Transparency: Anthropic Discloses Automation, Agent Oversight, And Safety Compute Metrics appeared first on Metaverse Post.
