|

Anthropic’s New Sonnet Model Nears Opus Performance but Uses More Tokens

Anthropic launched Claude Sonnet 5.5 on September 28 on the identical checklist value as Sonnet 5. The firm says the mannequin runs over 30% sooner and prices as much as 30% much less per process.

An impartial benchmarking agency reached a distinct conclusion on price as soon as it pushed the mannequin to its restrict.

Follow us on X to get the most recent information because it occurs

Anthropic’s Second 5.5 Model Lands Days After Opus

Sonnet 5.5 is the second mannequin within the Claude 5.5 household. Anthropic launched Opus 5.5 on September 22. On the identical day, OpenAI additionally launched GPT-6 Sol and Luna.

The new mannequin retains Sonnet 5’s costs of $2 per million enter tokens and $10 per million output tokens. Anthropic reported a 70.6% rating on Terminal-Bench 4.0, an agentic coding check. Sonnet 5 managed simply 10.3%, whereas Opus 5.5 reached 66.4%.

“Where Opus 5.5 is constructed for advanced work requiring cautious judgment, Sonnet 5.5 is strongest at well-scoped on a regular basis duties, fixing bugs, and creating polished paperwork, slides, and spreadsheets,” the team mentioned.

Anthropic mentioned Sonnet 5.5 doesn’t prolong its functionality frontier. Its alignment assessment subsequently centered on dangers akin to deceptive customers and aiding high-stakes misuse.

Sonnet 5.5 additionally ships with cyber safeguards and anti-distillation classifiers. These block makes an attempt to extract the mannequin’s reasoning for coaching rival programs. Claude Haiku 5.5 follows within the coming weeks.

Meanwhile, OpenAI shelved an upcoming mannequin, GPT-6.1 Astra, on security grounds. CNBC confirmed on Monday that the mannequin fell wanting the company’s standards.

Artificial Analysis Finds a Heavier Token Bill at Max Effort

Artificial Analysis, the benchmarking agency that ranked Grok 4.7 fourth, gave Sonnet 5.5 a 56 on its Intelligence Index. That places it second general, 2 factors behind Opus 5.5 at most effort.

The agency recorded 64% for Sonnet 5.5 on Terminal-Bench 4.0, in opposition to 60% for Opus 5.5 and GPT-6 Astra. On knowledge-work checks akin to GDPval-AA, it discovered Sonnet 5.5 roughly stage with Opus 5.5.

The agency mentioned Sonnet 5.5 used considerably extra tokens to achieve these outcomes.

“At max effort, the place it reaches efficiency nearing that of Opus 5.5, Claude Sonnet 5.5 used ~193k Output Tokens per Intelligence Index Task,” the post learn.

That quantity sits about 60% above Opus 5.5 and Sonnet 5 at max effort. It can be roughly 7 instances GPT-6 Astra’s max-effort depend. 

Early-access prospects quoted by Anthropic reported the other pattern in opposition to Sonnet 5, although on completely different workloads. Balyasny Asset Management noticed far decrease token use on its finance duties.

“On our non-public suite of two,441 finance duties overlaying Q&A, extraction, evaluation, and forecasting, Claude Sonnet 5.5 scored forward of Sonnet 5 and used about 121k tokens per reply, the place Sonnet 5 used 497k,”  Joe Poirier, Senior AI Engineer at Balyasny Asset Management, acknowledged.

Artificial Analysis examined a pre-release construct carrying a structured-output bug that Anthropic has since fixed. The agency plans to rerun the related evaluations quickly.

Subscribe to our YouTube channel to observe leaders and journalists present professional insights

The put up Anthropic’s New Sonnet Model Nears Opus Performance but Uses More Tokens appeared first on BeInCrypto.

Similar Posts