|

Zuckerberg’s Muse Code Loses to Anthropic on Meta’s Own Benchmark Charts

Mark Zuckerberg launched Muse Code in beta on Wednesday, Meta’s first synthetic intelligence (AI) coding agent. Anthropic’s Claude Opus 5 beats it in all 4 comparisons Meta printed at launch.

Meta launched these charts anyway. The firm is promoting a less expensive instrument somewhat than a greater one. Independent check knowledge suggests the hole is wider than Meta confirmed.

Follow us on X to get the most recent information because it occurs

Meta’s Own Charts Hand Anthropic Every Round

Muse Spark 1.2 is the mannequin inside Muse Code. It scored 82.9% on Terminal-Bench 2.1.

Claude Opus 5 scored 86.7% on the identical check. Terminal-Bench comes from the Laude Institute and Stanford researchers. It units 89 actual jobs spanning system restore, knowledge work, and safety.

Second place is respectable. Muse Code beat OpenAI’s Codex at 81.8% and Grok Build at 81.6%.

Benchmark comparison chart for Muse Spark 1.2
Benchmark comparability chart for Muse Spark 1.2, Source: Zuckerberg

The subsequent chart was harsher. DeepSWE 1.1 units 113 coding duties with web entry switched off throughout grading. Muse Spark 1.2 dropped to third at 59.3%.

Meta then printed a check it constructed itself, drawn from 440 actual pull requests by its personal engineers. Muse Spark 1.2 scored 70.6% there, roughly 9 factors behind Opus 5.

Muse Spark 1.2 ran a kernel optimization task for 24 hours on NVIDIA Hopper — over 1,000 tool calls — and kept finding real speedups long after the early exploration phase.
Muse Spark 1.2 ran a kernel optimization process for twenty-four hours on NVIDIA Hopper — over 1,000 instrument calls — and stored discovering actual speedups lengthy after the early exploration part.

That rating sits solely 2.3 factors above Muse Spark 1.1, the mannequin Meta shipped in July.

The Model Meta Left Off Its Coding Charts

Meta measured itself in opposition to GPT-5.6 Terra. OpenAI sells a stronger mannequin known as Sol, and Meta left it out of all three coding charts.

Sol tops the impartial Terminal-Bench 2.1 leaderboard at 89.5%. Opus 5 follows at 89.1%.

Both figures beat the 86.7% Meta reported for Opus 5. Meta picked a weaker setting of its strongest rival and nonetheless completed behind it.

Against Sol, the true chief, Muse Spark 1.2 trails by 6.6 factors somewhat than 3.8.

Meta did embrace Sol in a single place. On a graphics processing unit (GPU) kernel process operating previous 1,000 instrument calls, Sol improved on the baseline by 71.2%. Muse Spark 1.2 managed 68.7% and positioned fourth of six.

One caveat cuts the opposite approach. Muse Spark 1.2 doesn’t seem on that public leaderboard but, the place solely 26 of 183 tracked fashions have been examined. Its 82.9% stays a Meta determine.

“Muse Spark 1.2 is our subsequent step as we push towards frontier, with bigger, extra succesful fashions on the best way,” Zuckerberg said in a put up.

Zuckerberg May Soon Host the Model Beating His Own

Meta is reportedly in talks to lease compute to Anthropic. The deal may attain $10 billion over two years. Meta knowledge facilities would then assist run the Claude fashions Muse Code was constructed to unseat.

The management behind Muse Code was costly. Zuckerberg paid $14.3 billion in June 2025 for Scale AI and its founder Alexandr Wang, who now heads Meta Superintelligence Labs.

Price is the lever Wang has left. Rates match the July launch of Meta’s first paid API at $1.25 per million enter tokens and $4.25 per million output tokens.

A contributor tier prices greater than 10 occasions much less. Developers qualify by letting Meta prepare on their work. Wang declined to give adoption numbers for the Muse Spark line.

Meta’s accounts clarify the low cost. Revenue climbed 28% to $60.8 billion last quarter, but working revenue fell 8% to $18.8 billion.

Operating margin slid to 31% from 43% a 12 months earlier. Meta spent $31.08 billion on capital initiatives within the quarter alone, and guides to as a lot as $145 billion for the 12 months.

Muse Code does supply engineering Claude Code lacks. Background brokers maintain context throughout a session. Sub-agents work in remoted copies of a repository.

Meta has constructed a strong second-best coder and priced it like a finances choice. The beta will present whether or not builders commerce a number of factors of accuracy for a invoice roughly a tenth the scale.

The put up Zuckerberg’s Muse Code Loses to Anthropic on Meta’s Own Benchmark Charts appeared first on BeInCrypto.

Similar Posts