|

Meta Releases Muse Glimmer With Open Weights, Targeting On-Device AI Agents Via Novel Distillation Pipeline

Meta Releases Muse Glimmer With Open Weights, Targeting On-Device AI Agents Via Novel Distillation Pipeline
Meta Releases Muse Glimmer With Open Weights, Targeting On-Device AI Agents Via Novel Distillation Pipeline

Technology firm Meta launched Muse Glimmer, a 30-billion-parameter agentic AI mannequin distributed beneath the permissive Apache 2.0 license. Developed by Meta Superintelligence Labs, the mannequin is designed to function as a completely succesful autonomous agent—together with planning, software invocation, self-verification, and failure restoration—whereas remaining compact sufficient to run regionally on client {hardware} with as little as 24 GB of video reminiscence. 

The weights can be found instantly on Hugging Face, with integrations for well-liked inference engines and platforms resembling Ollama, LM Studio, vLLM, SGLang, Together AI, Fireworks AI, and OpenRouter scheduled to observe within the coming days. 

The launch extends Meta’s custom of open-sourcing foundational AI analysis, this time focusing on the rising demand for native, always-on agent workflows that don’t rely upon cloud connectivity or exterior infrastructure.

Muse Glimmer: Architecture, Training, and Local Optimisation

The new AI mannequin was constructed utilizing a bespoke structure and a novel distillation recipe meant to switch agentic reasoning from a considerably bigger trainer mannequin, known as Muse Spark, right into a extra environment friendly kind issue. The coaching pipeline comprised three phases: pre-training by way of logit distillation on the trainer’s outputs; mid-training on extended-context, agent-heavy information enriched with reasoning traces; and post-training combining supervised fine-tuning with on-policy distillation and reinforcement studying throughout basic, coding, and agentic domains. The mannequin was evaluated beneath Meta’s Advanced AI Scaling Framework earlier than launch.

Benchmark outcomes point out aggressive efficiency relative to equally sized counterparts, together with Gemma4-31B and Qwen3.6-27B, on duties resembling DeepSearch QA, MCP-Atlas, τ-Bench, and SWE-Bench. Beyond core reasoning, Muse Glimmer helps multimodal enter by means of a devoted notion encoder, multilingual operation throughout greater than 100 languages, and compatibility with agentic orchestration patterns resembling OpenClaw.

To allow sensible native deployment, Meta utilized quantisation strategies that compress the mannequin to roughly 4-bit precision, lowering its footprint to beneath 20 GB. This leaves enough reminiscence for the KV cache, picture encoder, and a light-weight speculative decoding drafter primarily based on DFlash, which proposes token blocks in parallel to speed up era with out altering output high quality. 

Meta validated the setup on MacBook M4-Max, M5-Max, and RTX-5090 {hardware}, reporting speeds appropriate for fluid dialog and real-time agent interplay totally on-device.

The put up Meta Releases Muse Glimmer With Open Weights, Targeting On-Device AI Agents Via Novel Distillation Pipeline appeared first on Metaverse Post.

Similar Posts