Alibaba Prices Qwen3.8-Flash API At $0.16 Per Million Tokens, Cutting Inference Costs For 125B-Parameter Model

The Qwen Team has launched Qwen3.8-Flash-Next, an open-weight multimodal mixture-of-experts mannequin that serves as an early architectural preview for the upcoming Qwen4 collection. The mannequin balances substantial capability with distinctive value effectivity, that includes 125 billion whole parameters and an extra 51 billion N-gram embedding parameters whereas activating solely 6 billion parameters per token.
The launch follows the precedent established by Qwen3-Next, which launched hybrid structure designs later adopted throughout the Qwen3.5 via Qwen3.8 households. Qwen3.8-Flash-Next will likely be accessible via the QwenCloud API at a fee of $0.16 per million enter tokens and $0.47 per million output tokens, positioning it as a extremely aggressive possibility for high-volume functions, coding assistants, and enterprise agentic workflows.
Benchmark outcomes point out sturdy efficiency throughout software program engineering and autonomous agent duties. The mannequin achieved 58.7% on DeepSWE 1.1, 62.5% on SWE-bench Pro, and 81.0% on the multilingual variant. In long-horizon workplace automation measured by CoWorkBench, it scored 73.9%, surpassing each Qwen3.7-Plus and Claude-Opus-4.6.
General reasoning capabilities stay strong, with scores of 91.7% on GPQA Diamond and 91.9% on StayCodeBench v6. Multimodal competence is equally stable, with 84.5% on AndroidWorld and 76.6% on LVBench for lengthy video understanding. Native context size reaches 262,144 tokens, extensible to 1 million tokens through YaRN.
Architectural Innovations and Developer Integration
The structure introduces 4 systematic upgrades. For consideration, the mannequin combines Gated DeltaNet with Qwen Sparse Attention, compressing historic context effectively whereas retrieving related data via micro-block indexing quite than token-level processing. This design achieves as much as 7.6× prefill speedup at a million tokens.
The Gated Residual mechanism expands the residual stream into 4 parallel branches with dynamic gating, enhancing cross-layer data circulation and coaching stability whereas supporting FP8 storage for decreased reminiscence site visitors. N-gram Embedding provides capability via local-context lookups that require minimal computation and could be asynchronously prefetched from host reminiscence. Training employs the Muon optimizer with refined orthogonalization and parameter-splitting methods, enabling steady convergence at bigger batch sizes with out conventional warmup procedures.
The mannequin weights can be found on HuggingFace and ModelScope, with API entry via QwenCloud supporting OpenAI-compatible Chat Completions and Anthropic-compatible protocols. Developers can combine the mannequin into current workflows via Claude Code, OpenAI Codex, Qoder CLI, Qwen Code, and OpenClaw, with reasoning effort configurable throughout low, medium, and xhigh ranges. An official manufacturing launch with built-in instruments and default one-million-token context is predicted to comply with shortly.
The put up Alibaba Prices Qwen3.8-Flash API At $0.16 Per Million Tokens, Cutting Inference Costs For 125B-Parameter Model appeared first on Metaverse Post.

Meet Qwen3.8-Flash, a multimodal MoE and an early preview of the Qwen4 structure, now open-weight!