|

Alibaba Expands Qwen Into Speech And Mobile Agents As Qwen 4 And Custom Silicon Signal Broader Ambitions

Alibaba Expands Qwen Into Speech And Mobile Agents As Qwen 4 And Custom Silicon Signal Broader Ambitions
Alibaba Expands Qwen Into Speech And Mobile Agents As Qwen 4 And Custom Silicon Signal Broader Ambitions

Alibaba’s Qwen group has launched two updates to its product lineup: Qwen Audio 3.1, a rebuilt speech stack, and Qwen Intelligence, a collection of cell brokers. Together they illustrate how the corporate is changing its mannequin momentum into deployable merchandise throughout modalities.

Qwen Audio 3.1 includes 5 fashions organized into two teams: an upgraded core lineup overlaying understanding, era, and interplay, and two new fashions geared toward creation and deeper comprehension. On the synthesis facet, the flagship TTS mannequin ranks first on the unbiased Artificial Analysis Text to Speech Leaderboard. It helps 16 languages and 20 Chinese dialect areas, produces as much as three minutes of steady speech in a single cross, and accepts free type pure language directions controlling emotion, tempo, timbre, and accent, alongside 86 effective grained inline tags for phrase stage results similar to pauses, respiration, laughter, and sighs. 

According to the technical report, a 12.5 Hz low body fee speech tokenizer reduces decoding value, whereas a 5 stage coaching pipeline coordinates the language and movement fashions for content material consistency, voice similarity, and robustness; the system generates usable speech even from noisy, reverberant, or degraded reference audio. A companion mannequin, TTS Next, combines a language mannequin with a diffusion framework to generate voice, sound results, and background audio in a single cross, focusing on audiobooks, podcasts, video games, and promoting.

On the understanding facet, the upgraded ASR provides stronger multilingual and dialect recognition with native transcript sprucing that removes fillers and repetitions mechanically. ASR Next extends this to multi speaker recognition with speaker labels, timestamps, and aligned transcripts, and might interpret feelings, ambient noise, and machine sounds, enabling sound captioning, occasion localization, and audio query answering. 

The Realtime mannequin helps simultaneous talking and listening with interruption at any second, and adjusts its tempo and tone when it detects a low temper within the caller. Alibaba paired the launch with substantial value reductions: TTS prices dropped roughly 70%, Realtime round 85%, and ASR by as much as 95%.

Qwen Intelligence targets a unique layer fully: autonomous job execution on cell gadgets. Its Planner agent decomposes and orchestrates complicated duties, rating first on the corporate’s MobilePA Bench, together with its enterprise and reminiscence variants. The Mobile Use agent executes actions primarily via APIs, falling again to graphical interface management when wanted, scoring 82.1 on MobileWorld, 92.2 on the true machine MobileWorld Real benchmark, and 97.2 on AndroidDaily, with a reported 90 % success fee on full finish to finish duties. 

The Creative agent generates a usable picture from a single sentence in about three seconds, roughly twice as quick as main alternate options, Alibaba claims. Notably, the corporate additionally opened its benchmark suite, together with MobilePA Bench, MobileWorld, MobileWorld Real, and MobileWorld Safety, to the analysis group, offering exterior reference factors for evaluating cell brokers.

Following the Conference Stage: Qwen 4 and the Infrastructure Announcements

These releases arrive within the wake of Alibaba’s Apsara Conference in Hangzhou on September 22, the place the Qwen 4 household was introduced in 4 tiers: Max because the flagship positioned in opposition to prime rival fashions, Flash for low latency and high quantity workloads, Plus as a balanced multimodal tier, and a 27B open weights variant for native use. None of the 4 has public specs, pricing, or a launch date, making the announcement an announcement of path fairly than a delivery product.

The convention’s different bulletins make clear that path. Chief Executive Eddie Wu said that the Qwen group plans to coach fashions of 5 to 10 trillion parameters, a goal Alibaba’s personal statements assign to Qwen 4.5 and Qwen 5, for tackling extra complicated, longer horizon duties. To help fashions of that scale, Alibaba’s T Head unit unveiled the Zhenwu V900 accelerator, which guarantees 3 times the efficiency of the M890 predecessor, carries 216 GB of reminiscence, and scales to clusters of as much as 500,000 chips, with mass manufacturing scheduled for the primary quarter of 2027. 

Wu set a goal of surpassing 20 gigawatts of worldwide cloud capability by 2032 and described a 3 layer cloud structure designed for agentic workloads. He additionally famous that demand for AI is outpacing provide, with AI supernodes coming on-line at business scale this quarter. Alibaba’s Hong Kong shares rose 5.1 % on the day. The home chip push unfolds in opposition to tightening US export curbs, whereas analysts level to August’s open sourced Qwen3.8 Flash Next, with its sparse consideration and n gram embedding strategies, because the closest accessible preview of the subsequent era structure, although Alibaba has not confirmed the connection.

A Broader Trajectory: From Open Models to an Integrated Stack

The latest releases cap an accelerating trajectory. Since its debut in 2023, Qwen has change into essentially the most extensively used open weights mannequin household, with greater than 300 million cumulative downloads and over 100,000 by-product fashions on Hugging Face, in accordance with Alibaba. 

The present flagship, Qwen3.8 Max, launched in August with open weights, comprises 2.4 trillion parameters and handles a million tokens of context, whereas September’s Qwen3.8 Omni Flash extends native processing throughout textual content, picture, audio, and video. The licensing construction has advanced in parallel: the flagship carries a clause requiring firms producing over $50 million in annual income to share that income with Alibaba, whereas smaller distilled fashions ship beneath permissive Apache 2.0 licenses. 

Several open questions will form how this technique unfolds: whether or not independently verified benchmarks affirm the seller reported outcomes for brokers and speech fashions, whether or not chip manufacturing timelines align with the mannequin roadmap, and the way the income sharing license is adopted in observe. With Qwen 4 specs nonetheless undisclosed and the subsequent era already in coaching, Alibaba has laid out an unusually broad agenda spanning silicon, cloud infrastructure, fashions, and client brokers. The coming quarters will present how these items converge in observe, and we are going to maintain watching how Qwen develops.

The put up Alibaba Expands Qwen Into Speech And Mobile Agents As Qwen 4 And Custom Silicon Signal Broader Ambitions appeared first on Metaverse Post.

Similar Posts