|

New DeepSeek V4.1-Flash Challenges Larger AI Systems On Coding And Agent Tasks

New DeepSeek V4.1-Flash Challenges Larger AI Systems On Coding And Agent Tasks
New DeepSeek V4.1-Flash Challenges Larger AI Systems On Coding And Agent Tasks

Chinese AI startup DeepSeek launched DeepSeek-V4.1-Flash, describing it because the smallest mannequin in a brand new structure household. The launch comes as the corporate is reportedly making ready for an preliminary public providing on Shanghai’s technology-focused STAR Market.

The multimodal mannequin has a 552-billion-parameter mixture-of-experts spine and helps contexts of as much as a million tokens. Its principal advance is effectivity relatively than sheer dimension. DeepSeek’s causal encoder-decoder structure divides 40 transformer layers into separate 20-layer encoding and decoding phases. As a outcome, solely eight billion parameters are activated for every enter token and 16 billion for every generated token.

The structure is designed for input-heavy agentic duties, the place processing massive paperwork, codebases or dialog histories can dominate computing prices. DeepSeek mentioned its SWA Bounded Replay methodology removes the necessity to persist chosen consideration states on solid-state storage, lowering the persistent KV-cache footprint to roughly one-eighth that of DeepSeek-V4-Flash.

A associated Compressed Sparse Attention 2 system assigns static consideration modes throughout layers and reuses chosen consideration indices. In mixture with FP4-formatted key-value caching, it reduces the worldwide KV-cache requirement to 890 bytes per token, about one-quarter of the earlier Flash mannequin. The mannequin additionally incorporates conditional reminiscence and speculative decoding whereas utilizing one shared and 384 routed consultants per MoE layer.

Competitive Results Across Reasoning, Coding and Agents

DeepSeek skilled V4.1-Flash from scratch on 45 trillion multimodal tokens, together with photographs processed natively alongside textual content from the start of pretraining. Sparse-attention coaching started at 64,000 tokens earlier than context size was prolonged to 1 million on the 34-trillion-token stage. Post-training mixed supervised fine-tuning, reinforcement studying and on-policy distillation, with reasoning effort configurable on a scale from one to 100.

Benchmarks point out robust efficiency relative to a lot bigger fashions. The base model scored 74.1 on MMLU-Pro, 79.4 on HumanEval and 93.0 on GSM8K, in contrast with 73.5, 76.8 and 92.6 for DeepSeek-V4-Pro-Base. At most reasoning effort, V4.1-Flash achieved 90.6 on Terminal-Bench 2.1 and 74.2 on DeepSWE v1.1, inserting it barely forward of main Opus and GPT fashions on these agent benchmarks. It additionally posted a 3,471 Codeforces ranking and matched one of the best listed outcome on MathArena Apex.

The outcomes usually are not uniformly dominant. V4.1-Flash remained behind the most important comparability fashions on GPQA Diamond, Humanity’s Last Exam and SimpleQA. Its scores on the newer Terminal-Bench 3.0 and 4.0 additionally trailed Opus-5.0, regardless of representing substantial positive factors over earlier DeepSeek fashions.

Multimodal efficiency included 56.5 on MMMU-Pro, 77.9 on CVBench and 95.6 on DocVQA. With agent instruments enabled, the mannequin scored 78.9 on Chartography and 89.6 on BabyVision.

The publish New DeepSeek V4.1-Flash Challenges Larger AI Systems On Coding And Agent Tasks appeared first on Metaverse Post.

Similar Posts