We present DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models — DeepSeek-V4.1-Flash with 552B parameters (8B / 16B activated) and DeepSeek-V4-Pro with 1.6T parameters (49B activated) — both supporting a context length of one million tokens.
Introduction#
DeepSeek-V4.1-Flash#
We introduce DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to one million tokens. The model natively processes images and text, and generates text autoregressively.
Architecture. DeepSeek-V4.1-Flash adopts a Causal Encoder-Decoder (CED) architecture: a 40-layer Transformer organized as a 20-layer causal encoder followed by a 20-layer decoder. With CED, the decoder’s global KV cache is projected from the final encoder hidden states rather than derived from each decoder layer’s own hidden states. This allows the model to activate only 8B parameters per token during prefill and 16B during decode, substantially improving cost efficiency for input-heavy agentic workloads. SWA Bounded Replay reconstructs missing SWA KV states by replaying only the most recent n_win tokens, avoiding the need to persist SWA KV to SSD and reducing the persistent KV cache footprint to roughly 1/8 of that of DeepSeek-V4-Flash.
Compressed Sparse Attention 2 (CSA2). DeepSeek-V4.1-Flash uses CSA2, which assigns each attention layer one of three static modes — Full, Reindex, or Reuse — to share main KV and indexer K across layers and reuse Top-K sparse-attention indices. In the decoder, a Hierarchical Sparse Indexer further restricts later indexing layers to a candidate pool constructed by the first Full Mode layer, bounding deeper indexer cost independently of context length. Combined with FP4 main KV caching (E2M1 format, one E4M3 scale per 16 channels), these designs reduce the global KV cache footprint to 890 bytes per token — roughly 1/4 of DeepSeek-V4-Flash.
Additional architectural components include Single-Pass mHC (revised residual-stream mixing with an efficient Mega-mHC kernel), Engram conditional memory (196B parameters, sparsely accessed via token-based lookup), and DSpark speculative decoding (semi-autoregressive draft generation with confidence-scheduled verification). The model uses 1 shared expert and 384 routed experts per MoE layer, activating 6 routed experts per token.
Multimodal architecture. A vision encoder (DeepSeek-ViT, trained from scratch with 2D-RoPE and 3×3 pixel-unshuffle downsampling) and a two-layer MLP projector convert images into visual embeddings, processed jointly with text embeddings from the start of language-model pre-training.
Pre-training. DeepSeek-V4.1-Flash is trained from scratch on a multimodal corpus comprising 45T tokens, with sparse attention trained at a sequence length of 64K and context extended to 1M tokens at 34T tokens.
Post-training. The post-training recipe follows the standard SFT → RL → on-policy distillation (OPD) paradigm without algorithmic modifications. All substantive changes lie instead in the data pipeline: large-scale automated synthesis of agent tasks and environments with progressive scaling of data, tasks, and rollouts. The model supports a continuously controllable reasoning effort setting (integer 1–100) that trades inference cost for accuracy.
DeepSeek-V4#
DeepSeek-V4 series incorporate several key upgrades in architecture and optimization:
- Hybrid Attention Architecture: We design a hybrid attention mechanism combining Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA) to dramatically improve long-context efficiency. In the 1M-token context setting, DeepSeek-V4-Pro requires only 27% of single-token inference FLOPs and 10% of KV cache compared with DeepSeek-V3.2.
- Manifold-Constrained Hyper-Connections (mHC): We incorporate mHC to strengthen conventional residual connections, enhancing stability of signal propagation across layers while preserving model expressivity.
- Muon Optimizer: We employ the Muon optimizer for faster convergence and greater training stability.
We pre-train both models on more than 32T diverse and high-quality tokens, followed by a comprehensive post-training pipeline. The post-training features a two-stage paradigm: independent cultivation of domain-specific experts (through SFT and RL with GRPO), followed by unified model consolidation via on-policy distillation, integrating distinct proficiencies across diverse domains into a single model.
DeepSeek-V4-Pro-Max, the maximum reasoning effort mode of DeepSeek-V4-Pro, significantly advances the knowledge capabilities of open-source models, firmly establishing itself as the best open-source model available today. It achieves top-tier performance in coding benchmarks and significantly bridges the gap with leading closed-source models on reasoning and agentic tasks. Meanwhile, DeepSeek-V4-Flash-Max achieves comparable reasoning performance to the Pro version when given a larger thinking budget, though its smaller parameter scale naturally places it slightly behind on pure knowledge tasks and the most complex agentic workflows.
Model Downloads#
| Model | #Total Params | #Activated Params | Context Length | Precision | Download |
|---|---|---|---|---|---|
| DeepSeek-V4.1-Flash | 552B | 8B / 16B | 1M | FP4 + FP8 Mixed* | HuggingFace | ModelScope |
| DeepSeek-V4-Flash-Base | 284B | 13B | 1M | FP8 Mixed | HuggingFace | ModelScope |
| DeepSeek-V4-Flash | 284B | 13B | 1M | FP4 + FP8 Mixed* | HuggingFace | ModelScope |
| DeepSeek-V4-Pro-Base | 1.6T | 49B | 1M | FP8 Mixed | HuggingFace | ModelScope |
| DeepSeek-V4-Pro | 1.6T | 49B | 1M | FP4 + FP8 Mixed* | HuggingFace | ModelScope |
*FP4 + FP8 Mixed: MoE expert parameters use FP4 precision; most other parameters use FP8.
Evaluation Results#
Base Model#
All base models are evaluated in our internal framework under the same evaluation settings. Scores within 0.3 of each other are considered equivalent.
| Benchmark (Metric) | # Shots | DeepSeek-V4-Pro-Base | DeepSeek-V4.1-Flash-Base | DeepSeek-V4-Flash-Base |
|---|---|---|---|---|
| Architecture | — | MoE | MoE | MoE |
| # Backbone Params | — | 1.6T | 552B | 284B |
| # Activated Params | — | 49B | 8B / 16B | 13B |
| World Knowledge | ||||
| AGIEval (EM) | 3–5-shot | 84.4 | 83.4 | 83.9 |
| MMLU-Pro (EM) | 5-shot | 73.5 | 74.1 | 68.3 |
| C-Eval (EM) | 5-shot | 93.1 | 92.1 | 92.1 |
| MultiLoKo (LLM-Judge) | 5-shot | 50.9 | 45.5 | 42.6 |
| SimpleQA-Verified (EM) | 25-shot | 55.2 | 42.3 | 30.1 |
| SuperGPQA (EM) | 5-shot | 53.9 | 53.1 | 46.5 |
| Language & Reasoning | ||||
| BBH (EM) | 3-shot | 87.5 | 86.1 | 86.9 |
| BBEH (EM) | 1-shot | 29.8 | 27.2 | 25.4 |
| DROP (F1) | 1-shot | 88.7 | 87.9 | 88.6 |
| HellaSwag (EM) | 0-shot | 88.0 | 87.2 | 85.7 |
| Code & Math | ||||
| BigCodeBench (Pass@1) | 3-shot | 59.2 | 60.6 | 56.8 |
| HumanEval (Pass@1) | 0-shot | 76.8 | 79.4 | 69.5 |
| GSM8K (EM) | 8-shot | 92.6 | 93.0 | 90.8 |
| MATH (EM) | 4-shot | 64.5 | 61.1 | 57.4 |
| MGSM (EM) | 8-shot | 84.4 | 80.2 | 85.7 |
| Long Context | ||||
| LongBench-V2 (EM) | 1-shot | 51.5 | 45.2 | 44.7 |
| Multimodal | ||||
| MMMU-Pro (EM) | 4-shot | — | 56.5 | — |
| CVBench (EM) | 4-shot | — | 77.9 | — |
| DocVQA (LLM-Judge) | 4-shot | — | 95.6 | — |
| RefCOCO-avg (Acc@0.5) | 0-shot | — | 86.0 | — |
Instruct Model#
DeepSeek-V4.1-Flash supports a continuously controllable reasoning effort from 1 to 100. All instruct results below use the maximum effort setting (reasoning_effort=100). Evaluations use temperature=1.0, top_p=0.95.
For code agent benchmarks (Terminal-Bench 2.1/3.0/4.0, DeepSWE v1.1, NL2Repo-Bench, ProgramBench), the model is evaluated with the Minimal mode of DeepSeek Harness and a 1M-token context window. To align with official setup requirements, the mini-SWE harness is used for DeepSWE v1.1, and the Claude Code harness for SEC-Bench Pro. Visual agent benchmarks (Chartography, BabyVision, ZeroBench) use the Claude Code harness with a 512k-token context window. Agent’s Last Exam and AutomationBench use their official scaffolds. All agentic evaluations use temperature=1.0, top_p=0.95.
Comparison with frontier models#
| Benchmark (Metric) | DS-V4.1-Flash | Opus-5.0 | GPT-5.6 Sol | Kimi-K3 | GLM-5.3 | DS-V4-Pro | DS-V4-Flash |
|---|---|---|---|---|---|---|---|
| Reasoning | |||||||
| GPQA Diamond (Pass@1) | 90.9 | 93.4 | 94.1 | 92.9 | 88.1 | 92.4 | 89.9 |
| HLE (Pass@1) | 36.8 (39.1†) | 56.3 | 44.5 | 43.5 | 42.0† | 42.7† | 37.8† |
| Codeforces (Rating) | 3471 | — | — | — | — | 3348 | 3289 |
| MathArena Apex (Pass@1) | 65.6 | — | — | 65.6 | — | 65.3 | 58.6 |
| Agentic | |||||||
| Terminal-Bench 2.1 (Pass@1) | 90.6 | 89.1 | 88.8 | 88.3 | 88.2 | 87.9 | 82.7 |
| Terminal-Bench 3.0 (Pass@1) | 30.0 | 43.3 | 34.4 | 17.7 | 28.3 | 11.8 | 7.6 |
| Terminal-Bench 4.0 (Pass@1) | 31.2 | 51.8 | 39.9 | 12.6 | 37.9 | 12.4 | 7.0 |
| DeepSWE v1.1 (Resolved) | 74.2 | 74.0 | 73.0 | 67.5 | 66.9 | 62.7 | 54.4 |
| ProgramBench (Almost@1) | 20.3 | 37.0 | 23.0 | 17.5 | 19.0 | 15.5 | — |
| NL2Repo-Bench (Score) | 64.0 | 75.3 | 56.8 | 58.0 | 58.0 | 61.5 | 54.2 |
| CyberGym (Pass@1) | 88.1 | — | 84.5 | 80.0 | 84.5 | 83.3 | 76.7 |
| SEC-Bench Pro (Pass@1) | 62.8 | — | 74.3 | — | — | 56.4 | 30.9 |
| ExploitGym (Pass@1) | 15.3 | 22.1 | 33.7 | — | 15.0 | 5.4 | 1.8 |
| HLE w/ tools (Pass@1) | 63.9 | 63.6 | — | 59.8 | 62.5 | 60.0 | 51.5 |
| AutomationBench (Pass@1) | 54.8 | 50.3 | 45.8 | 46.7 | 48.8 | 43.2 | 37.7 |
| Agent’s Last Exam (Pass@1) | 31.8 | 28.6 | 26.7 | 27.6 | 28.5 | 25.7 | 25.2 |
| Chartography w/ tools (Pass@1) | 78.9 | 84.0 | 79.9 | 68.1 | — | — | — |
| BabyVision w/ tools (Pass@1) | 89.6 | 94.1 | 88.9 | 85.7 | — | — | — |
| ZeroBench-main w/ tools (Pass@5) | 49.0 | 52.0 | 53.0 | 41.0 | — | — | — |
† Text-only subset of HLE.
Performance across agent scaffolds#
All scaffolds use N=8 samples per task on DeepSWE v1.1 and N=3 on Terminal-Bench 2.1, with Linux containers, temperature=1.0, top_p=0.95, a 1M-token context limit, and max_steps=500 per agent. Terminal-Bench 2.1 is evaluated without network access.
| Benchmark (Metric) | Claude Code | Codex | OpenCode | Pi | mini-SWE | DSH Minimal | DSH Standard | DSH PTC |
|---|---|---|---|---|---|---|---|---|
| DeepSWE v1.1 (Resolved) | 69.8 | 65.6 | 65.5 | 66.2 | 74.2 | 72.6 | 70.5 | 67.6 |
| Terminal-Bench 2.1 (Pass@1) | 88.0 | 84.1 | 85.0 | 86.1 | 90.3 | 90.6 | 85.8 | 85.8 |
Chat Template#
This release does not include a Jinja-format chat template. Instead, we provide a dedicated encoding folder with Python scripts and test cases demonstrating how to encode messages in OpenAI-compatible format into input strings for the model, and how to parse the model’s text output.
A brief example:
from encoding_dsv4 import encode_messages, parse_message_from_completion_text
messages = [
{"role": "user", "content": "hello"},
{"role": "assistant", "content": "Hello! I am DeepSeek.", "reasoning_content": "thinking..."},
{"role": "user", "content": "1+1=?"}
]
# messages -> string
prompt = encode_messages(messages, thinking_mode="thinking")
# string -> tokens
import transformers
tokenizer = transformers.AutoTokenizer.from_pretrained("deepseek-ai/DeepSeek-V4-Pro")
tokens = tokenizer.encode(prompt)pythonHow to Run Locally#
For local deployment, we recommend setting the sampling parameters to temperature = 1.0, top_p = 1.0. For the Think Max reasoning mode, we recommend setting the context window to at least 384K tokens.
License#
This repository and the model weights are licensed under the MIT License.
Citation#
@misc{deepseekai2026deepseekv4,
title={DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence},
author={DeepSeek-AI},
year={2026},
}bibtex