DeepSeek has released V4.1-Flash, the smallest model in its new architecture family, featuring native visual understanding and designed for greater capability, faster inference, and higher throughput.
Asymmetric Architecture Delivers Efficiency
V4.1-Flash is a 552-parameter mixture-of-experts (MoE) model built on a new Causal Encoder–Decoder architecture that uses just 8 billion active parameters for input processing and 16 billion for output generation. The model combines new pretraining methods with larger-scale reinforcement learning post-training to achieve benchmark results that exceed those of flagship models, including DeepSeek-V4-Pro.
| Specification | Value |
|---|---|
| Total Parameters | 552 (mixture-of-experts) |
| Active Parameters (Input) | 8 billion |
| Active Parameters (Output) | 16 billion |
| Architecture | Causal Encoder–Decoder |
Significant Cache Compression
Compared with the previous generation, V4.1-Flash reduces key-value (KV) cache requirements to one-quarter the HBM usage and one-eighth the SSD storage. Since cache-hit charges often represent a substantial portion of agent costs, this compression delivers meaningful savings for inference workloads.
Model Transitions and Pricing
DeepSeek is retiring V4-Flash and V4-Flash-Vision-Exp, with the old model identifiers temporarily routing to V4.1-Flash for compatibility. The company is also phasing out V4-Pro: beginning at 04:00 UTC on September 14, 2026, all deepseek-v4-pro API requests will route to V4.1-Flash at V4.1-Flash rates until V4.1-Pro launches.
Tests by multiple parties show V4.1-Flash outperforming V4-Pro on performance, cost, speed, and total runtime.
The more efficient architecture enables lower API pricing. Off-peak rates are 50% of peak rates, with peak/off-peak pricing continuing to balance demand. New pricing takes effect at 04:00 UTC on September 10, 2026.
Community Support and Deployment
Official partners WorkBuddy (including CodeBuddy) and OpenCode now fully support V4.1-Flash. DeepSeek is working closely with the open-source community on V4.1-Flash inference support and exploring additional deployment options. The company is also open to discussions about large-scale deployments involving 2,000 GPUs and storage clusters.
API Availability and Pricing
| Model | Price | Where to buy | As-of date |
|---|---|---|---|
| DeepSeek-V4.1-Flash | Off-peak: $0.003/$0.15/$0.60 per 1M; peak: $0.006/$0.30/$1.20 | DeepSeek API | Sep. 2026 |
| DeepSeek-V4.1-Flash | $0.112 input, $0.336 output, $0.0034 cached input/1M; promo tier | DeepInfra | Sep. 2026 |
| DeepSeek-V4.1-Flash | $0.02 input, $0.60 output, $0.02 cache read per 1M | OpenRouter | Sep. 2026 |
Most Read
