DeepSeek launches DeepSeek-V4.1-Flash with native multimodal support
The 552B-parameter MoE model uses 8B active input and 16B active output parameters, reducing KV cache demand to 1/4 HBM and 1/8 SSD storage.
- DeepSeek-V4.1-Flash is live on the DeepSeek API as deepseek-flash and available on Hugging Face.
- DeepSeek-V4-Flash and DeepSeek-V4-Flash-Vision-Exp are retired and automatically route to V4.1-Flash.
- Requests to deepseek-v4-pro will route to V4.1-Flash at V4.1-Flash rates starting September 14, 2026.
- Off-peak API rates remain at 50% of peak pricing, with updated pricing taking effect September 10, 2026.
DeepSeek API users and open-source deployers get higher inference throughput and lower cache storage costs while existing V4 models are phased out.

Sources
Read this as text
Back to the AI news