# DeepSeek launches DeepSeek-V4.1-Flash with native multimodal support

The 552B-parameter MoE model uses 8B active input and 16B active output parameters, reducing KV cache demand to 1/4 HBM and 1/8 SSD storage.

- DeepSeek-V4.1-Flash is live on the DeepSeek API as deepseek-flash and available on Hugging Face.
- DeepSeek-V4-Flash and DeepSeek-V4-Flash-Vision-Exp are retired and automatically route to V4.1-Flash.
- Requests to deepseek-v4-pro will route to V4.1-Flash at V4.1-Flash rates starting September 14, 2026.
- Off-peak API rates remain at 50% of peak pricing, with updated pricing taking effect September 10, 2026.

## Why it matters

DeepSeek API users and open-source deployers get higher inference throughput and lower cache storage costs while existing V4 models are phased out.

## Sources

- [DeepSeek: DeepSeek launches DeepSeek-V4.1-Flash with native multimodal support](https://x.com/deepseek_ai/status/2097930608790167907)

---

Summarized by dstilled on 2026-09-10. https://dstilled.ai/story/8885930a-72e1-4737-a25d-d6e6f3850f15
