DeepSeek launches DeepSeek-V4.1-Flash with native multimodal support

The 552B-parameter MoE model uses 8B active input and 16B active output parameters, reducing KV cache demand to 1/4 HBM and 1/8 SSD storage.

DeepSeek API users and open-source deployers get higher inference throughput and lower cache storage costs while existing V4 models are phased out.

DeepSeek launches DeepSeek-V4.1-Flash with native multimodal support

Sources

Read this as text

Back to the AI news