# Qwen3.8-Flash-Next: A New Architecture, Towards Ultimate Cost-Efficiency

The 125B multimodal MoE model activates 6B parameters per token and previews the architectural changes designed for the upcoming Qwen4 family.

- Introduces 51B offloadable N-gram embedding parameters, a four-branch Gated Residual stream, and a Gated DeltaNet plus Qwen Sparse Attention hybrid design.
- Supports 262,144 native context tokens, extensible to 1,000,000 tokens using YaRN.
- Available as open weights on Hugging Face and ModelScope, and hosted on QwenCloud at $0.16 per million input tokens and $0.47 per million output tokens.

## Why it matters

Developers get an open-weight multimodal model with 1M-token context capability and lower serving costs ahead of the full Qwen4 release.

## Sources

- [Qwen: Qwen3.8-Flash-Next: A New Architecture, Towards Ultimate Cost-Efficiency](https://qwen.ai/blog?id=qwen3.8-flash-next)

---

Summarized by dstilled on 2026-08-26. https://dstilled.ai/story/cac3a9f3-8d9a-4c92-ac59-3c66f62a2659
