Qwen3.8-Flash-Next: A New Architecture, Towards Ultimate Cost-Efficiency
The 125B multimodal MoE model activates 6B parameters per token and previews the architectural changes designed for the upcoming Qwen4 family.
- Introduces 51B offloadable N-gram embedding parameters, a four-branch Gated Residual stream, and a Gated DeltaNet plus Qwen Sparse Attention hybrid design.
- Supports 262,144 native context tokens, extensible to 1,000,000 tokens using YaRN.
- Available as open weights on Hugging Face and ModelScope, and hosted on QwenCloud at $0.16 per million input tokens and $0.47 per million output tokens.
Developers get an open-weight multimodal model with 1M-token context capability and lower serving costs ahead of the full Qwen4 release.
Sources
Read this as text
Back to the AI news