Qwen3.8-Flash-Next: A New Architecture, Towards Ultimate Cost-Efficiency

The 125B multimodal MoE model activates 6B parameters per token and previews the architectural changes designed for the upcoming Qwen4 family.

Developers get an open-weight multimodal model with 1M-token context capability and lower serving costs ahead of the full Qwen4 release.

Sources

Read this as text

Back to the AI news