Qwen3.8-Omni-Flash: Omni Senses. Agentic Delivery.
The native omnimodal model supports text, image, audio, and video inputs with a 1M-token context window alongside real-time streaming interaction.
- Reduces API pricing compared to Qwen3.5-Omni-Plus by over 98% per hour for audio input and over 93% for audio-visual input
- Introduces Qwen3.8-Omni-Flash-Realtime with WebSocket and WebRTC support, spatial audio perception, and real-time tool calling
- Uses an agentic coarse-to-fine evidence gathering mode that cuts token consumption by 45.7% on OmniVideoBench while increasing accuracy
- Open-sources Qwen-Live Harness for real-time agent workflows and updates Qwen-MM-Plugins with Video2Note and Omni Skill Creator tools
Developers get a cheaper, long-context omnimodal model capable of end-to-end audio-visual processing and real-time conversational agent execution.
Sources
Read this as text
Back to the AI news