Qwen3.8 Omni Flash is Alibaba Cloud Qwen's next-generation native multimodal model, supporting 1M-length sequences and able to directly accept text, images, audio, and video as input. It is based on the Qwen3.8-Flash-Next architecture. The model is designed for agent capabilities in real productivity scenarios: while offering agentic abilities such as programming, text knowledge work, and GUI operation, it also achieves significant results in audio/video-centered agentic applications—such as video editing, music video creation, film production and narration, audio/video-to-text-and-image summarization, and audio/video dialogue—that require integrated processing of text, images, audio, and video. It supports 2-channel and 4-channel spatial audio parsing, is compatible with DashScope and OpenAI protocols, and it is recommended to install the accompanying Qwen-MM-Plugins to facilitate agent frameworks' access to native multimodal capabilities.
← Models
Qwen3.8 Omni Flash
Qwen3.8 Omni Flash Compare
Compare pricing, specifications, performance, and benchmarks for up to four models.
+ Add model · 1/4
Qwen3.8 Omni Flash
Qwen · text, image, video, audio → text
Input$0.1126 /M
Output$0.38 /M
Cache read$0.0141 /M
Pick a second model to start comparing.
Popular comparisons
Related model match-ups readers also look at.