by Qwen
qwen-audio-3.0-tts-flash is a high-performance speech synthesis large model optimized for real-time interactive scenarios. Compared with the previous version, the model supports more low-resource languages and Chinese dialects, improves the authenticity of dialect pronunciation, and enhances free-style instruction following and fine-grained label control, enabling more flexible control of expression such as emotion, tone, character, speaking rate, and volume. At the same time, the model exhibits stronger robustness under complex acoustic conditions like noise and reverberation, improving sound quality, clarity, and overall expressiveness. The Flash version focuses on optimizing the real-time synthesis experience, keeping first-packet latency under 200 ms, making it suitable for low-latency interactive scenarios such as voice assistants, real-time dialogue, and intelligent customer service.
Use qwen-audio-3.0-tts-flash via the AIHubMix unified API — one interface for every major LLM.