qwen-audio-3.0-tts-plus

by Qwen

qwen-audio-3.0-tts-plus is a high-performance speech synthesis large model designed for high-quality speech generation scenarios. Compared with the previous version, the model supports more low-resource languages and Chinese dialects, significantly improving the authenticity of dialect pronunciation, and enhancing free-style instruction following and fine-grained label control capabilities, allowing more accurate control of emotion, tone, character, speaking rate, volume, and synthesis style. Meanwhile, the model exhibits stronger robustness under complex acoustic conditions such as noise and reverberation, further improving sound quality, clarity, resolution, and overall expressiveness. The Plus version places greater emphasis on synthesis quality and detail, making it suitable for professional scenarios that require higher sound quality, naturalness, and expressiveness, such as content creation, audiobooks, film and TV dubbing, brand voice design, and high-quality voice services.

API Pricing

Input$15 / 1M tokens
Output$15 / 1M tokens

Specifications

Modalitiestext

More from Qwen

Use qwen-audio-3.0-tts-plus via the AIHubMix unified API — one interface for every major LLM.