qwen-image-plus
- Input Tokens: $2.000 /M tokens
- Output Tokens: $0.000 /M tokens
- Cache Tokens: $0.000 /M tokens
- Text
- Vision
- Image
Image input: $0.00286 per image;Image generation: $0.02535 per image.
Qwen Image 3.0(qwen-image-3.0) is an image generation and editing model developed by Alibaba Cloud’s Qwen team. It supports text-to-image generation, reference-based creation, and image editing. It is well suited for social media, e-commerce, creative design, and everyday content production, offering a strong balance of image quality and speed. Compared with the Pro version, it is better suited for frequent and large-scale daily creation.
Image input: $0.00286 per image; 1K image generation: $0.03572 per image; 2K image generation: $0.07143 per image.
Qwen Image 3.0 Pro (qwen-image-3.0-pro) is Alibaba Cloud Qwen’s flagship image generation and editing model. It is designed for advertising, brand visuals, UI, presentations, product imagery, and professional design. Its strengths include complex layouts, accurate Chinese and English text rendering, realistic materials, and reference-based editing. Compared with the standard version, it delivers stronger detail, composition, and commercial-grade visual quality.
Qwen 3.8 Max(qwen3.8-max) is Alibaba Cloud’s flagship native vision-language model, built on a 2.4-trillion-parameter Mixture-of-Experts (MoE) architecture and supporting context windows of up to 1 million tokens. It is well suited for complex multimodal understanding, advanced reasoning, software development, agentic workflows, and long-context processing. At a similar price to Qwen3.7-Max, Qwen3.8-Max delivers significant improvements in reasoning, coding, and agent capabilities, with overall performance comparable to today’s leading models.
- Input: $ 0.338 /M Tokens
- Output: $ 1.014 /M Tokens
- Web Search: $0.000548/request
Qwen 3.8 Max Preview(Qwen3.8-Max-Preview) is the latest-generation foundation model in the Qwen family, packing 2.4T parameters and still evolving. Compared with the previous flagship Qwen 3.7 Max, it delivers major gains in core capabilities like Coding and Cowork (professional productivity), with world-leading performance on complex, long-horizon tasks such as full-stack development, data analysis, and Office workflows. Launch offer: Credits are consumed at just 20% of the standard rate, effectively 5× your usage. Limited time only.
$0.141 / 10K characters
qwen-audio-3.0-tts-flash is a high-performance speech synthesis large model optimized for real-time interactive scenarios. Compared with the previous version, the model supports more low-resource languages and Chinese dialects, improves the authenticity of dialect pronunciation, and enhances free-style instruction following and fine-grained label control, enabling more flexible control of expression such as emotion, tone, character, speaking rate, and volume. At the same time, the model exhibits stronger robustness under complex acoustic conditions like noise and reverberation, improving sound quality, clarity, and overall expressiveness. The Flash version focuses on optimizing the real-time synthesis experience, keeping first-packet latency under 200 ms, making it suitable for low-latency interactive scenarios such as voice assistants, real-time dialogue, and intelligent customer service.
$0.197 / 10K characters
qwen-audio-3.0-tts-plus is a high-performance speech synthesis large model designed for high-quality speech generation scenarios. Compared with the previous version, the model supports more low-resource languages and Chinese dialects, significantly improving the authenticity of dialect pronunciation, and enhancing free-style instruction following and fine-grained label control capabilities, allowing more accurate control of emotion, tone, character, speaking rate, volume, and synthesis style. Meanwhile, the model exhibits stronger robustness under complex acoustic conditions such as noise and reverberation, further improving sound quality, clarity, resolution, and overall expressiveness. The Plus version places greater emphasis on synthesis quality and detail, making it suitable for professional scenarios that require higher sound quality, naturalness, and expressiveness, such as content creation, audiobooks, film and TV dubbing, brand voice design, and high-quality voice services.
AIHubMix© 2023 - 2026 AIHubMix, LLC