icon

qwen-image-3.0

New
Alicloud latest image generation model, qwen-image-3.0. It supports up to 4.5k-token inputs, allowing complex image-and-text instructions to be generated in one pass. Stable text rendering: 10px small text, 12 languages, and 20+ fonts are clearly legible, making infographics and interfaces directly usable. Convenient for batch production: posters, web pages, and UIs can be generated in bulk with better cost efficiency, suitable for continuous creation.
Pricing Details
PricingImage InputImage Generation
No data
No data
Input Modalities
  • Text
  • Vision
Output Modalities
  • Image
Features
    Tags
    FAQ
    What is qwen-image-3.0?
    Alicloud latest image generation model, qwen-image-3.0. It supports up to 4.5k-token inputs, allowing complex image-and-text instructions to be generated in one pass. Stable text rendering: 10px small text, 12 languages, and 20+ fonts are clearly legible, making infographics and interfaces directly usable. Convenient for batch production: posters, web pages, and UIs can be generated in bulk with better cost efficiency, suitable for continuous creation.
    More from Qwen
    • Input: $ 2 /M Tokens
    • Output: $ 0 /M Tokens

    Qwen Image 3.0 Pro (qwen-image-3.0-pro) is Alibaba Cloud Qwen’s flagship image generation and editing model. It is designed for advertising, brand visuals, UI, presentations, product imagery, and professional design. Its strengths include complex layouts, accurate Chinese and English text rendering, realistic materials, and reference-based editing. Compared with the standard version, it delivers stronger detail, composition, and commercial-grade visual quality.

    icon
    New
    Copy ID
    Input:$ 1.69 /M Tokens
    Output:$ 5.07 /M Tokens
    Context:991K
    Latency:2.761 S
    Throughput:39 TPS

    Qwen 3.8 Max(qwen3.8-max) is Alibaba Cloud’s flagship native vision-language model, built on a 2.4-trillion-parameter Mixture-of-Experts (MoE) architecture and supporting context windows of up to 1 million tokens. It is well suited for complex multimodal understanding, advanced reasoning, software development, agentic workflows, and long-context processing. At a similar price to Qwen3.7-Max, Qwen3.8-Max delivers significant improvements in reasoning, coding, and agent capabilities, with overall performance comparable to today’s leading models.

    • Input: $ 0.169 /M Tokens
    • Output: $ 0.507 /M Tokens
    • Web Search: $0.000548/request

    Qwen 3.8 Max Preview(Qwen3.8-Max-Preview) is the latest-generation foundation model in the Qwen family, packing 2.4T parameters and still evolving. Compared with the previous flagship Qwen 3.7 Max, it delivers major gains in core capabilities like Coding and Cowork (professional productivity), with world-leading performance on complex, long-horizon tasks such as full-stack development, data analysis, and Office workflows. Launch offer: Credits are consumed at just 10% of the standard rate, effectively 10× your usage. Limited time only.

    Input:$ 14.2 /M Tokens
    Output:$ 14.2 /M Tokens
    Context:-
    Latency:-
    Throughput:-

    qwen-audio-3.0-tts-flash is a high-performance speech synthesis large model optimized for real-time interactive scenarios. Compared with the previous version, the model supports more low-resource languages and Chinese dialects, improves the authenticity of dialect pronunciation, and enhances free-style instruction following and fine-grained label control, enabling more flexible control of expression such as emotion, tone, character, speaking rate, and volume. At the same time, the model exhibits stronger robustness under complex acoustic conditions like noise and reverberation, improving sound quality, clarity, and overall expressiveness. The Flash version focuses on optimizing the real-time synthesis experience, keeping first-packet latency under 200 ms, making it suitable for low-latency interactive scenarios such as voice assistants, real-time dialogue, and intelligent customer service.

    Input:$ 15 /M Tokens
    Output:$ 15 /M Tokens
    Context:-
    Latency:-
    Throughput:-

    qwen-audio-3.0-tts-plus is a high-performance speech synthesis large model designed for high-quality speech generation scenarios. Compared with the previous version, the model supports more low-resource languages and Chinese dialects, significantly improving the authenticity of dialect pronunciation, and enhancing free-style instruction following and fine-grained label control capabilities, allowing more accurate control of emotion, tone, character, speaking rate, volume, and synthesis style. Meanwhile, the model exhibits stronger robustness under complex acoustic conditions such as noise and reverberation, further improving sound quality, clarity, resolution, and overall expressiveness. The Plus version places greater emphasis on synthesis quality and detail, making it suitable for professional scenarios that require higher sound quality, naturalness, and expressiveness, such as content creation, audiobooks, film and TV dubbing, brand voice design, and high-quality voice services.

    Quality
    standard
    720x1280
    1280x720
    960x960
    832x1088
    1088x832
    1920x1080
    1080x1920
    1440x1440
    1632x1248
    1248x1632
    standard
    $0.1395
    $0.1395
    $0.1395
    $0.1395
    $0.1395
    $0.1395
    $0.1691
    $0.1691
    $0.1691
    $0.1691
    $0.1691

    HappyHorse-1.1-I2V supports image-to-video generation, further enhancing visual texture, dynamic performance, and cross-segment consistency. The model can more accurately understand input images and carry forward the creative intent, delivering significant improvements in character skin texture, ID consistency across segments, motion smoothness, text rendering stability, and audio–visual synchronization, producing higher-quality videos that are more realistic and natural, rich in detail, and more consistent.