icon

qwen-image-3.0-pro

New
Alicloud latest image generation model., qwen-image-3.0-pro. It supports up to 4.5k token inputs and dense in-image information layout, enabling complex layouts such as newspapers, storyboards, menus, and exam papers to be generated in one pass. Realistic details: supports precise rendering of 10px small text, vividly reproduces micro-expressions, pores, hair strands, and other details, approaching the texture of real photography. Rich knowledge: supports 12 languages and native rendering of 20+ fonts, simulates mainstream web, game, and live-streaming interfaces, and incorporates external knowledge.
Pricing Details
PricingImage InputImage Generation
No data
No data
Input Modalities
  • Text
  • Vision
Output Modalities
  • Image
Features
    Tags
    FAQ
    What is qwen-image-3.0-pro?
    Alicloud latest image generation model., qwen-image-3.0-pro. It supports up to 4.5k token inputs and dense in-image information layout, enabling complex layouts such as newspapers, storyboards, menus, and exam papers to be generated in one pass. Realistic details: supports precise rendering of 10px small text, vividly reproduces micro-expressions, pores, hair strands, and other details, approaching the texture of real photography. Rich knowledge: supports 12 languages and native rendering of 20+ fonts, simulates mainstream web, game, and live-streaming interfaces, and incorporates external knowledge.
    More from Qwen
    icon
    New
    Copy ID
    • Input: $ 2 /M Tokens
    • Output: $ 0 /M Tokens

    Qwen Image 3.0(qwen-image-3.0) is an image generation and editing model developed by Alibaba Cloud’s Qwen team. It supports text-to-image generation, reference-based creation, and image editing. It is well suited for social media, e-commerce, creative design, and everyday content production, offering a strong balance of image quality and speed. Compared with the Pro version, it is better suited for frequent and large-scale daily creation.

    icon
    New
    Copy ID
    Input:$ 1.69 /M Tokens
    Output:$ 5.07 /M Tokens
    Context:991K
    Latency:2.761 S
    Throughput:39 TPS

    Qwen 3.8 Max(qwen3.8-max) is Alibaba Cloud’s flagship native vision-language model, built on a 2.4-trillion-parameter Mixture-of-Experts (MoE) architecture and supporting context windows of up to 1 million tokens. It is well suited for complex multimodal understanding, advanced reasoning, software development, agentic workflows, and long-context processing. At a similar price to Qwen3.7-Max, Qwen3.8-Max delivers significant improvements in reasoning, coding, and agent capabilities, with overall performance comparable to today’s leading models.

    • Input: $ 0.169 /M Tokens
    • Output: $ 0.507 /M Tokens
    • Web Search: $0.000548/request

    Qwen 3.8 Max Preview(Qwen3.8-Max-Preview) is the latest-generation foundation model in the Qwen family, packing 2.4T parameters and still evolving. Compared with the previous flagship Qwen 3.7 Max, it delivers major gains in core capabilities like Coding and Cowork (professional productivity), with world-leading performance on complex, long-horizon tasks such as full-stack development, data analysis, and Office workflows. Launch offer: Credits are consumed at just 10% of the standard rate, effectively 10× your usage. Limited time only.

    Input:$ 14.2 /M Tokens
    Output:$ 14.2 /M Tokens
    Context:-
    Latency:-
    Throughput:-

    qwen-audio-3.0-tts-flash is a high-performance speech synthesis large model optimized for real-time interactive scenarios. Compared with the previous version, the model supports more low-resource languages and Chinese dialects, improves the authenticity of dialect pronunciation, and enhances free-style instruction following and fine-grained label control, enabling more flexible control of expression such as emotion, tone, character, speaking rate, and volume. At the same time, the model exhibits stronger robustness under complex acoustic conditions like noise and reverberation, improving sound quality, clarity, and overall expressiveness. The Flash version focuses on optimizing the real-time synthesis experience, keeping first-packet latency under 200 ms, making it suitable for low-latency interactive scenarios such as voice assistants, real-time dialogue, and intelligent customer service.

    Input:$ 15 /M Tokens
    Output:$ 15 /M Tokens
    Context:-
    Latency:-
    Throughput:-

    qwen-audio-3.0-tts-plus is a high-performance speech synthesis large model designed for high-quality speech generation scenarios. Compared with the previous version, the model supports more low-resource languages and Chinese dialects, significantly improving the authenticity of dialect pronunciation, and enhancing free-style instruction following and fine-grained label control capabilities, allowing more accurate control of emotion, tone, character, speaking rate, volume, and synthesis style. Meanwhile, the model exhibits stronger robustness under complex acoustic conditions such as noise and reverberation, further improving sound quality, clarity, resolution, and overall expressiveness. The Plus version places greater emphasis on synthesis quality and detail, making it suitable for professional scenarios that require higher sound quality, naturalness, and expressiveness, such as content creation, audiobooks, film and TV dubbing, brand voice design, and high-quality voice services.

    Quality
    standard
    720x1280
    1280x720
    960x960
    832x1088
    1088x832
    1920x1080
    1080x1920
    1440x1440
    1632x1248
    1248x1632
    standard
    $0.1395
    $0.1395
    $0.1395
    $0.1395
    $0.1395
    $0.1395
    $0.1691
    $0.1691
    $0.1691
    $0.1691
    $0.1691

    HappyHorse-1.1-I2V supports image-to-video generation, further enhancing visual texture, dynamic performance, and cross-segment consistency. The model can more accurately understand input images and carry forward the creative intent, delivering significant improvements in character skin texture, ID consistency across segments, motion smoothness, text rendering stability, and audio–visual synchronization, producing higher-quality videos that are more realistic and natural, rich in detail, and more consistent.