Qwen3 Max
Qwen logo

Qwen3 Max

qwen3-maxllms.txt
Qwen
The Tongyi Qianwen 3 series Max model has undergone special upgrades in intelligent agent programming and tool invocation compared to the preview version. The officially released model this time reaches SOTA level in the field and is adapted to more complex intelligent agent scenarios.

Pricing

TierPricingWeb SearchCache WriteCache Read
0<Input<=32K
$0.451$2.705
$0.000548/request$0.5635/M tokens$0.09016/M tokens
32K<Input<=128K
$0.902$5.412
$0.000548/request$1.1275/M tokens$0.1804/M tokens
128K<Input<=256K
$1.352$8.113
$0.000548/request$1.69025/M tokens$0.27044/M tokens

Input Modalities

  • Text
  • Vision

Output Modalities

  • Text

Capabilities

  • Tools
  • Tool calling
  • Structured outputs

Providers

Alibaba Cloud qwen3-max
Pricing$0.451$2.705
Web Search$0.000548/request
Cache Write$0.5635/M tokens
Cache Read$0.09016/M tokens
Pricing$0.902$5.412
Web Search$0.000548/request
Cache Write$1.1275/M tokens
Cache Read$0.1804/M tokens
Pricing$1.352$8.113
Web Search$0.000548/request
Cache Write$1.69025/M tokens
Cache Read$0.27044/M tokens
Context262K
Max output65K
Latency2.1S
Throughput20.0TPS
Uptime
99.94% uptime 2 days ago
100.00% uptime yesterday
100.00% uptime today

Performance for qwen3-max

Uptime is the percentage of requests that succeeded over the past 72 hours. AIHubMix continuously monitors every provider and automatically retries with the next-best provider when one returns an error or responds too slowly; Latency is total round-trip time (lower is better); Throughput is how fast the model writes (tokens per second, higher is better).

Uptime
Loading...
Latency
Loading...
Throughput
Loading...

Try this model

Python
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["AIHUBMIX_API_KEY"],
    base_url="https://aihubmix.com/v1",
)

response = client.chat.completions.create(
    model="qwen3-max",
    messages=[
      {
        "role": "user",
        "content": "Hello, how are you?"
      }
    ],
    max_tokens=1024,
    stream=False,
)

print(response.choices[0].message.content)

Frequently asked questions

What is Qwen3 Max?

The Tongyi Qianwen 3 series Max model has undergone special upgrades in intelligent agent programming and tool invocation compared to the preview version. The officially released model this time reaches SOTA level in the field and is adapted to more complex intelligent agent scenarios.