Qwen3-Next-80B-A3B-Thinking is a reasoning-first chat model in the Qwen3-Next line that excels by outputting structured 'thinking' traces (Chain-of-Thought) by default.
Designed for hard, multi-step problems, it is ideal for tasks like math proofs, code synthesis, logic puzzles, and agentic planning. Compared to other Qwen3 variants, it offers greater stability during long reasoning chains and is tuned to follow complex instructions without getting repetitive or off-task.
This model is perfectly suited for agent frameworks, tool use (function calling), and benchmarks where a step-by-step breakdown is required. It leverages throughput-oriented techniques for fast generation of detailed, procedural outputs.
Pricing
- Input Tokens: $0.142 /M tokens
- Output Tokens: $1.420 /M tokens
Input Modalities
- Text
- Vision
Output Modalities
- Text
Capabilities
- Thinking
- Tools
- Tool calling
- Structured outputs
Providers
Alibaba Cloud qwen3-next-80b-a3b-thinking
Pricing$0.142$1.420
Context124K
Max output32K
Latency0.5S
Throughput217.6TPS
Uptime
0.00% uptime 2 days ago
100.00% uptime yesterday
100.00% uptime today
Performance for qwen3-next-80b-a3b-thinking
Uptime is the percentage of requests that succeeded over the past 72 hours. AIHubMix continuously monitors every provider and automatically retries with the next-best provider when one returns an error or responds too slowly; Latency is total round-trip time (lower is better); Throughput is how fast the model writes (tokens per second, higher is better).
Uptime
Loading...
Latency
Loading...
Throughput
Loading...
Try this model
Python
