Qwen3 VL 30B A3B Thinking
Qwen · text, image, video → text
The Qwen3-VL series’ second-largest MoE model Thinking version offers fast response speed, stronger multimodal understanding and reasoning, visual agent capabilities, and ultra-long context support for long videos and long documents; it features comprehensive upgrades in image/video understanding, spatial perception, and universal recognition abilities, making it capable of handling complex real-world tasks.
Input$0.10 /M
Output$1.03 /M
