Qwen3 VL 30B A3B Instruct
Qwen · text, image, video → text
The Qwen3-VL series’ second-largest MoE model Instruct version offers fast response speed and supports ultra-long contexts such as long videos and long documents; it features comprehensive upgrades in image/video understanding, spatial perception, and universal recognition abilities; it also provides visual 2DD/3D localization capabilities, making it capable of handling complex real-world tasks.
Input$0.10 /M
Output$0.41 /M
