The Qwen3 series of compact visual-understanding models achieves an effective fusion of thinking mode and non-thinking mode, outperforming the open-source Qwen3-VL-30B-A3B with faster response speeds. It comprehensively upgrades image and video understanding, supporting ultra-long contexts such as long videos and long documents, spatial awareness, and universal object recognition; it also possesses visual 2D/3D localization capabilities and is capable of handling complex real-world tasks.
Pricing
Input Modalities
- Text
- Vision
- Video
Output Modalities
- Text
Capabilities
- Tools
- Tool calling
- Structured outputs
Try this model
Python
