Qwen/Qwen2.5-VL-32B-Instruct
QwenQwen2.5-VL-32B-Instruct is an advanced multimodal model from the Tongyi Qianwen team that can recognize objects, analyze text and graphics in images, operate tools, locate objects in images, and generate structured outputs. Through reinforcement learning, it has improved mathematics and problem-solving capabilities, with a more concise and natural response style.
Pricing
- Input Tokens: $0.240 /M tokens
- Output Tokens: $0.240 /M tokens
Input Modalities
- Text
- Vision
- Video
Output Modalities
- Text
Capabilities
- Tools
- Tool calling
- Structured outputs
Try this model
Python
