Models

Qwen2.5 VL 72B Instruct Compare

Compare pricing, specifications, performance, and benchmarks for up to four models.

QwenQwen2.5 VL 72B Instruct
Qwen logo
Qwen2.5 VL 72B Instruct
Qwen · text, image, video → text

The model provider is the Sophon platform. Qwen2.5-VL-72B-Instruct is the latest vision-language model released by the Qwen team. This model excels not only at recognizing common objects such as flowers, birds, fish, and insects, but also at efficiently analyzing text, charts, icons, graphics, and layouts within images. As a visual agent, it is capable of reasoning and dynamically guiding tool usage, supporting both computer and mobile operations. Moreover, it can understand videos longer than one hour and accurately locate relevant video segments.

Input$0.62 /M
Output$0.62 /M

Pick a second model to start comparing.

Popular comparisons

Related model match-ups readers also look at.