gme-qwen2-vl-2b-instruct
Qwen logo

Gme Qwen2 VL 2B Instruct

gme-qwen2-vl-2b-instruct
Qwen
The GME-Qwen2VL series is a unified multimodal Embedding model trained based on the Qwen2-VL multimodal large language model (MLLMs). The GME model supports three types of inputs: text, images, and image-text pairs. All these input types can generate universal vector representations and exhibit excellent retrieval performance.

Pricing

  • Input Tokens: $0.138 /M tokens
  • Output Tokens: $0.138 /M tokens

Input Modalities

  • Text
  • Vision
  • Video

Frequently asked questions

What is gme-qwen2-vl-2b-instruct?

The GME-Qwen2VL series is a unified multimodal Embedding model trained based on the Qwen2-VL multimodal large language model (MLLMs). The GME model supports three types of inputs: text, images, and image-text pairs. All these input types can generate universal vector representations and exhibit excellent retrieval performance.