Gme Qwen2 VL 2B Instruct
Qwen · text, image, video → text
The GME-Qwen2VL series is a unified multimodal Embedding model trained based on the Qwen2-VL multimodal large language model (MLLMs). The GME model supports three types of inputs: text, images, and image-text pairs. All these input types can generate universal vector representations and exhibit excellent retrieval performance.
Input$0.14 /M
Output$0.14 /M
