MiMo V2 Omni
Xiaomi logo

MiMo V2 Omni

mimo-v2-omni
Xiaomi
MiMo-V2-Omni is designed for complex real-world multimodal interaction and execution scenarios. We've built an all-modal foundation from the ground up that fuses text, vision, and speech, and use a unified architecture to deeply bind "perception" and "action." This not only breaks the traditional models' limitation of emphasizing understanding over execution, but also natively equips the model with multimodal perception, tool invocation, function execution, and GUI operation capabilities. MiMo-V2-Omni can seamlessly integrate with major agent frameworks, achieving a leap from understanding to manipulation and significantly lowering the barrier to deploying full-modal agents.

Pricing

PricingCache ReadWeb Search
$0.440$2.200
$0.088/M tokens$0.005/request

Input Modalities

  • Text
  • Vision
  • Audio
  • Video

Output Modalities

  • Text

Capabilities

  • Web

Providers

Xiaomi xiaomi-mimo-v2-omni
Pricing$0.440$2.200
Cache Read$0.088/M tokens
Web Search$0.005/request
Context256K
Max output0
Latency1.7S
Throughput41.3TPS
Uptime
0.00% uptime 2 days ago
0.00% uptime yesterday
0.00% uptime today

Performance for mimo-v2-omni

Uptime is the percentage of requests that succeeded over the past 72 hours. AIHubMix continuously monitors every provider and automatically retries with the next-best provider when one returns an error or responds too slowly; Latency is total round-trip time (lower is better); Throughput is how fast the model writes (tokens per second, higher is better).

Uptime
Loading...
Latency
Loading...
Throughput
Loading...

Try this model

Python
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["AIHUBMIX_API_KEY"],
    base_url="https://aihubmix.com/v1",
)

response = client.chat.completions.create(
    model="mimo-v2-omni",
    messages=[
      {
        "role": "user",
        "content": "Hello, how are you?"
      }
    ],
    max_tokens=1024,
    stream=False,
)

print(response.choices[0].message.content)

Frequently asked questions

What is MiMo V2 Omni?

MiMo-V2-Omni is designed for complex real-world multimodal interaction and execution scenarios. We've built an all-modal foundation from the ground up that fuses text, vision, and speech, and use a unified architecture to deeply bind "perception" and "action." This not only breaks the traditional models' limitation of emphasizing understanding over execution, but also natively equips the model with multimodal perception, tool invocation, function execution, and GUI operation capabilities. MiMo-V2-Omni can seamlessly integrate with major agent frameworks, achieving a leap from understanding to manipulation and significantly lowering the barrier to deploying full-modal agents.