MiMo-V2-Omni is designed for complex real-world multimodal interaction and execution scenarios. We've built an all-modal foundation from the ground up that fuses text, vision, and speech, and use a unified architecture to deeply bind "perception" and "action." This not only breaks the traditional models' limitation of emphasizing understanding over execution, but also natively equips the model with multimodal perception, tool invocation, function execution, and GUI operation capabilities. MiMo-V2-Omni can seamlessly integrate with major agent frameworks, achieving a leap from understanding to manipulation and significantly lowering the barrier to deploying full-modal agents.
Pricing
Input Modalities
- Text
- Vision
- Audio
- Video
Output Modalities
- Text
Capabilities
- Web
Providers
Xiaomi xiaomi-mimo-v2-omni
Pricing$0.440$2.200
Cache Read$0.088/M tokens
Web Search$0.005/request
Context256K
Max output0
Latency1.7S
Throughput41.3TPS
Uptime
0.00% uptime 2 days ago
0.00% uptime yesterday
0.00% uptime today
Performance for mimo-v2-omni
Uptime is the percentage of requests that succeeded over the past 72 hours. AIHubMix continuously monitors every provider and automatically retries with the next-best provider when one returns an error or responds too slowly; Latency is total round-trip time (lower is better); Throughput is how fast the model writes (tokens per second, higher is better).
Uptime
Loading...
Latency
Loading...
Throughput
Loading...
Try this model
Python

