MiMo-V2.5 is a native, fully multimodal large model designed for agent scenarios; it can see, hear, and read, and translate understanding into action. It has over 1 trillion total parameters (42B active parameters), employs an innovative hybrid-attention architecture, and supports an ultra-long 1M context length. Built on a powerful model base, we continuously scale compute across broader agent scenarios, further expanding the agent’s action space and achieving an important generalization from coding to claw.
xiaomi-mimo-v2.5 vs xiaomi-mimo-v2.5-free
Context lengths are 256,000 tokens on xiaomi-mimo-v2.5 and 256,000 tokens on xiaomi-mimo-v2.5-free. Time to first token (TTFT) measured on AIHubMix is 1.7s on xiaomi-mimo-v2.5 and 1.7s on xiaomi-mimo-v2.5-free. Measured output throughput is 41.3 tok/s on xiaomi-mimo-v2.5 and 41.3 tok/s on xiaomi-mimo-v2.5-free.
xiaomi-mimo-v2.5
xiaomi-mimo-v2.5-freexiaomi-mimo-v2.5-free is the open free version of xiaomi-mimo-v2.5. To maintain reliable service, each account is limited to 5 requests per minute, 500 requests per day, and 1 million tokens per day.
Pricing & Specifications
Prices are per million tokens. Time to First Token and throughput are rolling averages measured on AIHubMix.
Promotional prices show the discounted rate; see each model page for promotion windows.
Activity Past 30 Days
Daily traffic served through AIHubMix — how demand for each model is trending.
Tokens / day
Requests / day
Performance Past 3 Days
Measured on real AIHubMix traffic, hourly buckets. Gaps mean no traffic in that hour.
Throughput (tok/s)
TTFT (s)
Uptime (%)
Cost calculator
Estimate your monthly bill for the same workload on each model.
Monthly = daily × 30. Discounted rates applied where a promotion is active.
FAQ
Which responds faster?
xiaomi-mimo-v2.5: 1.7s time to first token measured on AIHubMix; see the live performance charts above for how each model behaves across the day.
How large is each context window?
xiaomi-mimo-v2.5 accepts 256,000 and xiaomi-mimo-v2.5-free accepts 256,000 input tokens.
Which one generates tokens faster?
xiaomi-mimo-v2.5 at 41.3 tok/s and xiaomi-mimo-v2.5-free at 41.3 tok/s, measured as output throughput on AIHubMix — a separate metric from time to first token.
What inputs and capabilities does each model support?
xiaomi-mimo-v2.5 accepts text, image, video and audio input and supports web search; xiaomi-mimo-v2.5-free accepts text, image, video and audio input and supports web search.
Can I call xiaomi-mimo-v2.5 and xiaomi-mimo-v2.5-free with the same API key?
Yes. AIHubMix serves every model on this page behind one OpenAI-compatible endpoint, so switching between them is a one-line change to the model field — no second account, key or SDK.
Popular comparisons
Related model match-ups readers also look at.
