MiMo-V2-Flash is a mixture of experts (MoE) language model with a total of 309 billion parameters and 15 billion activated parameters. It is designed for high-speed inference and proxy workflows, adopting a novel hybrid attention architecture and multi-token prediction (MTP), significantly reducing inference costs while achieving state-of-the-art performance.
mimo-v2-flash vs mimo-v2-flash-free
Time to first token (TTFT) measured on AIHubMix is 3.9s on mimo-v2-flash and 1.6s on mimo-v2-flash-free. Measured output throughput is 58.6 tok/s on mimo-v2-flash and 130.4 tok/s on mimo-v2-flash-free. On the LMArena coding leaderboard mimo-v2-flash scores 1446 and mimo-v2-flash-free scores 1445.
mimo-v2-flash
mimo-v2-flash-freeMiMo-V2-Flash is an open-source foundation language model developed by Xiaomi. It adopts a MoE architecture with 309B total parameters and 15B active parameters per inference, balancing performance and efficiency. The model features a hybrid attention architecture, supports a hybrid-thinking toggle, and offers a 256K context window, enabling strong capabilities in complex reasoning, code generation, and agent-based scenarios. On SWE-bench Verified and SWE-bench Multilingual, MiMo-V2-Flash ranks #1 among open-source models globally, delivering performance comparable to Claude Sonnet 4.5 while costing only about 3.5% as much.
Pricing & Specifications
Prices are per million tokens. Time to First Token and throughput are rolling averages measured on AIHubMix.
Promotional prices show the discounted rate; see each model page for promotion windows.
Activity Past 30 Days
Daily traffic served through AIHubMix — how demand for each model is trending.
Tokens / day
Requests / day
Performance Past 3 Days
Measured on real AIHubMix traffic, hourly buckets. Gaps mean no traffic in that hour.
Throughput (tok/s)
TTFT (s)
Uptime (%)
LMArena Benchmarks
LMArena ratings by capability (Bradley-Terry, commonly called Elo). Higher is better.
Source: LMArena (arena.ai) leaderboard, imported by AIHubMix. Models without published ratings are omitted per chart.
Cost calculator
Estimate your monthly bill for the same workload on each model.
Monthly = daily × 30. Discounted rates applied where a promotion is active.
FAQ
How do their coding arena scores compare?
mimo-v2-flash: 1446; mimo-v2-flash-free: 1445 (LMArena coding leaderboard).
Which responds faster?
mimo-v2-flash-free: 1.6s time to first token measured on AIHubMix; see the live performance charts above for how each model behaves across the day.
Which one generates tokens faster?
mimo-v2-flash-free at 130.4 tok/s and mimo-v2-flash at 58.6 tok/s, measured as output throughput on AIHubMix — a separate metric from time to first token.
What inputs and capabilities does each model support?
mimo-v2-flash accepts text input and supports web search; mimo-v2-flash-free accepts text input and supports web search.
Can I call mimo-v2-flash and mimo-v2-flash-free with the same API key?
Yes. AIHubMix serves every model on this page behind one OpenAI-compatible endpoint, so switching between them is a one-line change to the model field — no second account, key or SDK.
Popular comparisons
Related model match-ups readers also look at.
