MiMo V2 Flash
Xiaomi · text → text
MiMo-V2-Flash is a mixture of experts (MoE) language model with a total of 309 billion parameters and 15 billion activated parameters. It is designed for high-speed inference and proxy workflows, adopting a novel hybrid attention architecture and multi-token prediction (MTP), significantly reducing inference costs while achieving state-of-the-art performance.
Input$0.19 /M
Output$0.58 /M

MiMo V2 Flash