MiMo-V2-Flash is a mixture of experts (MoE) language model with a total of 309 billion parameters and 15 billion activated parameters. It is designed for high-speed inference and proxy workflows, adopting a novel hybrid attention architecture and multi-token prediction (MTP), significantly reducing inference costs while achieving state-of-the-art performance.
Pricing
- Input Tokens: $0.192 /M tokens
- Output Tokens: $0.575 /M tokens
- Cache Read: $0.038 /M tokens
Input Modalities
- Text
Output Modalities
- Text
Capabilities
- Web
Try this model
Python

