Xiaomi Mimo V2.5
Xiaomi · text, image, video, audio → text
MiMo-V2.5 is a native, fully multimodal large model designed for agent scenarios; it can see, hear, and read, and translate understanding into action. It has over 1 trillion total parameters (42B active parameters), employs an innovative hybrid-attention architecture, and supports an ultra-long 1M context length. Built on a powerful model base, we continuously scale compute across broader agent scenarios, further expanding the agent’s action space and achieving an important generalization from coding to claw.
Input$0.16 /M
Output$0.31 /M

Xiaomi Mimo V2.5