A Mixture-of-Experts model that activates only 4B parameters per inference,delivering high-performance reasoning with a fraction of the memory cost - idealfor cost-efficient, high-throughput server deployments.
Pricing
- Input Tokens: $0.088 /M tokens
- Output Tokens: $0.385 /M tokens
- Cache Read: $0.011 /M tokens
Input Modalities
Output Modalities
- Text
Try this model
Python
