inclusionAI/Ling-flash-2.0
InclusionAILing-flash-2.0 is a language model from inclusionAI with a total of 100 billion parameters, of which 6.1 billion are activated per token (4.8 billion non-embedding). As part of the Ling 2.0 architecture series, it is designed as a lightweight yet powerful Mixture-of-Experts (MoE) model. It aims to deliver performance comparable to or even exceeding that of 40B-level dense models and other larger MoE models, but with a significantly smaller active parameter count. The model represents a strategy focused on achieving high performance and efficiency through extreme architectural design and training methods.
Pricing
- Input Tokens: $0.136 /M tokens
- Output Tokens: $0.544 /M tokens
Input Modalities
- Text
Output Modalities
- Text
Capabilities
- Tools
- Tool calling
- Structured outputs
Try this model
Python

