Ling-flash-2.0 is a language model from inclusionAI with a total of 100 billion parameters, of which 6.1 billion are activated per token (4.8 billion non-embedding). As part of the Ling 2.0 architecture series, it is designed as a lightweight yet powerful Mixture-of-Experts (MoE) model. It aims to deliver performance comparable to or even exceeding that of 40B-level dense models and other larger MoE models, but with a significantly smaller active parameter count. The model represents a strategy focused on achieving high performance and efficiency through extreme architectural design and training methods.
← Models
inclusionAI/Ling-flash-2.0
inclusionAI/Ling-flash-2.0 Compare
Compare pricing, specifications, performance, and benchmarks for up to four models.
inclusionAI/Ling-flash-2.0+ Add model · 1/4
inclusionAI/Ling-flash-2.0
InclusionAI · text → text
Input$0.14 /M
Output$0.54 /M
Pick a second model to start comparing.
Popular comparisons
Related model match-ups readers also look at.
