Models

Mercury 2.5 Preview Compare

Compare pricing, specifications, performance, and benchmarks for up to four models.

InceptionMercury 2.5 Preview
80% off
Inception logo
Mercury 2.5 Preview
Inception · text → text

Mercury 2.5 is the latest diffusion-based large language model (dLLM) released by Inception. It is the fastest inference LLM; unlike the sequential token-by-token generation approach, Mercury 2.5 can generate and optimize multiple tokens in parallel, achieving a generation speed of 1,107 tokens per second on standard GPUs. Compared to Mercury 2, its intelligence has increased by more than 10 percentage points, and its quality rivals leading cost-optimized frontier models such as GPT-5.6 Luna (Low), Gemini 3.5 Flash-Lite, and Claude Haiku 4.5.

Input$0.20$0.04 /M
Output$0.75$0.15 /M
Cache read$0.02$0.00 /M

Pick a second model to start comparing.

Popular comparisons

Related model match-ups readers also look at.