80% off
Mercury 2.5 Preview
Inception · text → text
Mercury 2.5 is the latest diffusion-based large language model (dLLM) released by Inception. It is the fastest inference LLM; unlike the sequential token-by-token generation approach, Mercury 2.5 can generate and optimize multiple tokens in parallel, achieving a generation speed of 1,107 tokens per second on standard GPUs. Compared to Mercury 2, its intelligence has increased by more than 10 percentage points, and its quality rivals leading cost-optimized frontier models such as GPT-5.6 Luna (Low), Gemini 3.5 Flash-Lite, and Claude Haiku 4.5.
Input$0.20$0.04 /M
Output$0.75$0.15 /M
Cache read$0.02$0.00 /M

Mercury 2.5 Preview