Mercury 2.5
Inception logo

Mercury 2.5

mercury-2.5llms.txt
Inception
New
Mercury 2.5 is Inception's diffusion language model (dLLM), designed for text generation and reasoning tasks. Instead of sequential generation, it produces and refines multiple tokens in parallel to achieve rapid reasoning, supported by an expansive 260,000-token context window.

Pricing

  • Input Tokens: $0.2 /M tokens
  • Output Tokens: $0.75 /M tokens
  • Cache Read: $0.02 /M tokens

Input Modalities

  • Text

Output Modalities

  • Text

Context length

  • 260K tokens

Providers

Inception inception-mercury-2.5
Pricing$0.2$0.75
Cache$0.02
Context260K
Max output0
Latency-
Throughput-
Uptime
0.00% uptime 2 days ago
0.00% uptime yesterday
100.00% uptime today

Performance for mercury-2.5

Uptime is the percentage of requests that succeeded over the past 72 hours. AIHubMix continuously monitors every provider and automatically retries with the next-best provider when one returns an error or responds too slowly; Latency is total round-trip time (lower is better); Throughput is how fast the model writes (tokens per second, higher is better).

Uptime
Loading...
Latency
Loading...
Throughput
Loading...

Try this model

Python
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["AIHUBMIX_API_KEY"],
    base_url="https://aihubmix.com/v1",
)

response = client.chat.completions.create(
    model="mercury-2.5",
    messages=[
      {
        "role": "user",
        "content": "Hello, how are you?"
      }
    ],
    max_tokens=1024,
    stream=False,
)

print(response.choices[0].message.content)

Frequently asked questions

What is Mercury 2.5?

Mercury 2.5 is Inception's diffusion language model (dLLM), designed for text generation and reasoning tasks. Instead of sequential generation, it produces and refines multiple tokens in parallel to achieve rapid reasoning, supported by an expansive 260,000-token context window.