Mercury 2.5

inception/mercury-2.5

Mercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception. Instead of generating tokens sequentially, Mercury 2.5 produces and refines multiple tokens in parallel, achieving...

Modalities
Text
In / Out price
$0.04 / $0.15 per 1M
Context
260K · 66K out
Released
Sep 8, 2026

Providers

Live list prices and uptime across 1 provider serving this model.

ProviderInput /MOutput /MCache read /MContextMax outputUptime 24h
Inception$0.04$0.15$0.004260K66K99.91%

No implicit prompt caching; cache reads require explicit cache control.

Capabilities

Tool calling
Reasoning
Include reasoning
Max tokens
Reasoning effort
Response format
Stop
Structured outputs
Temperature
Tool choice

Pricing detail

Input
$0.04
Output
$0.15
Cache read
$0.004
Cache write
—

USD per 1M tokens.