Mercury 2.5
inception/mercury-2.5
Mercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception. Instead of generating tokens sequentially, Mercury 2.5 produces and refines multiple tokens in parallel, achieving...
Modalities
Text
In / Out price
$0.04 / $0.15 per 1M
Context
260K · 66K out
Released
Sep 8, 2026
Providers
Live list prices and uptime across 1 provider serving this model.
| Provider | Input /M | Output /M | Cache read /M | Context | Max output | Uptime 24h |
|---|---|---|---|---|---|---|
| Inception | $0.04 | $0.15 | $0.004 | 260K | 66K | 99.91% |
No implicit prompt caching; cache reads require explicit cache control.
Capabilities
Tool calling
Reasoning
Include reasoning
Max tokens
Reasoning effort
Response format
Stop
Structured outputs
Temperature
Tool choice
Pricing detail
- Input
- $0.04
- Output
- $0.15
- Cache read
- $0.004
- Cache write
- —
USD per 1M tokens.