Inception: Mercury 2.5

Inception mercury-2-5
Model Information
Slug mercury-2-5
LLMs.txt View
Release Date September 8, 2026 New
Organization
Model Description
Mercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception.
Instead of generating tokens sequentially, Mercury 2.5 produces and refines multiple tokens in parallel, achieving 1,107 tokens/sec on standard GPUs. It delivers a 10+ point jump in intelligence over Mercury 2, comparable quality to cost-optimized frontier models like GPT-5.6 Luna (Low), Gemini 3.5 Flash-Lite, and Claude Haiku 4.5.
Mercury 2.5 supports tunable reasoning levels, parallel tool calls, and schema-aligned JSON output. It's built for production workloads where latency compounds: search agents, voice pipelines, and coding subagents. Read more in the [blog post](https://www.inceptionlabs.ai/blog/introducing-mercury-2-5).
Available at 7 Providers
Provider Type Model Name Original Model Input ($/1M) Output ($/1M) Free Actions
Nous Research
Nous Research
Inception: Mercury 2.5
inception/mercury-2.5 $0.032 $0.12
Cline
Cline
Code
Inception: Mercury 2.5
inception/mercury-2.5 $0.04 $0.15
OpenRouter
OpenRouter
Chat Code
Mercury 2.5
inception/mercury-2.5 $0.04 $0.15
Vercel AI Gateway
Vercel AI Gateway
Mercury 2.5
inception/mercury-2.5 $0.04 $0.15
Kilo Code
Kilo Code
Code
Inception: Mercury 2.5
inception/mercury-2.5 $0.20 $0.75
Krater
Krater
Inception: Mercury 2.5
mercury-2-5 - -
ValorGPT
ValorGPT
IN
inception-mercury-2.5 - -