Skip to content
TrustList
All launches
AI model launch

Mercury 2.5

A diffusion language model that writes text in parallel — 1,107 tokens a second, 260K context, $0.20/$0.75.

Launched 2026-09-08AI model

About this launch

Mercury 2.5 was launched by Inception on 8 September 2026. It is a diffusion language model: rather than producing one token after another, it generates and refines chunks of text in parallel, which is where its speed comes from. Inception states it runs at 1,107 tokens per second on widely available NVIDIA GPUs and describes it as the fastest reasoning model in production; that is a vendor-reported figure that has not been independently reproduced. The context window is 260K tokens. Regular pricing is $0.20 per million input tokens and $0.75 per million output, and at launch it was offered at 80% off ($0.04 and $0.15) with no end date published; the regular rate is recorded here. Inception claims a 40% increase in intelligence over Mercury 2 and quality comparable to cost-optimised frontier models such as Claude Haiku 4.5 and Gemini 3.5 Flash-Lite, but names no benchmark behind the 40% figure, so no score is stored. The two customer results in the announcement are also vendor-reported: Augment Code cutting a compaction step from roughly 150 seconds to 27, and OpenCall reporting a median response latency close to 170 milliseconds. Mercury 2.5 is available through the Inception API, Baseten and OpenRouter, with dedicated enterprise deployments. No licence or parameter count is stated.

Discussion

Questions & answers about Mercury 2.5.

No questions yet. Be the first to ask about Mercury 2.5.