Megadose AI progress, ranked and analyzed.

Inception launches Mercury 2.5 at 1,107 tokens per second

· TestingCatalog ·
Inception is pitching Mercury 2.5 as a diffusion LLM built for production systems where latency is the constraint.

The company says the model is 40% stronger than Mercury 2 while keeping the same low-cost, low-latency profile. It reports 1,107 tokens per second on widely available NVIDIA GPUs, a 260K-token context window, tunable reasoning, parallel tool calls, and schema-aligned JSON. The piece cites deployments in live calling and coding workflows, including OpenCall latency near 170 milliseconds and Augment Code cutting compaction latency from about 150 seconds to 27 seconds. Mercury 2.5 is available through Inception’s chat product and API, Baseten, and OpenRouter, with discounted launch pricing and 100 million free API tokens through Inception. TestingCatalog's note

score 6

Categories: Model Releases