Mercury 2.5 Is Fast. That Is the Claim You Have to Test

Inception's diffusion LLM hit 1,107 tokens per second and a 260K window. The launch pricing is cheap. The architecture is the bet, not another Flash-class incremental drop.

Server room aisle with cool white lighting and a laptop on a cart showing a terminal, no readable text or logos

Most model launches this year have been a bigger card of the same recipe. Inception’s Mercury 2.5 is the one that still insists the recipe is wrong.

On September 8, the company said in a Business Wire release that Mercury 2.5 is the most capable diffusion large language model it has shipped, and the fastest reasoning LLM in production, at over 1,100 tokens per second. TestingCatalog filled in 1,107 tokens per second on widely available NVIDIA GPUs, a 40 percent intelligence jump over Mercury 2, a 260K context window, and launch pricing that looks like a fire sale.

Last week’s GPT-6 Astra note was about a model whose thoughts got harder to read. This one is about a model that generates text by a different process and wants you to care about throughput first.

The numbers they chose to lead with

Redwood City, first commercial dLLMs, now a 2.5. Business Wire’s intelligence claim is a 10-point jump over Mercury 2, enough, they say, to sit with cost-optimized frontier models: GPT-5.6 Luna, Gemini 3.5 Flash-Lite, Claude Haiku 4.5. TestingCatalog repeats the comparison and adds Stefano Ermon’s tweet from the same day: most capable diffusion LLM on the market, 40 percent jump, over 1,100 tokens per second.

Those two percentages will get copy-pasted. They are not the same measurement. “10 points” is a scoreboard move. “40 percent” is a relative jump. Neither article publishes the benchmark table in the excerpts we have. Treat both as vendor language until you see the eval suite.

Throughput is the number they want in the headline. 1,107 tokens per second is specific enough to be testable. Ryan Cole at shattered.io is right that they led with production speed instead of a single quality score, and that the word “production” is doing work. Lab tokens per second die on messy prompts and shared GPUs. If you route real traffic, you will find out whether 1,107 survives your mix of short chat, long JSON, and tool calls.

Context: 260K, up from 128K. That is a real change if you were stuffing retrieval into 128K and spilling. It is also the same class of number every lab now prints. We already covered how context windows stopped being the whole story. A bigger window is not a memory architecture. It is permission to spend more on the prompt.

Product surface, from the release: tunable reasoning, native tool use, JSON mode, parallel tool calls. That is the checklist you would expect if they want to steal Flash-class agent traffic, not just demo autocomplete.

What the price is doing

List: $0.20 per million input tokens, $0.75 per million output. Launch: 80 percent off, $0.04 / $0.15. TestingCatalog and Business Wire agree on those figures.

Cheap input is how you get people to try a weird architecture. Cheap output is how you keep them if the model is verbose. Diffusion models that generate in parallel can, in theory, make that output cheaper to produce. Theory is not your invoice. The discount expires; the architecture either holds the latency or it does not.

Do not build a unit-econ slide on $0.04. Build it on $0.20 / $0.75 and treat the launch rate as a trial coupon. If Mercury 2.5 only wins because of the coupon, you do not have a model decision. You have a promo.

Why diffusion is the actual news

Autoregressive models emit a token, then another. Diffusion language models iterate toward a sequence. Inception has been selling that difference since Mercury 2: speed as a property of the method, not a smaller student model distilled from a giant teacher.

TestingCatalog says the company is calling 2.5 the largest dLLM ever trained. “Largest” without a parameter count is a press phrase. What you can use: they are not positioning this as a tiny fast head on top of someone else’s weights. They are asking you to believe a full diffusion stack can land in the Luna / Flash-Lite / Haiku quality band and still print four-digit tokens per second on GPUs you can actually buy.

That is the bet shattered.io named. The other launches this week are bigger or more open versions of an established recipe. Mercury 2.5 wants the recipe changed. If you evaluate it like another 8B instruct drop, you will measure the wrong thing. Measure:

  • tokens per second on your prompt mix, not their slide
  • quality on the tasks you pay for (JSON schema hits, tool-call validity, long-context retrieval)
  • tail latency, not averages
  • what happens when tunable reasoning is turned up — speed claims tend to be for the fast setting

We wrote about how thin some safety circuitry is. A new generation method does not inherit those papers. If you put Mercury 2.5 on customer traffic, run your own refusal and injection tests. Do not assume Haiku-class comparison means Haiku-class behavior.

What to do if you actually might buy it

You do not need a research intern. You need a weekend harness.

Take 200 prompts from production: the short ones, the ugly ones, the ones that already fail JSON mode. Run them against your current Flash-class endpoint and against Mercury 2.5 at list-equivalent settings, not only at the coupon. Log tokens per second, schema-valid rate, and the p95. If they will not give you an SLA or a rate limit in writing, write that down too. shattered.io noted that early coverage was thin on those operational details. Missing SLA is a fact. It is not a smear.

If you do not have 200 prompts, you are not in the market yet. Read the Business Wire and wait for someone else to burn the coupon.

For teams that live in long context, 260K is the experiment: one-shot a repo chunk you currently split. If quality falls off at 200K, you learned something the window size did not tell you.

For agent tooling, parallel tool calls and JSON mode are the features that either save you glue or become another dialect. Test the dialect. Native tool use in a press release has meant five different wire formats this year.

What this is not

It is not a reason to rip out GPT-5.6 Luna or Gemini Flash-Lite on Tuesday. It is not proof that autoregressive decoding is over. It is not a safety paper. It is a speed claim attached to a quality claim attached to a discount.

Inception’s earlier rounds, the same Business Wire page reminds you, were Mercury 2 as a fast reasoning LLM and a $50 million raise to make diffusion cheap enough for wearables and data centers. 2.5 is the “we still mean it” release. The honest response from an engineering org is to put a slice of traffic on it or to ignore it. Sitting in a meeting repeating 1,107 is neither.

A note on the comparison set. GPT-5.6 Luna, Gemini 3.5 Flash-Lite, and Claude Haiku 4.5 are the names Inception wants next to 2.5. Those models are not interchangeable with each other, and they are not interchangeable with a diffusion stack you have not run. If your current default is already one of those three, the only comparison that matters is yours against yours. If your default is a larger reasoning model, Mercury 2.5 is not a drop-in. Tunable reasoning is the feature that might close that gap. It is also the feature that usually destroys a tokens-per-second headline. Measure both knobs.

Hardware is the other quiet claim. TestingCatalog’s 1,107 figure is on widely available NVIDIA GPUs, not a custom inference chip. That is helpful if your cluster is already full of the same cards. It is not a promise about batch size, speculative tricks, or whether 260K context still prints four digits when the KV cache is hot. If Inception will not show a latency-vs-context curve, generate a crude one with 4K, 32K, and 128K prompts. The shape of that curve is the product.

Procurement will ask about training data, safety reports, and whether “largest dLLM ever trained” comes with a model card. As of the launch write-ups, you have a press release, a tweet, and a couple of explainers. That is enough to trial. It is not enough to put in a regulated workflow. Keep it in the sandbox with the other speed models until the card exists. Ask for the card in the same email as the rate limit.

Diffusion also changes how you think about streaming UX. Autoregressive UIs can show a token as soon as it exists. A diffusion pass that refines a block may want a spinner, a draft dump, or a two-stage reveal. If your product is a blinking cursor, Mercury 2.5 might feel worse even when the wall clock is better. If your product is “give me the JSON,” wall clock is the UX.

On cost, watch the output side. $0.75 per million at list is still cheap next to many reasoning models, but agent loops are output-heavy. Parallel tool calls can multiply that. The launch coupon hides it. A week of traces will not.

Open-source diffusion LLMs will show up in someone’s Discord as a counter. That is a different procurement problem: you still need GPUs, a serving stack, and someone who understands the sampler. Inception is selling that stack as an API. Compare API to API first. Compare weights later if you have the people.

Customer support will try to turn 1,107 into a guaranteed SLO. It is not. Put the vendor number in the appendix. Put your p95 in the decision. If they cannot give you a region list and a rate limit, you are not buying production. You are buying a beta with a price list.

If the production number holds on your mix, you just got a new default for high-volume, low-stakes generation: drafts, rewrites, tool-loop chatter. If it does not hold, you lost a week and kept your old bill. That is a cheap experiment even at list price. Run it like an experiment. Do not write architecture decision records based on a tweet from the inventor of the method.