Token volume is not a receipt.
On September 21, Nathan Lambert’s Interconnects newsletter put a number on open-weight traffic through OpenRouter: about 1 trillion tokens a week in a week of September 2025, about 80 trillion tokens a week now. Chinese models moved from roughly 70 percent of that usage to more than 80 percent. Four days earlier, Vercel’s AI Gateway Production Index, dated September 17 and cited in The New Stack, said open-weight models were 56 percent of gateway tokens and 14 percent of estimated spend. Closed models were 44 percent of tokens and 86 percent of the bill. Anthropic, all models, was 64 percent of estimated spend.
Those two tables are the week’s LLM paper, even though neither one is a paper. Usage moved. The invoice did not.
We already argued that open-weight models are not a charity drop. This week is the accounting. If your slide still says “frontier or nothing,” you are arguing with OpenRouter’s meter and then paying Anthropic anyway.
80 trillion is a router number
OpenRouter is a single interface over open and closed models from the US and China. Lambert says the platform is known for people trying different open-weight models, and that it has published top-model usage since January 1, 2025. The 1T-to-80T jump is that series, not Hugging Face download counts. Downloads are souvenirs. Tokens are inference.
He also names the industrial anecdote layer: software engineering, legal, financial services. Platforms that sell inference on open models — Together, OpenRouter, Fireworks, Baseten — are the first winners he lists in a post-training economy. Fine-tuning APIs such as Thinking Machines’ Tinker sit in the next layer. Staff in the industry, he writes, talk about GLM-5.3 as an alternative to Claude or GPT because of speed, lower prices, customizable offerings, and privacy.
Read that list as four different products. Speed is a latency budget. Price is a token budget. Customizable is a weights-on-your-cluster budget. Privacy is a “the prompt does not leave the VPC” budget. A blog post that collapses those into “open source is winning” is doing marketing. A team that needs one of the four can name which one in a design doc.
The New Stack piece adds an August cut of OpenRouter: open-weight models accounted for roughly 60 percent of US-originating token consumption, with Chinese models making up most of the volume. DeepSeek V4 Pro, Kimi K3, GLM 5.2 are the names in that paragraph. Origin of the weights and location of the GPUs are different questions. The same article notes that OpenRouter is selling control over the second one, and that a Deloitte global survey found 77 percent of companies factor an AI solution’s country of origin into vendor selection.
If procurement asks “is it Chinese,” ask them whether they mean the checkpoint or the region pin. Those have been mixed on purpose in slack threads all year. This week’s numbers make the mix expensive.
Stripe’s announced acquisition of OpenRouter, reported at about $8 billion in that New Stack extract, is a price on the routing layer, not a verdict on any one model. Do not treat a reported deal as a closed trade. Do treat it as a signal that the meter Lambert is reading has a buyer.
Vercel still bills the closed lab
Vercel’s September 17 Production Index is the other meter. Open weights crossed into the majority of that gateway’s token volume for the first time, up from a reported 7 percent in December 2025. Vercel says the current open-weight classification is broader than the one used in earlier reports. That footnote matters. A definition change can look like a landslide.
The spend column did not landside. Open-weight 56 percent of tokens, 14 percent of estimated spend. Closed-weight 44 percent of tokens, 86 percent of estimated spend. Anthropic 64 percent of estimated spend. GPT-6 Astra, in the first 12 days after its September 3 launch, 7.7 percent of estimated spend. Spending is estimated from labs’ published list prices; actual bills may differ. Gateway traffic only.
We already wrote that Astra’s thoughts got harder to read. This index is the other Astra fact: it showed up on a bill fast, and it still did not dethrone Anthropic on that gateway. A 12-day slice is not a quarter. It is also not zero.
Among teams running more than ten million tokens in both comparison months, the median team paid 7.6 percent less per token, after July’s 2.9 percent decline. That is a unit-price story, not a “we stopped using Claude” story. You can burn more tokens at a lower rate and still send most of the money to the same lab. The 86 percent closed-spend number is how.
If you are comparing Interconnects to Vercel, do not average them. OpenRouter is a try-everything router with a documented open-model habit. Vercel’s gateway is production traffic from teams already in that ecosystem, with a classification caveat. Both can be true: open weights can be most of the tokens and a minority of the dollars.
The research implication is dull and useful. Benchmarks that only report quality at the frontier miss the deployment that is actually growing. Benchmarks that only report open-model Elo miss the invoice. Publish both or you are describing a different market than the one on the meter.
A 9 billion parameter model won a domain bench
On September 17, Insilico Medicine described a Cell cover study that is the week’s actual paper-shaped object. Researchers trained a family of five compact, open-source Longevity Large Language Models, 0.6 billion to 9 billion parameters, fine-tuned on aging-specific clinical and multi-omics data in Insilico’s MMAI Gym for Science. The family was built on open architectures from Liquid AI and Alibaba: LFM2, Qwen3, and Qwen3.5.
The specialized models matched or exceeded all 16 frontier systems evaluated on LongevityBench. The best of them, L-Qwen3.5-9B, had the highest overall score among 26 AI systems in the study and outperformed every tested frontier model, including Google’s Gemini 3.1-Pro, at a fraction of the parameters. The smallest specialized model, about 0.6 billion parameters, outperformed most of the frontier systems tested.
That is not “9B beats Gemini at everything.” It is a domain bench with domain weights. The company is also selling a toolkit story; read the leaderboard, not the press-release headline, before you put L-Qwen3.5-9B in a grant. The transferable claim is smaller and matches Lambert’s usage chart: specialization plus open bases is a research program you can run without a frontier training run.
We already watched Paper2Agent wrap a methods paper as an MCP server and still want a human in the loop. LongevityBench is the other direction. Instead of turning a paper into a chat, they turned a literature into a small model and then published the comparison against 16 frontier systems. If your lab’s default is “call GPT and RAG the PDFs,” this paper is the objection: the 0.6B specialist beat most of those calls on that bench.
Insilico points at a Hugging Face collection, a GitHub repo, and a public leaderboard. Use those. Do not scrape a marketing page for a methods section. If you cannot reproduce the split, you do not have a result. You have a cover.
GPT-5-Thinking tied subspecialists on reports, not scans
AuntMinnie summarized a Radiology paper on September 16 from a team led by Jie Li, MD, at The Second Hospital of Jilin University. GPT-5-Thinking processed narrative MRI reports and produced benign-versus-malignant classifications plus top-three differential diagnoses. It did not read pixels. Most deployed imaging tools, the authors note, are image-centric and narrow. This one is a text model sitting on the report.
The study included 1,000 patients, mean age 51 ± 17 years, 520 of them men. Top-one accuracy versus subspecialists was 76.7 percent versus 78.3 percent for orbital tumors (p = 0.64) and 67.5 percent versus 68.0 percent for head-and-neck tumors (p = 0.92). Versus routine generalists, GPT-5-Thinking was higher: 74.0 versus 68.7 percent for orbital (p = 0.03) and 64.0 versus 27.0 percent for head-and-neck (p < 0.001). Seven generalists then read selected reports twice, without and with the model. Assistance moved orbital top-one accuracy from 61.4 to 70.3 percent and head-and-neck from 47.0 to 61.0 percent (both p < 0.001).
Those p-values are why this belongs in llm-research and not in a clinic newsletter. A non-significant gap versus subspecialists on report text is not a license to skip the subspecialist. A large gap versus generalists on head-and-neck reports is a statement about how much specialty language was already in the note, and how much the generalist pass was leaving on the table. The paper’s own sentence is cautious: reasoning-oriented LLMs as subspecialty-informed aids for generalist radiology practice.
Do not take medical advice from this paragraph. Do take the experimental design. The model saw language. The comparison arms were named. The assistance study used the same readers twice. That is more LLM science than another arena screenshot.
Put it next to Insilico anyway. One week, a 9B domain model beats frontier chatbots on a longevity bench, and a frontier reasoning model ties subspecialists on MRI prose. The shared claim is not “AI doctors.” It is that the interesting LLM result this month is a task definition, not a parameter count.
What to measure if you are not Vercel
If you run models, copy the columns, not the brands. Tokens versus spend. Open versus closed. Origin of weights versus region of inference. Domain bench versus general chat. Report-level tasks versus pixel-level tasks.
Lambert’s risk paragraph is the one closed labs will quote: open weights are entering capability levels where new risks, including cybersecurity, can be enabled by many available models, and policy is uncertain. He also says researchers use open models because closed ones such as Claude and GPT often refuse work those researchers consider legitimate, including some defensive cybersecurity and biology. That is a tension, not a tutorial. This site is not going to walk either side’s exploit path. The research question is whether refusal and capability are being measured in the same paper. Right now they usually are not.
We already treated Mercury 2.5’s speed claim as something you have to test. Test this week’s claims the same way. If OpenRouter is 80T, your own gateway logs are the replication. If L-Qwen3.5-9B won LongevityBench, the leaderboard is the replication. If GPT-5-Thinking tied a subspecialist on reports, the paper is the replication, and your hospital is not.
The tokens moved. The bill stayed with the closed models. The 9B specialist still posted a number. Pick which of those three you are actually arguing about before you update the slide.