Mellum2.1 Is 2.5B Active. SWE-bench Is Self-Reported

JetBrains shipped Mellum2.1 on Oct 8 on the same 12B MoE. The Hugging Face card lists 47% SWE-Verified. GGUF is still coming soon.

Developer workstation with a JetBrains-style IDE on one monitor and a printed model card showing 12B total and 2.5B active parameters, cool desk lamp, no logos

JetBrains posted Mellum2.1 as the next 12B mixture-of-experts drop, Apache 2.0, same skeleton as the June Mellum2 open-source. shattered.io dates the ship to October 8 and pulls 47.0 percent on SWE-bench Verified off the Hugging Face card for JetBrains/Mellum2.1-12B-A2.5B-Thinking. That number is JetBrains’ card. It is not an independent rerun. GGUF builds for llama.cpp, Ollama, and LM Studio are listed as coming soon. If your plan was to drop a Q4 into LM Studio tonight, the blog already told you no.

We already treated Claude Code’s hole versus Copilot as an IDE-agent story and Codex’s cloud desk as a laptop you can close. This week is the small model you host. 2.5 billion parameters active per token is the inference claim. 47 percent is the agent claim. Keep them on separate lines.

The architecture did not move. Post-training did

JetBrains is blunt about what did not change. Mellum2.1 is still a compact MoE with 2.5B active parameters. The June post called Mellum2 a model for routing, Q&A, sub-agents, and private AI in software engineering systems. The October post keeps that sentence and adds a narrower one: trained with reinforcement learning in real environments, built for coding agents and fast sub-agents that run on your own hardware.

What changed is everything after pre-training. RL went from a short final stage to the main event. They added tasks in math, competitive programming, science, tool use, and software engineering, mixing open RL datasets with in-house work. Open data, they say, often arrives with broken tests, unverifiable answers, or tasks that are too easy or impossible, so they filtered before training. They built infrastructure to run thousands of RL environments and launched millions of sandboxes.

That is a training-ops story. It is not a new tokenizer. It is not a new expert count. Tech Insider still describes 64 experts, 8 active, 28 layers, a Qwen3-MoE-derived layout with per-layer-type rotary embeddings and interleaved sliding-window attention, and a 131,072-token context on the Thinking family. shattered.io adds hidden size 2,304, grouped-query attention with 32 query heads and 4 key-value heads, and sliding-window attention on three of every four layers with a 1,024-token local window. Those details are how you get 12B on the card and 2.5B in the forward pass. They are also how you should read “fast.” Speed is the sparse path, not a marketing adjective.

Checkpoints on Hugging Face: Base, Instruct, Thinking. The Thinking variant is the one the published scores refer to. It emits chain-of-thought before the answer. If you wanted autocomplete, that is still closer to original Mellum, the 2024 dense completion model JetBrains later put on Apache 2.0 and Bedrock Marketplace. Do not install Thinking and then complain it talks too much in the IDE.

47 percent is a card. 17.4 percent is also a card

shattered.io tabulated the Hugging Face self-report. HumanEval+ 91.5 percent. LiveCodeBench v6 82.0. MBPP+ 79.4. GSM-Plus 88.3. AIME 25/26 83.3. MMLU-Redux 87.8. GPQA Diamond 64.6. BFCL v4 62.3. ToolHop 49.1. WorkBench 44.6. SWE-bench Verified 47.0. SWE-bench Pro 28.0. Terminal-Bench 2.1 17.4.

Read the drop. Single-function coding is where a 12B MoE looks like a product. SWE-bench Pro and Terminal-Bench are where it looks like a sub-agent. JetBrains’ own blog says the biggest improvement versus Mellum2 is agentic coding, and that they compared Mellum2.1 with Mellum2, Qwen3.5-9B, and Gemma 4 E4B under the same setup. The charts are images. The 47 percent figure is the card, via shattered, not a sentence in the JetBrains HTML. If you cite it, cite the card. If you buy it, rerun SWE-Verified on your harness. Single-run vendor tables are still single-run vendor tables.

Tech Insider, on October 8, explicitly said JetBrains had not published head-to-head tables against Codestral, Qwen Coder, StarCoder, or DeepSeek Coder on HumanEval, MBPP, or SWE-bench, and that any “beats X” claim should wait. That caution still applies to the 47 percent. It is a number on a model page. It is not a reason to delete Claude from the hard tickets.

JetBrains also says that under heavy load Mellum2.1 is the fastest in that comparison group and serves almost twice as many tokens as Qwen3.5-9B, and that multi-token prediction makes a single request about 1.6 times faster. MTP for speculative decoding in vLLM is coming soon, same bucket as the GGUF files. Today’s speed claim is the MoE path they already had in June. Tomorrow’s 1.6x is a head that is not in your cluster yet.

Coming soon is the install

The collection is JetBrains/mellum21 on Hugging Face. The Thinking repo name is the long A2.5B string. Quantized GGUF builds are mentioned in secondary writeups as intended, including a Q4_K_M, and the official post groups GGUF for llama.cpp, Ollama, and LM Studio with the MTP head as not ready.

If your constraint is an air-gapped GPU box with vLLM, you can start from the HF weights. If your constraint is a laptop with Ollama, you wait. That split is the actual product boundary this week. Apache 2.0 does not mean a one-click Modelfile.

JetBrains lists three jobs: a worker inside agentic systems (root-cause a failing test, draft a fix, check it); a general assistant for everyday questions and hard math; private self-hosted deployment so code stays inside the building. The first and the third are the reason a bank talks to JetBrains. The second is how you demo it. Do not demo SWE-Pro 28 percent as if it were Verified 47.

Original Mellum, per shattered, was a roughly 4B dense completion model, 8,192 context, trained on about 4 trillion tokens of permissively licensed code. It finished the line you were typing. Mellum2 on June 2, 2026, was the 12B rebuild. Four months of RL is the delta you are being asked to trust. Millions of sandboxes is a compute receipt. It is not your repo.

The June technical report on arXiv, 2605.31268, is still the architecture paper. It is the one to keep. Mellum2.1 does not replace it. If you are writing an internal model-risk note, attach both: the June report for the MoE, the October blog for the RL recipe, the Hugging Face card for the numbers you have not rerun. A screenshot of the SWE-Verified chart from the blog is not a citation. The card is.

A Gemini coworker with Gmail is a different SKU

The same Thursday, VentureBeat reported Google Cloud’s persistent Gemini agents. Thomas Kurian’s pitch: a universal workplace agent with a common interface and API across web, iOS, Android, Windows, Mac, CLI, Google Workspace, Microsoft 365, and Slack. Configurations include a personal assistant and a persistent digital coworker. The coworker can receive its own Workspace account — email, calendar, Drive, directory presence. You @mention it in Chat. Work can run for days. Sub-agents can be spun up. Execution is in the cloud, so you start on one device and resume on another.

Google told VentureBeat the two configurations share architecture, memory, business context, personalization, and security controls. Google has not said every agent gets a Workspace mailbox, or whether extra Workspace licenses are required, or how admins toggle Gmail versus Calendar. Agents get a cryptographically attested identity. Complex tasks run in isolated containers with restricted outbound traffic. An Agent Gateway applies company rules. A spokesperson said agents can access only files explicitly shared with them. Actions are logged against the agent identity. Prompts and docs are supposed to stay inside Gemini Enterprise even when a third-party model such as Claude is in the mix.

That is an identity and mailbox problem. Mellum2.1 is a weights problem. If you need a bot that owns a calendar, you are not choosing a 12B GGUF. If you need a sub-agent that never leaves the VPC, you are not choosing a Gemini coworker until legal finishes the mailbox question.

VentureBeat’s competitive frame is the rest of September and early October: Microsoft’s updated Copilot on September 25 (Home, Cowork, Code, Autopilot in preview), OpenAI Dots at DevDay on September 29, Anthropic’s Claude for Google Workspace on October 6. Those are product shells. Mellum2.1 is a checkpoint. Mixing them in one “AI coworker” roundup is how you buy the wrong thing.

We already warned that a plugin install is not saved by a SHA pin. A self-hosted 12B still executes the patch it writes. Sandboxes in JetBrains’ training are not the sandbox in your CI. Put the agent in a runner that cannot push to main.

What to run on Monday

Download the Thinking weights if you have a GPU and a vLLM box. Do not tell the team it is 47 percent until you have your own SWE-Verified slice, even a small one. Compare against Qwen3.5-9B on the same box, because that is the comparison JetBrains already picked. Measure tokens per second under load. If you cannot beat “almost 2x Qwen” in your stack, the blog’s speed chart is not your chart.

If you only have Ollama, wait for the GGUF. Using a one-off conversion you found in a Discord is how you debug a tokenizer for a week.

If the actual request from the business is “an agent with an email address,” read the Gemini piece and the unanswered license questions, then talk to Workspace admin. Mellum will not grow a Gmail.

If the actual request is “stop sending the monorepo to a frontier API for routine edits,” Mellum2.1 is in the right size class. Route the hard tickets to the model you already pay for. Keep the 2.5B-active path on the high-volume, low-horizon jobs: failing test, small patch, self-check. That is the sentence JetBrains wrote. The 47 percent is the sentence the internet will repeat. They are not the same sentence.

One more operational limit. Sliding-window attention on three of four layers, with a 1,024-token local window, is why a 131K context is affordable. It is also why “paste the whole monorepo” is a worse prompt than “open the failing test and the two files it imports.” A sub-agent that is fast on a scoped ticket will still wander if you give it the company as context. JetBrains Context, the repository-intelligence layer they put in early access in July, is a different product page. Do not assume 2.1 includes it. Weights are weights. Indexing is indexing.