SWE-2 Is on Devin's $20 Plan. The $48B Number Is a Different Product

Cognition raised at $48 billion on September 8, then shipped SWE-2 two days later. The model is a cost claim inside Devin. The round is a bet that coding is not winner-take-all.

Desk with a laptop showing a pull-request diff and a second screen with a terminal, coffee mug, no vendor logos

Cognition spent four days being two companies. On September 8, TechCrunch reported a $2 billion round at a $48 billion valuation, four months after a $26 billion one. Andreessen Horowitz led, with Accel, Founders Fund, General Catalyst, and Avenir. Annualized run-rate revenue, Cognition’s number, went from $492 million in May to $900 million. Two days later the same shop shipped SWE-2, the coding model inside Devin, and put it on Desktop and CLI.

If you write software, the round is gossip. The model is a tool with a price story. This piece is the tool.

What SWE-2 actually claims

Cognition did not pre-train a secret foundation model from scratch. AI/TLDR’s launch record matches the blog: SWE-2 is post-trained from Moonshot’s Kimi K3, which Cognition describes as a 2.8-trillion-parameter model that had already been reinforced for agentic coding. SWE-1.7 used the same recipe on Kimi K2.7.

The headline score is 50.0% on FrontierCode 1.1 Main, Cognition’s own benchmark, against 50.9% for Anthropic’s Fable 5.1 and 53.3% for GPT-6 Astra. Cognition says that near-tie comes at 64% lower cost than Fable, and about a quarter of Astra. Versus SWE-1.7, the company reports 58% fewer turns and 81% lower average cost.

Those numbers are vendor-reported. CellCog said the quiet part: no independent eval, no public per-token price, no model card, no published context window as of launch day. Weights stay inside Devin. There is no standalone API.

The rest of Cognition’s table is more interesting than the one-point FrontierCode photo. SWE-2 leads their published sheet on Terminal-Bench 2.1 at 92.8% and on DeepSWE 1.1 at 73.0%. Terminal-Bench 4, the hard row, is 27.3% against 57.9% for Astra. If your job looks like a long, messy repo, the first two rows matter. If it looks like a hostile benchmark suite, the fourth row is the one that should slow you down.

The $20 seat is the product test

SaaSCity and Cognition’s own launch thread described a month of SWE-2 on Pro, Max, and Teams without chewing the usual quota. Treat a promo as a promo. It can vanish. While it lasts, the experiment is cheap enough to be honest.

Devin is not an autocomplete tab. It is an agent that takes a ticket and comes back with a branch. We already covered Cursor parking cloud agents in Cloudflare sandboxes and the AI-native IDE pile-up. Cursor is still the editor that sits in your keystrokes. Devin is the intern you do not watch. SWE-2 is the intern’s brain, not a new window in VS Code.

That split decides who should bother. If you want inline completions, this launch is not for you. If you already send Devin chores — dependency bumps, test scaffolding, “make the CI go green” — SWE-2 is a model swap with a cost claim attached. Run the same ticket twice, once on the old default if you still have it, once on SWE-2. Count turns. Count the diff you threw away. Ignore FrontierCode until that notebook has rows.

What the $48 billion is buying

TechCrunch’s valuation story is a market argument, not a changelog. Cursor talked at $50 billion in April, then sold to SpaceX for $60 billion, in part because it was compute-starved. Cognition leases an Nvidia cluster that, per The Information via TechCrunch, can run to hundreds of millions a year and help push cash burn toward $800 million. a16z backing both the Cursor outcome and this Cognition round is the tell: investors still think coding assistants are a multi-horse race.

Cognition is trying to get off OpenAI and Anthropic by training on open bases. SWE-2 is that strategy in a shippable file. Kimi K3 in, Devin out. The $900 million run-rate is how they talk to the cap table. Mercedes-Benz, NASA, Goldman, Citi are how they talk to procurement.

None of that changes your laptop. It does change the failure mode. A company burning that much on GPUs will push usage. Unlimited SWE-2 for a month is not charity. It is a load test and a habit test.

How to try it without joining the valuation

Pick one repository you own and one ticket you would have done yourself. Not a greenfield demo. A flaky test, a dependency that needs a pin, a docs page that has been wrong since March. Give Devin CLI the ticket. Time it. Read every file it touches.

The YouTube circuit already ran the Astro-plus-auth demo. Demos compile. Your repo has a house style, a linter you forgot you configured, and a reviewer who hates drive-by refactors. SWE-2’s Terminal-Bench 2.1 number says it can operate a shell. Terminal-Bench 4 says the hard environment still belongs to the frontier labs on Cognition’s own scoreboard.

If the agent opens a pull request, your job is the same as last year: the merge button is yours. Cognition’s FrontierCode metric is “would a maintainer merge this.” That is a useful question. It is not a substitute for you being the maintainer.

For teams, the useful comparison is not Fable versus SWE-2 on a blog table. It is Devin versus a Cursor Cloud Agent on the same internal ticket, with the same secrets policy. Cursor now has a sandbox story. Devin has a model that is cheaper on Cognition’s spreadsheet. Price only exists if both tools are allowed to touch the repo.

What I would not do with this week

I would not migrate an enterprise fleet because a round landed at $48 billion. I would not trust a 50.0% on a first-party benchmark as “on par with Fable.” I would not assume the $20 unlimited window is a price. It is a sample.

I would put SWE-2 on the chores Devin already half-finishes, and I would keep a human on Terminal-Bench-4-shaped work: new architecture, incident response, anything where a wrong edit is expensive. The model is built for long-horizon agent loops. That is a real niche. It is also how you get a 200-line cleanup you did not ask for.

Cognition wants you to feel the cost story. Fair. Run the month. Keep the receipts. If the only thing you learn is that Devin still needs a tight ticket, you learned it on a promo instead of a new SKU.

Secrets, evals, and the missing API

SWE-2 shipping Desktop and CLI first, with Web and Fusion “rolling out,” is a product tell. Cognition wants the model in the loop that already has your repo checkout, not in a public endpoint you could wrap. CellCog noted there was no standalone API on launch day. For a security team that is almost a feature. For a platform team that wanted to A/B SWE-2 against Fable inside their own agent, it is a wall.

If you cannot point the model at a held-out set of internal tickets, you are not evaluating SWE-2. You are demoing Devin. Write ten tickets you already merged by hand. Run them. Score with the only rubric that matches FrontierCode’s marketing: would you merge this without a rewrite. Count rewrites. That number will do more than 50.0%.

Keep secrets out of the agent’s environment unless you already did that work for Cursor Cloud Agents. A cheaper brain with the same GitHub token is not a cheaper incident.

What “post-trained on Kimi” changes in a review

SWE-1.7 sat on Kimi K2.7. SWE-2 sits on K3. That is not a footnote for people who already fight model routing. If Devin’s older runs felt like a chat model pretending to be a janitor, the new base is supposed to have been reinforced for long agent loops before Cognition even started. You still cannot inspect the weights. You can inspect whether the agent stops after the ticket is done.

Watch for three failure modes that chat models always had and agent models still have. One: it “finishes” by editing a test until the test agrees. Two: it refactors a neighbor file because the neighbor was ugly. Three: it leaves a half-applied migration because the shell hung. Terminal-Bench 2.1 says Cognition thinks they got better at the shell. Terminal-Bench 4 says they know they are not done.

If you are a tech lead, write those three failure modes on the review checklist. Do not add “seems smart.” Add “did it touch files outside the ticket,” “did the new tests assert behavior or mock it,” and “can I revert in one commit.” SWE-2 being 81% cheaper than SWE-1.7 on Cognition’s sheet does not make a bad revert cheaper for you.

The $900 million run-rate and the $48 billion round will be in every vendor email this quarter. Delete that paragraph. Keep the checklist. The tool is a model inside Devin, available on Desktop and CLI as of September 10, with a promo on the cheap seat. That is enough product to test. It is not enough product to re-architect a company around.

The round says investors do not think Cursor-plus-SpaceX closed the category. SWE-2 says Cognition’s answer is a cheaper in-house brain, not another chat sidebar. Use the brain. Ignore the ticker.