Decisions API Answers in 150ms. It Does Not Write the Email

OpenAI’s DevDay Decisions API uses Luna on a finite list and returns in about 150 milliseconds. Codex Cloud is a different SKU. So is the $500 Pro plan.

Laptop showing a short multiple-choice list next to a stopwatch, office daylight, no logos or brand marks

InfoQ’s October 2 recap put a Decisions API in the DevDay pile next to GPT-6.1 Sol, computer use on the Agents API, cloud Codex, and ChatGPT plugins. Startup Fortune is the piece that isolates the product: announced September 29, constrained GPT-6 Luna, a question with a finite set of answers, a probability-scored choice in about 150 milliseconds against roughly 1.6 seconds for a standard Luna call.

We already wrote that Codex got a cloud desk. This is not that desk. A classifier that returns in a tenth of a chat call is a routing tool. If you paste it into the same demo as “the agent wrote my PR,” you will measure the wrong thing and then blame the model.

Finite answers, Luna, 150 milliseconds

Neowin repeats OpenAI’s own framing: using LLMs for decision-making is not an efficient approach, so Decisions API focuses Luna on user-defined questions with finite predefined answers. Context can be text or images. The jobs named are classification, routing, and choosing an agent’s next action. InfoQ says the API is in limited preview and uses the smaller Luna model.

Those constraints are the product. Unlimited generation is ChatGPT. A closed list is a switch. 150 milliseconds, as Startup Fortune reports it, is the number you put next to a request router, not next to a writing assistant. 1.6 seconds for standard Luna is the comparison they printed. If your current chain waits on a 2-second chat completion to decide “refund / escalate / ignore,” this is the SKU that wants that row.

Limited preview means you should not write the migration runbook as if every teammate has a key. It also means latency numbers from a launch post are launch numbers. Re-time them on your traces. If your classifier already returns in 40 milliseconds from a fine-tuned small model you host, 150 milliseconds is not a win. If your classifier is GPT-4-class chat with a JSON prayer, 150 milliseconds is the whole pitch.

Image context is easy to overclaim. The docs in these recaps say you can set context using text or images. That is “the ticket screenshot is an input,” not “the model understands your brand.” Keep the answer set tiny. Four routes. Not forty. Forty is a generation problem wearing a classification hat.

Jev shipped two weeks earlier

Startup Fortune’s news value is the calendar. TypeSafe AI came out of two years in stealth on September 15 with Jev, a model that does this kind of thing and nothing else. Founder: Diogo Almeida, former OpenAI researcher. OpenAI’s Decisions API landed at DevDay on September 29. Two weeks is not a scandal by itself. It is the reason you should compare APIs, not vibes.

What Fortune printed about OpenAI’s side is the latency split (150ms vs 1.6s) and the constrained Luna. What it printed about Jev is the specialty: exactly this, nothing else. If you are buying a router, a one-trick model from a startup and a limited-preview endpoint from OpenAI are the same category. Price, quota, and whether the answer set is actually closed will decide it. A keynote will not.

Do not invent Jev latency. Fortune did not put Jev’s milliseconds in the extract we have. Put OpenAI’s 150ms next to your current chat call. Put Jev on a weekend spike if you can get access. Shipping a router because a keynote had a slide is how you get two classifiers and no owner.

If Almeida’s ex-OpenAI line makes your board nervous, that is a procurement feeling, not a benchmark. Record it as a feeling. Then measure overlap on your actual labels: intent, toxicity bin, next-tool pick. The overlap is the only comparison that matters. Everything else is origin story.

The rest of DevDay is a different SKU

InfoQ’s Codex paragraph is a list you should refuse to fold into Decisions. Cloud environments, CLI voice input, an /agents interface for delegated tasks, a code-review workflow on GitHub pull requests and GitLab merge requests, Codex Security Cloud that scans repos and commits and prepares fixes. We covered the cloud desk. Voice and /agents are CLI features. Security Cloud is a scanner. None of that returns a labeled enum in 150ms.

Agents API computer use is also next door, not inside. InfoQ says applications can operate software through graphical interfaces, with multi-agent capabilities from Codex, tool search, tool calling, and context compaction, OpenAI managing execution infrastructure. Available through the API and in Codex and ChatGPT Work for selected plans. That is a computer-use runtime. It will be slow compared with a classifier because it is doing a different job. If you route with Decisions and act with computer use, you have a pipeline. If you ask computer use to also pick the route, you paid for a GUI loop to flip a boolean.

GPT-6.1 Sol, per InfoQ, is an update aimed at coding, computer use, and professional tasks. OpenAI says it approaches GPT-6 Astra on several evaluations at one-fifth of Astra’s standard input and output token prices. Cached input: $0.10 per million tokens. Available through the API, ChatGPT Work, and Codex. That is a generation model with a price slide. It is not Decisions. Do not put Sol in the router slot because the name is smaller than Astra.

Neowin adds ChatGPT tokens into participating third-party tools: sign-in with ChatGPT across 16 partners, including Cognition’s Devin, Notion, Vercel, T3, OpenClaw, and Dactyl. Subscribers control how much each tool consumes. Eligible usage counts toward existing plan limits. That is billing portability. It does not make Notion a classifier. If your worry last month was plugin installs without a pin, portable tokens do not replace a hash. They replace a second invoice.

$500 Pro is not a classifier

Neowin: ChatGPT Pro 500 is $500 per month, joining $100 and $200 Pro tiers, including access to Astra Ultrafast. OpenAI paused Pro 200, then accepted new Pro 200 subscriptions again with a lower usage allowance. Existing eligible subscribers keep the previous allowance through October 29, 2026. Read that twice if you run seats. New 200 is not old 200. 500 is a different SKU with Ultrafast in the sentence.

@ChatGPT in Slack and Microsoft Teams, per Neowin, lets people ask for help using connected company tools. An individual ChatGPT license is not required to participate in enabled channels. That is a seat-policy grenade. If your legal team thought “no license, no bot,” the channel setting now disagrees. Decisions API is not what they will ping with @ChatGPT. They will ping the chat model and paste a customer email. Keep the products apart in the rollout mail or you will spend November explaining why the $500 plan does not classify refunds faster.

Price math on Sol is the only generation figure worth taking from InfoQ this week if you already have Astra: one-fifth standard token prices, cached input at $0.10 per million. Decisions pricing is not in these recaps. Do not guess it. Put “limited preview, ask for the rate card” in the ticket. A 150ms call that costs like a full chat completion is not the efficiency story OpenAI told.

If you are the person who has to justify Pro 500, attach a trace that needs Ultrafast, not a feeling that $500 sounds like priority. If you cannot show the trace, stay on the allowance you still have through October 29 and re-read the 200-tier change. Seats are not vanity.

InfoQ’s Codex Security Cloud line is another ticket you should not staple to Decisions. It scans repositories and new commits, investigates findings, removes duplicates, and prepares fixes. That is a security product with a write-back path. A 150ms enum does not open a pull request. If your CISO asks whether DevDay “fixed GitHub,” the honest answer is: there is a scanner SKU, there is a review SKU for GitHub PRs and GitLab MRs, and there is a classifier SKU. Buy the one that matches the ticket title.

Where routing belongs in an agent

A useful agent loop, in the boring version, is: classify, then call a tool, then write. Decisions wants the first step. Computer use and Codex want the second. Sol or Astra want the third. Claude Code’s parallel agents already taught people that more workers without a router is just more concurrent mess. OpenAI shipping a router does not mean you delete the workers. It means the workers should not vote on the enum.

Write the enum down before you send a preview request. Refund. Escalate. Ignore. Ask a human. If a product manager wants “be nice,” that is generation. If they want “which queue,” that is Decisions or Jev or the small model you already host. Mixing those in one prompt is how you get a confident paragraph that still went to the wrong queue.

Log the probability. Fortune says OpenAI returns a probability-scored choice. If you throw away the score, you built a coin flip with extra latency. Set a threshold. Below it, human. That is the whole reliability story, and it does not need a keynote.

Do not point this at open-ended moderation of people. A finite list is still a policy. If the list is “allow / block / shadowban,” you own the list. The 150ms call just applies it faster. Faster is not kinder.

What to actually test this week

If you have preview access, time one existing JSON-classifier prompt against Decisions on the same labels. Same inputs. Wall-clock and error mix. If you do not have preview access, time your current chat classifier anyway so the 150ms claim has a baseline when the key arrives.

Keep Codex Cloud, Sol, Pro 500, and Slack @ChatGPT on separate tickets. They shipped in one recap. They do not share a success metric. Cloud desks close laptops. Sol changes a token bill. Pro 500 is a seat. @ChatGPT is a channel permission. Decisions is a list.

If TypeSafe will take a trial, run Jev on the same labels. Two weeks of head start is enough time for a startup to have real bugs and real docs. Use both. Pick one owner. Delete the other from the default path.

And if someone on the team says they will “just use Sol for routing because we already pay for it,” send them the 1.6-second Luna number and the 150-millisecond one. Same family, different product. The email writer can wait. The router should not.