OpenAI's Decisions API vs Jev: Inside the Decision-Model Architecture

Hiba FathimaHiba Fathima
Sep 30, 2026

TL;DR

  • OpenAI's Decisions API is a limited-preview endpoint that constrains GPT-6 Luna to developer-defined questions with pre-defined answers. Context goes in as text or images, and a selection comes back.
  • OpenAI's own DevDay slide claims 150 ms against 1.6 s for a standard GPT-6 Luna call, a 10x speedup. The baseline is OpenAI's own cheapest model.
  • The contract is not public. As of September 30, 2026, OpenAI has published no request schema, response schema, endpoint path, SDK method, or pricing.
  • Jev, from TypeSafe AI, is the incumbent. It is a purpose-built decision model, generally available, with documented Choice, Score, and Noul primitives, per-option probabilities, and free output tokens at $0.042 per million input tokens.
  • The architectural split is constrained LLM versus purpose-built model. It shows up in pricing, in whether output can ever be off-schema, and in whether the model can also write prose.

Your agent needs to know one thing: is this shell command reversible? That is a yes or no answer, and getting it costs a full LLM round trip. Build a prompt, generate tokens, parse the reply, validate the shape, retry when it drifts. Do that at every branch in the loop and the control flow costs more than the work.

On September 29, 2026, OpenAI put a name on the fix. The Decisions API, announced at DevDay 2026, focuses GPT-6 Luna on a bounded question with a finite list of allowed answers and returns a selection your code can branch on. It arrived fourteen days after TypeSafe AI shipped Jev, a model built only for this job.

What is OpenAI's Decisions API?

OpenAI's Decisions API is an endpoint that answers one narrow question about supplied context by selecting from a list of answers you define in advance. It returns a value your application branches on and does no drafting, summarizing, or explaining.

OpenAI's entire public description is one paragraph in the DevDay 2026 recap:

Decisions API enables real-time decision-making by focusing Luna's intelligence on a specific set of user-defined questions with finite pre-defined answers. Developers supply context using text or images, and get back answers they can use to classify content, route requests, or choose an agent's next action.

OpenAI's Decisions API key art from the DevDay 2026 recap, white title text on a starfield background

The three named jobs (classifying content, routing requests, and choosing the next action) are all control flow. Access is limited for now. The recap says the API is "available in limited preview today with a broad release planned in the coming days," and @OpenAIDevs clarified that "preview access is limited to selected API customers for testing."

Sam Altman showed it on the DevDay keynote stage driving a computer-use agent:

The decisions API lets the model respond in a fraction of a second... You can see here it using a computer and how quickly it's moving through these steps. This is not sped up. This works by giving our Luna model a predefined set of options to choose from.

Why does an agent need a decision API at all?

Many of an agent's turns go to small questions: which tool to call next, whether a page is relevant, whether an action needs confirmation, which model tier can handle a task. None of those need prose, but a chat model answers each one with generated text anyway.

The chat-shaped pipeline for a yes-or-no question runs five steps before anything happens:

prompt -> generated text -> parser -> schema validation -> retry -> action

A decision-shaped pipeline runs three, and the retry branch mostly disappears:

state + bounded question -> typed decision -> policy check -> action

Besides tokens, you drop the parser, the retry path, and the class of bugs where a model answers the right question in the wrong format. Prompt-and-parse code tends to grow into a small string-handling codebase that nobody wanted to own.

The obvious objection is that a decision API is just structured outputs under a new name. Structured outputs constrain the shape of generated text after the model has decided what to say. A decisions endpoint bounds the answer space before inference runs, which is why the round trip gets shorter as well as tidier.

Hacker News was not entirely convinced by that distinction. In the thread that predicted this launch a week early, alex_sf pushed back on the whole category: "It's not guaranteed to be correct: it's guaranteed to be formatted in a particular way. You can get the same thing with grammars on any LLM."

How does the Decisions API work?

Even without a published schema, the design reduces to four steps.

  1. Capture the state. The evidence the decision needs: a ticket, a scraped page, a proposed tool call, a screenshot. Send what a human reviewer would need to answer this one question, and nothing else.
  2. Define a bounded question. Specific enough that two engineers would agree on what a correct answer is. "Which of these four approved routes applies?" beats "What should we do?"
  3. Select from the answer space. You supply the options. The model picks one rather than inventing a next action.
  4. Enforce the result in code. Treat the decision as input to your policy code, which keeps ownership of permissions, resource scope, rate limits, and business policy.

A decision model can judge meaning, but authority has to live in your code. A 0.97 on "this looks safe" does not grant an account permission it did not already have.

What OpenAI has not published

As of September 30, 2026, there is no public contract for the Decisions API.

I checked OpenAI's API guides index, its endpoint reference index, its 5 MB combined documentation export, the public OpenAPI spec, and the Python, Node, and Go SDKs, and none of them include it. In the DevDay recap itself, Decisions API is the only announcement in its section with no "Learn more" link. Vercel, an OpenAI launch partner, notes that OpenAI's announcement "does not specify whether Decisions API returns option probabilities, rubric scores, or yes-or-no probabilities."

Every field name, response shape, and price is unknown for now, so any blog post showing an OpenAI Decisions request body is made up.

QuestionStatus as of Sep 30, 2026
Endpoint path, auth, SDK methodNot published
Request and response schemaNot published
Probabilities or confidence returnedNot published
Multiple questions per callNot published
Image inputConfirmed
Context limit for DecisionsNot published
Pricing, output-token billingNot published

OpenAI's published latency claim

OpenAI DevDay slide comparing Decisions API at 150 ms against the GPT-6 Luna API at 1.6 seconds for task completion time

OpenAI's DevDay slide, reproduced by The Decoder, reads: "Helps apps and agents take action 10x faster than GPT-6 Luna," with bars at 150 ms for the Decisions API and 1.6 s for the GPT-6 Luna API.

The comparison is against OpenAI's own cheapest model doing the same job the slow way. That is a fair internal measurement, but it says little about how the Decisions API compares with other vendors. OpenAI's product lead Tibo Sottiaux described it more loosely as "less than a few hundreds of milliseconds end to end."

For comparison, OpenRouter's production telemetry puts Jev's P50 at 0.21 s, P95 at 0.34 s, and P99 at 0.75 s as of October 1, 2026. OpenAI's 150 ms is a vendor claim under unstated conditions, while Jev's 210 ms is measured across real traffic, so I would treat them as the same ballpark until someone benchmarks both.

How does OpenAI's Decisions API compare with Jev?

Jev launched on September 15, 2026 from TypeSafe AI, founded by former OpenAI researcher Diogo Almeida. It is a "System One" model, and TypeSafe's pitch is that it never learned to write. It takes state in and returns typed decisions.

TypeSafe's Jev quick start docs showing a Noul question defined as JSON against a sample support-ticket state

Jev answers three question types, documented on OpenRouter and mixable in a single call:

  • Choice returns the selected option, a probability for every option, and a confidence value.
  • Noul returns the probability that a condition holds.
  • Score returns a probability-weighted position on an ordered rubric, plus the full distribution.

OpenRouter's Jev documentation page describing the TypeSafe decision model and its two API surfaces

DimensionOpenAI Decisions APITypeSafe Jev
AnnouncedSep 29, 2026Sep 15, 2026
AvailabilityLimited preview, selected customersGenerally available, no waitlist
EngineA constrained version of GPT-6 LunaPurpose-built System One model
Public API contractNone publishedFull schemas, Python and JS SDKs
Question types"Questions with finite pre-defined answers"Choice, Score, Noul, mixable per call
ProbabilitiesNot documentedPer-option, plus derived confidence
InputText or imagesText only: string, JSON object, or array
Context limitNot published32K on OpenRouter, 64K per TypeSafe's docs
Published latency150 ms vs 1.6 s (OpenAI slide)0.21 s P50 measured by OpenRouter
Input price per 1M tokens$0.10 for GPT-6 Luna$0.042
Output price per 1M tokens$0.50 for GPT-6 Luna$0
Can generate textUnderlying model canArchitecturally cannot

Three of these differences matter most.

Constrained LLM versus purpose-built model. OpenAI is narrowing a generative model's job. Jev is a different model class trained with what TypeSafe calls Reinforcement Learning for Calibrated Decisions. Jev generates no text, so it bills nothing for output tokens. On Hacker News, HarHarVeryFunny did the arithmetic: "They say it's built on Luna, which costs $0.10M/in, vs Jev which only costs $0.04M/in." Both figures check out against the vendors' own pages.

Image input. This is OpenAI's clear win, and it follows from the use case OpenAI is targeting. Altman's demo was a computer-use agent, and OpenAI's Romain Huet framed it on stage around robots taking "fast actions based on what they see in near real time." Jev's docs list text only. If your decision point is a screenshot, Jev cannot see it.

The name collision. OpenRouter already ships an endpoint called the Decisions API, at POST https://openrouter.ai/api/alpha/decisions, and it currently serves Jev. It predates OpenAI's product, and the two names will be easy to mix up in search results and docs.

How did the developer community react to the Decisions API?

The Decisions API's Hacker News submission drew six points, and commenters noticed the gap against the enthusiasm Jev had generated two weeks earlier.

Part of the reason is that developers had already seen the idea several times. In the fourteen days between the two launches, Hacker News front-paged six community reimplementations of Jev, including Jev in 25 Lines of Python at 691 points and Jeff, a set of Jev-compatible 0.8B models trained at home, at 566. OpenAI's version was the seventh, and the first without a public schema.

The @OpenAIDevs announcement tweet for the Decisions API, powered by GPT-6 Luna, available in limited preview

Most of the discussion was about what the launch meant for TypeSafe. One Hacker News commenter argued that Jev's backers had just learned there is no durable moat here, since OpenAI's version arrived two weeks to the day after Jev launched. The same commenter suspected Jev "was intended to be an acquisition play" that OpenAI passed on, a theory others had floated in the thread that predicted the launch. On Reddit, an r/AI_Agents thread titled "Jev Is Dead" opened with "With OpenAI's new Decision API, it honestly feels like Jev is going to have a very hard time surviving."

Other commenters read the same facts the opposite way. In their view, OpenAI building a decision endpoint at all proves the category Jev created is real, and that validation is worth more to a young company than two weeks of exclusivity.

Diogo Almeida, TypeSafe's founder, responded ninety minutes after the keynote, in a post that drew 647 likes:

Some of the frustration was aimed at OpenAI in general. The top replies under Sottiaux's post asked whether DevDay amounted to copying competitors and said there is no room left for smaller players. Both replies cleared 200 likes.

Press coverage focused on the timing. Simon Willison, live-blogging from the room, wrote: "Sounds like their response to Jev, which came out of stealth less than two weeks ago!" The New Stack's Frederic Lardinois said OpenAI "probably rushed the announcement ahead of its DevDay."

A smaller group questioned both products. In the Hacker News thread that predicted the launch a week early, hbrn pointed out that Jev's probabilities shift when you reorder the answer list, and that its confidence value is computed from those probabilities, so it adds no new information. That criticism applies to any decision model sold on calibration, including OpenAI's.

Where does a decision layer belong in an agent stack?

The decision layer sits in the same place whichever vendor you pick. Agent routing, tool gating, and passage filtering all use the same slot in a workable agent harness, which has five layers:

  1. Orchestrator holds the task, context, retries, and the loop.
  2. Generative model interprets intent, plans, and writes anything a human reads.
  3. Decision model answers narrow questions about route, risk, priority, or completion.
  4. Policy and permissions enforce what the user, agent, and tool may do.
  5. Action and audit execute the approved call and record what happened.

Layer three should never quietly become layer four. Keeping them apart makes the loop auditable, because the planner explains the goal, the decision layer answers one bounded question, and code executes only what policy permits. If you are designing that loop, loop engineering and context engineering cover the surrounding decisions.

For anything that touches the live web, a decision layer helps in four places:

  • Route scraped pages. Classify a crawled page as pricing, docs, changelog, or noise before it costs a frontier model any tokens. It is a cheaper first pass than full LLM extraction on every page.
  • Filter retrieval candidates. Score each passage for relevance and contradiction, then threshold in code rather than trusting cosine similarity.
  • Catch prompt injection in scraped content. Ask a narrow "does this text contain instructions aimed at the agent" question on every fetched page. A page can rank first by embedding and still be an attack.
  • Gate tool calls. Before a write, a payment, or a delete, ask whether the action is reversible and on-task, and combine that with an allowlist. This is the same surface that makes CLI-first agents risky to run unattended.

Every one of those depends on clean input. A decision model reading a page full of nav chrome, cookie banners, and ad markup burns its context on noise, and TypeSafe's own jaggedness page warns that accuracy falls as irrelevant state grows. Firecrawl turns a live page into clean markdown or typed JSON, which is the input a decision model needs. Our writeups on agentic search and training versus retrieved versus live web data go deeper on why the retrieval step decides the quality of everything downstream.

Should you build on the Decisions API today?

Reasons to wait. Access is restricted to selected API customers. Endpoint naming, response semantics, limits, and pricing can all change before broad release, and OpenAI told The New Stack it will share more "at broad rollout." You cannot forecast unit economics on a price that does not exist yet, and you cannot write a threshold policy against a confidence field nobody has documented.

Reasons to try it. If your decisions are visual, it is the only one of the two that can see. If you are already consolidated on OpenAI billing and safety review, a second vendor means more procurement and review work. And if the on-stage computer-use demo reflects real throughput, tight agent loops are the use case it was built for.

If you do integrate now, keep it behind a thin adapter so a preview contract change touches one file. Then roll it out in stages:

historical examples -> offline evaluation -> shadow traffic
  -> thresholded automation -> monitored production

Start in shadow mode whenever the decision touches money, access, safety, or reputation. Compare the model's pick against a human or a trusted rule for long enough to see the error distribution, then choose a threshold. Store the raw probabilities, the input version, and the eventual outcome, because the threshold you pick today is the one you will need to re-tune after the next model version.

Keep the decision model to fast judgments, and leave policy enforcement and conversation to other parts of the stack. By TypeSafe's own account, the teams getting value from Jev today run it alongside a frontier model.

Decision models are now a product category

Two weeks ago, "decision model" described one startup's product. Now OpenAI sells one and OpenRouter serves Jev through its own Decisions endpoint.

Whichever vendor ends up with the market, the way you build on them is the same. You write a bounded question and a fixed set of answers, let the model pick one, and keep permissions and policy in your own code. You also stop paying a generative model to write a paragraph when your code only needed one word.

The decision is only as good as the state you send with it. If your agent makes decisions about the live web, start with Firecrawl, which turns pages into clean markdown or JSON so the decision layer sees the content without the navigation, banners, and ads.

Frequently Asked Questions

What is OpenAI's Decisions API?

It is a limited-preview endpoint announced at DevDay 2026 that focuses GPT-6 Luna on developer-defined questions with a finite set of pre-defined answers. You supply context as text or images and get back a selection your code can branch on, for classifying content, routing requests, or choosing an agent's next action.

Is the Decisions API a new model?

No. OpenAI describes it as focusing Luna's intelligence on a bounded question. Sam Altman said on stage that it works by giving the Luna model a predefined set of options to choose from. It is a constrained interface over an existing generative model.

How is the Decisions API different from structured outputs or JSON mode?

Structured outputs constrain the format of generated text after the fact. A decisions endpoint bounds the answer space before inference and optimizes the round trip around selecting from it. In practice the difference shows up as latency, because the model is not generating and you are not parsing.

Has OpenAI published the Decisions API request schema?

Not as of September 30, 2026. There is no entry for it in OpenAI's API guides index or endpoint reference, no path in the public OpenAPI spec, and no method in the Python, Node, or Go SDKs. Any request payload you find online claiming to be the OpenAI Decisions contract is invented.

How much does the Decisions API cost?

OpenAI has not published Decisions-specific pricing. The underlying GPT-6 Luna model is listed at $0.10 per million input tokens and $0.50 per million output tokens. An OpenAI spokesperson told The New Stack the company will share more at broad rollout.

Is the Decisions API faster than Jev?

On published numbers they are close. OpenAI's DevDay slide claims 150 ms for the Decisions API against 1.6 s for a standard GPT-6 Luna call. OpenRouter's measured P50 latency for Jev is 0.21 s, with P95 at 0.34 s as of October 1, 2026. OpenAI's figure is a vendor claim under unknown conditions; OpenRouter's is production telemetry.

Can the Decisions API read images?

Yes, and this is its clearest advantage over Jev. Image input is confirmed in OpenAI's recap, in the @OpenAIDevs announcement, and twice on the DevDay keynote stage. Jev's documentation lists text only, as a string, JSON object, or array.

Does the Decisions API return confidence scores?

OpenAI has not documented the response shape. The New Stack reports that it returns predefined answers with confidence scores, and OpenAI has not published a contract that confirms it. Jev, by contrast, documents per-option probabilities plus a derived confidence value for Choice and Score questions.

Should I build on the Decisions API today?

Only behind an adapter, and only for experiments you can afford to rewrite. Access is limited to selected API customers, and endpoint naming, response semantics, limits, and pricing can all still change before broad release. For a decision layer you can ship this week, Jev is generally available through TypeSafe or OpenRouter with no waitlist.

Does a decision model replace an LLM?

No. It replaces the prompt-and-parse step where you were already asking an LLM a narrow question and extracting a label. Planning, reasoning, tool orchestration, and anything user-facing still belong to a generative model. The common pattern is a strong model that plans and a fast decision model that picks.