The OpenAI Decisions API is a new endpoint that does not write answers. It returns probabilities. OpenAI's Decisions guide describes it as public beta, available at POST /v1/decisions, and for now it accepts a single model, gpt-6-luna. Reports date the public beta to October 6, 2026, after the feature was introduced at DevDay on September 29, according to Vercel's explainer. It follows the wave we covered in The Decision Models Wave: Perplexity, Cloudflare, AWS vs Jev, which did not include OpenAI. This piece focuses on what OpenAI shipped, what it costs, and where its claims remain unverified.
How the Decisions API works
A request carries a model, an input (text, images, or both) and a list of named questions. Each question is one of three types, according to the OpenAI guide:
- predicate: returns a probability from 0 to 1 that a condition is true.
- choice: returns one of the values you supplied, a probability for each, and a confidence figure.
- score: takes ordered levels and returns the probability-weighted average of the level indices, which can fall between levels.
The guide's own score example turns level probabilities of 0.1, 0.7 and 0.2 into a score of 1.1. Independent questions can share one request, while a decision that depends on an earlier answer needs a second request. Any answer can also come back as a refusal, so applications should check the answer type before reading numbers, as Vercel's write-up notes.
Why a probability is different from generated text
A chat model asked to classify a support ticket writes a word, and the word has no attached uncertainty. If you ask it to also write "confidence: 0.8", that number is more text, produced by the same process as the rest of the sentence. A decision endpoint instead returns a distribution over the options you gave it. The guide advises setting thresholds from labeled examples from your own application and weighing the cost of false positives against false negatives.
This matters because a probability supports a policy. A team can auto-approve above one threshold, send mid-range cases to a larger model, and route low-confidence cases to a person. That only works if the numbers are calibrated, meaning a 0.9 is right about nine times in ten. OpenAI's guide does not publish calibration results for Luna, so that property has to be measured on each workload. The score type adds one more nuance: because it is an average, a result of 1.1 can mean "mostly level 1" or "split between levels 0 and 2", and the separate confidence value and per-level probabilities are what distinguish them.
A probability supports a policy: auto-approve above one threshold, escalate in the middle, hand to a person below another.
Analysis
Pricing math: what 1,000 classifications cost
OpenAI lists Decisions at $0.10 per million input tokens, with no cache-read, cache-write or output-token charges; regional processing premiums and long-context multipliers still apply. For comparison, the gpt-6-luna listing prices ordinary generation at $0.10 input and $0.50 output per million tokens, and gpt-6-sol at $2 input and $10 output (see our GPT-6 Sol and Luna pricing post).
The figures below are an illustrative calculation, not a measurement. They assume 500 input tokens per classification, no caching, no regional premium, and output lengths that we chose.
Illustrative cost of 1,000 classifications
Assumes 500 input tokens per call, list prices only, no caching or regional premiums. Output lengths are assumptions. Not a measurement.
Per 1,000 calls: 500,000 input tokens. Sources: OpenAI Decisions guide, gpt-6-luna and gpt-6-sol model pages.
| Scenario (1,000 calls, 500 input tokens each) | Input cost | Output cost | Total |
|---|---|---|---|
| Decisions API, gpt-6-luna | $0.050 | $0 | $0.050 |
| Luna generation, 50 output tokens | $0.050 | $0.025 | $0.075 |
| Luna generation, 300 output tokens | $0.050 | $0.150 | $0.200 |
| Sol generation, 50 output tokens | $1.000 | $0.500 | $1.500 |
The key reading is that Decisions input costs the same per token as Luna generation input. The saving against Luna comes only from dropping output tokens, so it is modest for short labels (about a third in the 50-token case) and larger when a generation call would emit reasoning or long JSON. The bigger gap appears against a larger model, which is a model-choice effect as much as an endpoint effect. Prompt size dominates the bill either way, so a 5,000-token prompt costs ten times more than the 500-token case.

Where it fits: routing, moderation, triage and evals
OpenAI's guide lists classifying content, routing requests and prioritizing work, with examples such as checking product photos for visible damage, routing customer complaints to departments and rating issue severity. Those map onto four common jobs:
- Routing: a choice question picks which model or queue handles a request. Multi-model workspaces such as Metir face exactly this question on every request, which is why routing is one of the most natural fits for typed decision calls.
- Moderation: predicate questions such as "does this contain a threat" give a probability to threshold per policy.
- Triage: a score question ranks urgency; a choice question assigns a department.
- Evals: a score question can act as a grader with a rubric, though grader bias still needs checking.
The guide recommends a fallback value such as "other" when categories may not cover an input, and writing questions around observable criteria.
Limits and open questions
Several constraints are documented, and one claim is not independently checked.
- Single model, beta: only gpt-6-luna is supported, and OpenAI expects general availability "in the coming weeks" without a date. Details may change before then.
- Input restrictions: images must be inline base64 data URLs. Hosted image URLs and file IDs are not supported, and a request allows up to 128 image parts, per Mixed News.
- Speed: OpenAI says the endpoint is about 10x faster than the Responses API. Mixed News notes that OpenAI published no benchmark behind the claim, and Let's Data Science says no workload definition or latency distribution accompanied it. Treat 10x as a hypothesis and time your own traffic.
- Fixed answer sets: a correct outcome can be missing from your list, and a classification does not authorize any downstream action.
- No outside knowledge: the model sees only what you send, not your order records or policies.
Mixed News also reported that the gpt-6-luna endpoint-support table did not yet list /v1/decisions when it checked on October 6, a documentation gap rather than a functional one, since the guide itself documents the endpoint.
What to watch
The most useful signals will be independent calibration and latency tests, the GA date, and whether OpenAI adds models beyond Luna. The pricing also sets a reference point for the rivals covered in our earlier piece: TypeSafe's Jev and the open-weight entrants priced near four cents per million input tokens. OpenAI's $0.10 is higher per token, so the competition may turn on accuracy, calibration, compliance options (OpenAI cites Zero Data Retention and HIPAA support for eligible customers) and ecosystem rather than list price alone.
Sources:
- Decisions API guide | OpenAI
- GPT-6 Luna model page | OpenAI
- GPT-6 Sol model page | OpenAI
- OpenAI Decisions API beta, endpoint table | Mixed News
- OpenAI opens Decisions API public beta | Let's Data Science
- What is OpenAI Decisions API | Vercel
Image credits
Header image: the Pioneer Building in San Francisco's Mission District, OpenAI's longtime headquarters, photographed in 2019 by HaeB, via Wikimedia Commons, licensed under CC BY-SA 4.0. In-body: Sam Altman at TechCrunch Disrupt San Francisco 2019, by TechCrunch, via Wikimedia Commons, licensed under CC BY 2.0.
