metir
metir
Docs
Download on App StoreGet it on Google PlayLog inSign up
Back to Blog
OpenAI
Decisions API
GPT-6 Luna
Model Routing
AI Pricing

OpenAI Decisions API: GPT-6 Luna Beta, Pricing and Limits

OpenAI's Decisions API is in public beta on gpt-6-luna at $0.10 per 1M input tokens with no output charge. How it works, the cost math and its limits.

Metir AI TeamOctober 9, 20267 min read
OpenAI Decisions API: GPT-6 Luna Beta, Pricing and Limits

The OpenAI Decisions API is a new endpoint that does not write answers. It returns probabilities. OpenAI's Decisions guide describes it as public beta, available at POST /v1/decisions, and for now it accepts a single model, gpt-6-luna. Reports date the public beta to October 6, 2026, after the feature was introduced at DevDay on September 29, according to Vercel's explainer. It follows the wave we covered in The Decision Models Wave: Perplexity, Cloudflare, AWS vs Jev, which did not include OpenAI. This piece focuses on what OpenAI shipped, what it costs, and where its claims remain unverified.

$0.10Input price per 1M tokensno output, cache-read or cache-write charge
3Question typespredicate, choice, score
10xOpenAI's speed claim vs Responses APIunverified
1Supported modelgpt-6-luna
OpenAI logoOpenAI
OpenAI is the latest large lab to ship a typed decision endpoint.

How the Decisions API works

A request carries a model, an input (text, images, or both) and a list of named questions. Each question is one of three types, according to the OpenAI guide:

  • predicate: returns a probability from 0 to 1 that a condition is true.
  • choice: returns one of the values you supplied, a probability for each, and a confidence figure.
  • score: takes ordered levels and returns the probability-weighted average of the level indices, which can fall between levels.

The guide's own score example turns level probabilities of 0.1, 0.7 and 0.2 into a score of 1.1. Independent questions can share one request, while a decision that depends on an earlier answer needs a second request. Any answer can also come back as a refusal, so applications should check the answer type before reading numbers, as Vercel's write-up notes.

Why a probability is different from generated text

A chat model asked to classify a support ticket writes a word, and the word has no attached uncertainty. If you ask it to also write "confidence: 0.8", that number is more text, produced by the same process as the rest of the sentence. A decision endpoint instead returns a distribution over the options you gave it. The guide advises setting thresholds from labeled examples from your own application and weighing the cost of false positives against false negatives.

This matters because a probability supports a policy. A team can auto-approve above one threshold, send mid-range cases to a larger model, and route low-confidence cases to a person. That only works if the numbers are calibrated, meaning a 0.9 is right about nine times in ten. OpenAI's guide does not publish calibration results for Luna, so that property has to be measured on each workload. The score type adds one more nuance: because it is an average, a result of 1.1 can mean "mostly level 1" or "split between levels 0 and 2", and the separate confidence value and per-level probabilities are what distinguish them.

“

A probability supports a policy: auto-approve above one threshold, escalate in the middle, hand to a person below another.

Analysis

Pricing math: what 1,000 classifications cost

OpenAI lists Decisions at $0.10 per million input tokens, with no cache-read, cache-write or output-token charges; regional processing premiums and long-context multipliers still apply. For comparison, the gpt-6-luna listing prices ordinary generation at $0.10 input and $0.50 output per million tokens, and gpt-6-sol at $2 input and $10 output (see our GPT-6 Sol and Luna pricing post).

The figures below are an illustrative calculation, not a measurement. They assume 500 input tokens per classification, no caching, no regional premium, and output lengths that we chose.

Illustrative cost of 1,000 classifications

Assumes 500 input tokens per call, list prices only, no caching or regional premiums. Output lengths are assumptions. Not a measurement.

Per 1,000 calls: 500,000 input tokens. Sources: OpenAI Decisions guide, gpt-6-luna and gpt-6-sol model pages.

Scenario (1,000 calls, 500 input tokens each)Input costOutput costTotal
Decisions API, gpt-6-luna$0.050$0$0.050
Luna generation, 50 output tokens$0.050$0.025$0.075
Luna generation, 300 output tokens$0.050$0.150$0.200
Sol generation, 50 output tokens$1.000$0.500$1.500

The key reading is that Decisions input costs the same per token as Luna generation input. The saving against Luna comes only from dropping output tokens, so it is modest for short labels (about a third in the 50-token case) and larger when a generation call would emit reasoning or long JSON. The bigger gap appears against a larger model, which is a model-choice effect as much as an endpoint effect. Prompt size dominates the bill either way, so a 5,000-token prompt costs ten times more than the 500-token case.

Portrait of Sam Altman, OpenAI's chief executive, speaking on stage with a headset microphone
Sam Altman, OpenAI's chief executive, at TechCrunch Disrupt San Francisco in 2019. The photo is a file portrait and does not depict the Decisions API launch. Photo: TechCrunch, via Wikimedia Commons, CC BY 2.0.

Where it fits: routing, moderation, triage and evals

OpenAI's guide lists classifying content, routing requests and prioritizing work, with examples such as checking product photos for visible damage, routing customer complaints to departments and rating issue severity. Those map onto four common jobs:

  • Routing: a choice question picks which model or queue handles a request. Multi-model workspaces such as Metir face exactly this question on every request, which is why routing is one of the most natural fits for typed decision calls.
  • Moderation: predicate questions such as "does this contain a threat" give a probability to threshold per policy.
  • Triage: a score question ranks urgency; a choice question assigns a department.
  • Evals: a score question can act as a grader with a rubric, though grader bias still needs checking.

The guide recommends a fallback value such as "other" when categories may not cover an input, and writing questions around observable criteria.

Limits and open questions

Several constraints are documented, and one claim is not independently checked.

  • Single model, beta: only gpt-6-luna is supported, and OpenAI expects general availability "in the coming weeks" without a date. Details may change before then.
  • Input restrictions: images must be inline base64 data URLs. Hosted image URLs and file IDs are not supported, and a request allows up to 128 image parts, per Mixed News.
  • Speed: OpenAI says the endpoint is about 10x faster than the Responses API. Mixed News notes that OpenAI published no benchmark behind the claim, and Let's Data Science says no workload definition or latency distribution accompanied it. Treat 10x as a hypothesis and time your own traffic.
  • Fixed answer sets: a correct outcome can be missing from your list, and a classification does not authorize any downstream action.
  • No outside knowledge: the model sees only what you send, not your order records or policies.

Mixed News also reported that the gpt-6-luna endpoint-support table did not yet list /v1/decisions when it checked on October 6, a documentation gap rather than a functional one, since the guide itself documents the endpoint.

What to watch

The most useful signals will be independent calibration and latency tests, the GA date, and whether OpenAI adds models beyond Luna. The pricing also sets a reference point for the rivals covered in our earlier piece: TypeSafe's Jev and the open-weight entrants priced near four cents per million input tokens. OpenAI's $0.10 is higher per token, so the competition may turn on accuracy, calibration, compliance options (OpenAI cites Zero Data Retention and HIPAA support for eligible customers) and ecosystem rather than list price alone.

Compare models side by side on Metir

Try GPT-6 Luna and other models in one workspace and see how each handles your own classification prompts.

Open Metir chat

Sources:

  • Decisions API guide | OpenAI
  • GPT-6 Luna model page | OpenAI
  • GPT-6 Sol model page | OpenAI
  • OpenAI Decisions API beta, endpoint table | Mixed News
  • OpenAI opens Decisions API public beta | Let's Data Science
  • What is OpenAI Decisions API | Vercel

Image credits

Header image: the Pioneer Building in San Francisco's Mission District, OpenAI's longtime headquarters, photographed in 2019 by HaeB, via Wikimedia Commons, licensed under CC BY-SA 4.0. In-body: Sam Altman at TechCrunch Disrupt San Francisco 2019, by TechCrunch, via Wikimedia Commons, licensed under CC BY 2.0.

Ready to experience AI that adapts to you?

metir brings together the world's best AI models in one seamless experience. Start for free today.

Get Started Free
metir

Agentic Operating System for Professionals buried in meetings, emails and docs.

© 2026 metir. All rights reserved.

Product

  • Features
  • Pricing
  • Research
  • Docs
  • Blog
  • Enterprise

Company

  • Docs
  • Support
  • Careers

Legal

  • Terms of service
  • Privacy policy

Personalisation is powerful. Privacy is non-negotiable.

Status: All systems operational