OpenAI Decisions API thumbnail with bold title and probability bars on a blue gradient

OpenAI Decisions API Explained: Typed Answers at $0.10 per Million Tokens, With Caveats

The OpenAI Decisions API is a new endpoint that skips the chat and hands your app a typed answer: a probability, a pick from your list, or a score. It is in public beta, works only with GPT-6 Luna, and costs $0.10 per million input tokens, according to StackOne’s guide.

If your app sorts support tickets, flags risky messages, or routes requests, this is built for that job. If you need written explanations or tool arguments, it is not.

Here are the facts you can scan in ten seconds:

  • Status: public beta, with general availability expected “in the coming weeks” per StackOne
  • Endpoint: POST /v1/decisions
  • Model: gpt-6-luna only, for now
  • Price: $0.10 per million input tokens, with output not billed
  • Best for: classifying, routing, gating and scoring

Because the beta is new, limits and pricing can still change before launch.

What the OpenAI Decisions API actually does

A normal chat call returns text, and you then parse that text to find out what the model decided. The Decisions API removes that step. You send an input and a list of named questions, and you get back one typed answer per question.

Each question needs a unique name, and the response returns an answers array in the same order. That makes the result easy to plug straight into an if-statement or a queue.

OpenAI showed the idea at DevDay, which we covered in our DevDay 2026 preview. It later opened the beta to all developers, as AI Weekly reported.

The three question types and when to use each

The API supports three kinds of questions. Pick the one that matches the shape of the answer you need.

Question type What you get back Good for
Predicate A probability from 0 to 1 that a condition is true Does this passage answer the question? Is this message spam?
Choice One option from your list, a probability for each option, and a confidence value Routing a ticket to billing, tech support or sales
Score A probability-weighted position on ordered levels, so results can fall between levels Rating ticket severity or lead quality

Per Vercel’s explainer, answers can also come back as a refusal. Check the answer type before you read any numbers, and send refusals to a review path.

Price, speed and limits you should know

The pricing is unusual because you only pay for what goes in. StackOne reports these details:

  • $0.10 per million input tokens
  • Output, cache reads and cache writes are not billed
  • Regional processing and long-context requests carry price multipliers
  • Zero data retention is available to qualifying customers

That flat cost per call is handy, since the price does not change with the answer. OpenAI also says it responds about 10x faster than the Responses API, though you should time it on your own workload.

It lands in the same cheap-input range as Anthropic’s latest small model, which we broke down in our Claude Haiku 5.5 pricing guide. Low-cost classification is clearly where the market is heading.

A few hard limits apply. Images must be inline base64 data, not hosted URLs or file IDs. Native decision calls also ignore the reasoning options from the Responses API. Vercel notes its gateway version accepts text only.

Decisions API vs Structured Outputs vs function calling

These three tools look similar but solve different problems. This table shows where each one fits.

Tool Returns Use it when
Decisions API Typed answers with probabilities You need to classify, gate, route or score
Structured Outputs Text that follows your JSON schema You need to extract fields or generate written explanations
Function calling A tool call with generated arguments You want the model to trigger an action

The Decisions API only chooses from answers you define. It never writes tool arguments, so you still need your own code to act on the result.

Early tests say go slowly

The first outside results are mixed, and all of them are single early-beta tests. StackOne cites a HiringCafe benchmark where OpenAI cost about 2x more than TypeSafe AI’s Jev and was 5 to 10% less accurate at relevance scoring. A forum tester also reported more confident wrong answers on nuanced calls, while narrow yes/no checks came out about even.

Keep in mind that StackOne sells competing agent products, so read those numbers as a prompt to test, not a verdict. Two more cautions come from the guides: the order of options in a choice question can change the answer, and a fixed list can leave out the correct outcome.

The safe approach is to test on a reviewed sample, set your own probability thresholds, and keep a human review path for low-confidence results. Also keep your own authorization checks on anything an answer triggers.

Frequently Asked Questions

Is the OpenAI Decisions API free?

No. It is billed at $0.10 per million input tokens, according to StackOne, but output is not charged. Regional and long-context requests cost more.

Which model does the Decisions API use?

Only GPT-6 Luna (gpt-6-luna) is supported during the beta. Vercel’s AI Gateway also lists other decision models, including Jev, Laya and Liquid d1.

Can the Decisions API read images?

Yes, when you call OpenAI directly, but images must be inline base64 data URLs. Hosted URLs and file IDs are not supported, and Vercel’s gateway accepts text only.

Is it ready for production?

It is still a beta, and OpenAI expects general availability soon. Use it for low-risk routing first and confirm your organization qualifies before sending personal or health data.

Our take: worth a pilot, not a rewrite

If you already use a chat model to sort or score things and then parse the reply, try the Decisions API on one workflow this week. The typed answers and input-only pricing make it a clean fit. But the early accuracy reports are thin and mixed, so benchmark it against your current setup before you move anything important.


Comments

Leave a Reply

Your email address will not be published. Required fields are marked *