METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

OpenAI Opens Decisions API in Beta

OpenAI has opened its Decisions API, which uses GPT-6 Luna to return only probabilities, choices, and scores, to all developers in public beta. Input costs $0.10 per 1M tokens with no charge for output tokens, and the company said it expects general availability within weeks.

OpenAI Opens Decisions API in Beta

Image: @OpenAIDevs (X) (video still)

Summary

  • OpenAI released the Decisions API to all developers in public beta on October 6 (US time), saying it makes decisions up to 10x faster than calling GPT-6 Luna through the Responses API.
  • It supports three question types, predicates that estimate the probability a condition is true, choices that pick one option, and scores against a rubric, over text and image inputs.
  • Pricing is $0.10 per 1M input tokens with no output or cache charges, and Zero Data Retention and HIPAA use are supported for eligible customers.
Let your app choose the right model, tool, or action in near real-time with Decisions API

OpenAI released the Decisions API, which returns only typed answers to a question, to all developers in public beta on October 6 (US time). On its official developer account, the company said the API lets "your app choose the right model, tool, or action in near real-time," and that it makes decisions up to 10x faster than GPT-6 Luna through the Responses API. GPT-6 Luna is the only supported model, served through a dedicated POST /v1/decisions endpoint. In its developer documentation, OpenAI wrote that it expects general availability (GA) in the coming weeks.

The release comes one week after the API appeared in limited preview at DevDay 2026 on September 29. METAL previously reported that OpenAI unveiled the Decisions API in limited preview, when per-call pricing and data policies had not been disclosed. The public beta fills in those blanks. GPT-6 Luna is the reasoning model OpenAI released alongside GPT-6 Sol on September 22.

A request has three fields. The model field names the evaluating model, and the input field carries shared evidence for the questions as a text string or as user messages mixing text and images. The questions field lists each question's type, instructions, and any allowed choices or score levels. The response comes back as an answers array, and the unique name given to each question is echoed back so developers can tell which answer belongs to which question.

Decisions API 요청의 세 칸(model, input, questions)과 각 칸의 역할을 정리한 오픈AI 개발자 문서 표

There are three question types. A predicate estimates the probability, from 0 to 1, that a condition is true. A choice selects one of the options a developer supplies and returns a probability for each option along with a separate confidence value. A score rates an input against levels ordered from lowest to highest and returns the probability-weighted average of the level indices, so a score can fall between two levels.

The examples in the developer guide that METAL reviewed put numbers on this structure. A predicate question checking a product photo for a crack, tear, or dent returned a probability of 0.92. A choice question routing a customer complaint about being "charged twice for my order" picked billing with a probability of 0.95 and a confidence of 0.93. A score question rating an export that fails only in Safari against three severity levels produced a score of 1.1 with a confidence of 0.55 from a distribution of 0.1, 0.7, and 0.2. OpenAI described these values as illustrative response excerpts.

Inputs come with conditions. Images must be sent as inline base64 data URLs, and hosted HTTP or HTTPS image URLs and file_id inputs are not accepted. Independent questions can share one request's questions array to evaluate the same input in a single call. Questions that depend on an earlier answer must be sent as separate requests. The guide gives the example of checking for damage first and then using the result to decide whether to ask for a repair category.

The pricing structure is also out for the first time. Using the Decisions API with GPT-6 Luna costs $0.10 per 1M input tokens, and only input tokens are billed. Cache reads, cache writes, and output tokens carry no charge. Regional processing premiums and long-context input multipliers still apply, and the rate covers only /v1/decisions calls. On data controls, the API supports Zero Data Retention (ZDR) and HIPAA use for eligible customers, with data residency and regional processing available in the United States and Europe (EEA and Switzerland).

The development environment is in place as well. Running the guide's examples requires Python SDK 3.26.0, JavaScript 7.30.0, Go 3.73.0, Ruby 0.101.0, or Java 4.78.0 or later. Developers can test questions and inputs in the Playground before writing code. Voice apps can pair the API with client delegation in the Live API to choose actions from voice requests and report the results to the user.

The 71-second official demo video put speed up front. An OpenAI staff member in the video showed loosely written user inputs being sorted into sales leads, said "this is not sped up," and explained that the Decisions API responded in under 100 milliseconds on the server side. Further demos showed a game in which the API takes frames of the road and obstacles and steers a car into the correct lane, and an animated character whose voice is handled by GPT Live 1 while the Decisions API picks its facial expression. At the time of writing, the post had 577,000 views, 2,538 likes, and 1,322 bookmarks.

Seen through an AI engineer's lens, the API targets branching, not generation. The guide draws the line itself, pointing developers to Structured Outputs in the Responses API when they need an object that follows their own JSON schema or a written explanation, and to function calling when a model must request a tool call with arguments. What remains are short judgments such as classification, routing, and prioritization, the calls an agent makes when choosing its next action. A design that does not bill output tokens fits a use case whose answers end in a few probabilities.

Operations will hinge on thresholds. In the guide, OpenAI recommends using "labeled examples from your application to set thresholds for routing, filtering, or review," and choosing thresholds based on the cost of false positives and false negatives. It also advises including a fallback option such as "other" for inputs the categories do not cover. An API that returns probabilities does not make the call for you, and deciding where to cut those probabilities remains the developer's job. The public beta is the first chance to tune that line against real traffic.

Comments