Skip to content
ANNOUNCED AT DEVDAY 2026 · POWERED BY LUNA

Constrained Decisionsat the speed of inference

OpenAI Decisions API delivers deterministic, zero-hallucination choice selection in ~150ms. Purpose-built for autonomous agent routing, real-time moderation, and policy evaluation with multimodal text and vision support.

OpenAI Documentation

DevDay 2026 · Luna Engine · ~150ms Latency · Multimodal Text & Vision

Live Decisions Stream~150ms Latency
1Multimodal context payload received
2Predefined choices strictly bounded
3Luna ultra-fast inference (142ms)
4Guaranteed exact choice returned
POST /v1/decisions→ choice: "auto_refund"

Engineered for deterministic agent execution

Traditional LLMs generate unconstrained text or complex JSON. The Decisions API is an entirely new primitive: given context and candidate choices, it returns exactly one choice with zero hallucination.

Luna Model

OpenAI's smallest, most cost-effective model designed specifically for high-speed constrained classification.

Sub-150ms Latency

Benchmarked at ~150ms response times, ideal for latency-sensitive customer triage and multi-agent routing loops.

Multimodal Input

Natively evaluates both text context and visual image payloads against candidate decisions.

Guaranteed Invariant

Mathematical guarantee that the response is strictly one of your declared choices without regex or parsing overhead.

SDK & cURL

Developer Integration Quickstart

Inspect request structures with multimodal context, candidate choices, and ~150ms deterministic output.

import OpenAI from 'openai';

const client = new OpenAI();

// Fast, constrained decision making with Luna
const decision = await client.decisions.create({
  model: 'luna',
  context: [
    {
      role: 'user',
      content: 'Customer asks for an immediate refund on duplicate charge #8921.',
    },
  ],
  choices: ['escalate_to_human', 'auto_refund', 'request_more_info'],
});

console.log(decision.choice); // "auto_refund" (guaranteed invariant)
console.log(decision.latency_ms); // ~142ms

Four production architectural patterns

How high-throughput teams integrate OpenAI Decisions API across agent pipelines, edge workers, and mission-critical systems.

1. Agent Next-Action Routing

Select tools and next-best actions in autonomous agent loops without expensive reasoning token overhead.

2. Real-Time Content Triage

Classify incoming tickets, customer inquiries, and moderation events in ~150ms at scale.

3. Multimodal Visual Verification

Pass photos or screenshots with candidate status labels for instant automated inspection and categorization.

4. Edge Guardrail Enforcement

Deploy policy enforcement directly in Cloudflare Workers to block jailbreaks and route requests at the edge.

Decisions API vs Alternatives

Understand when to use the new Decisions API over Structured Outputs, Chat Completions, or dedicated classifier models.

  1. 1

    vs Structured Outputs

    Structured Outputs enforce JSON schemas; Decisions API is 5x faster and selects 1 discrete choice at a fraction of the token cost.

  2. 2

    vs Chat Completions

    Chat models generate conversational prose and risk hallucination; Decisions API guarantees an exact choice with 0% token waste.

  3. 3

    vs Custom Classifiers (BERT)

    Zero training data or dedicated GPU server maintenance needed; update decision sets dynamically in production.

  4. 4

    vs Function Calling

    Function calling requires schema parsing and argument extraction; Decisions API routes branching in a single rapid hop.

Official Reference

github.com

OpenAI Platform / Decisions API

Constrained Decisions at the speed of inference

OpenAI Documentation

Frequently Asked Questions

Announced at DevDay 2026, the Decisions API is an API primitive engineered for ultra-fast, constrained decision-making tasks rather than open-ended prose generation. Given context and predefined choices, it returns exactly one choice.

Deploy deterministic decisions to your agents today

Eliminate latency bottlenecks and hallucinated tool calls with OpenAI's fastest decision primitive.

OpenAI Documentation