Luna Model
OpenAI's smallest, most cost-effective model designed specifically for high-speed constrained classification.
OpenAI Decisions API delivers deterministic, zero-hallucination choice selection in ~150ms. Purpose-built for autonomous agent routing, real-time moderation, and policy evaluation with multimodal text and vision support.
OpenAI DocumentationDevDay 2026 · Luna Engine · ~150ms Latency · Multimodal Text & Vision
Traditional LLMs generate unconstrained text or complex JSON. The Decisions API is an entirely new primitive: given context and candidate choices, it returns exactly one choice with zero hallucination.
OpenAI's smallest, most cost-effective model designed specifically for high-speed constrained classification.
Benchmarked at ~150ms response times, ideal for latency-sensitive customer triage and multi-agent routing loops.
Natively evaluates both text context and visual image payloads against candidate decisions.
Mathematical guarantee that the response is strictly one of your declared choices without regex or parsing overhead.
Inspect request structures with multimodal context, candidate choices, and ~150ms deterministic output.
import OpenAI from 'openai';
const client = new OpenAI();
// Fast, constrained decision making with Luna
const decision = await client.decisions.create({
model: 'luna',
context: [
{
role: 'user',
content: 'Customer asks for an immediate refund on duplicate charge #8921.',
},
],
choices: ['escalate_to_human', 'auto_refund', 'request_more_info'],
});
console.log(decision.choice); // "auto_refund" (guaranteed invariant)
console.log(decision.latency_ms); // ~142msHow high-throughput teams integrate OpenAI Decisions API across agent pipelines, edge workers, and mission-critical systems.
Select tools and next-best actions in autonomous agent loops without expensive reasoning token overhead.
Classify incoming tickets, customer inquiries, and moderation events in ~150ms at scale.
Pass photos or screenshots with candidate status labels for instant automated inspection and categorization.
Deploy policy enforcement directly in Cloudflare Workers to block jailbreaks and route requests at the edge.
Understand when to use the new Decisions API over Structured Outputs, Chat Completions, or dedicated classifier models.
Structured Outputs enforce JSON schemas; Decisions API is 5x faster and selects 1 discrete choice at a fraction of the token cost.
Chat models generate conversational prose and risk hallucination; Decisions API guarantees an exact choice with 0% token waste.
Zero training data or dedicated GPU server maintenance needed; update decision sets dynamically in production.
Function calling requires schema parsing and argument extraction; Decisions API routes branching in a single rapid hop.
github.com
Constrained Decisions at the speed of inference
Announced at DevDay 2026, the Decisions API is an API primitive engineered for ultra-fast, constrained decision-making tasks rather than open-ended prose generation. Given context and predefined choices, it returns exactly one choice.
Eliminate latency bottlenecks and hallucinated tool calls with OpenAI's fastest decision primitive.
OpenAI Documentation