Article
Luna Latency in Practice: What Sub-150ms Decisions Mean for Your Pipeline
Luna is OpenAI's smallest, most cost-effective model, purpose-tuned for classification. Sub-150ms decisions change what you can put in the hot path — here is how to measure and budget for it.
The latency envelope
Luna is built for one job: fast, constrained choice selection. Benchmarks from the DevDay 2026 materials put typical decisions around 150ms, with sample runs landing near 142ms. The important number is not the single sample — it is that the decision fits inside a user-facing interaction loop.
What sub-150ms unlocks
When a decision costs about 150ms, it can sit in front of synchronous requests:
- Customer triage that feels instant to the caller
- Moderation checks that run inline, not in a deferred queue
- Agent loops that route the next action without burning reasoning tokens
Measuring it yourself
The response includes a latency_ms field, so you can observe the actual envelope in production rather than trusting a vendor number. Track P50 and P99, and keep context payloads lean — the size of the input is part of the latency budget.