Skip to content

Article

Luna Latency in Practice: What Sub-150ms Decisions Mean for Your Pipeline

Luna is OpenAI's smallest, most cost-effective model, purpose-tuned for classification. Sub-150ms decisions change what you can put in the hot path — here is how to measure and budget for it.

The latency envelope

Luna is built for one job: fast, constrained choice selection. Benchmarks from the DevDay 2026 materials put typical decisions around 150ms, with sample runs landing near 142ms. The important number is not the single sample — it is that the decision fits inside a user-facing interaction loop.

What sub-150ms unlocks

When a decision costs about 150ms, it can sit in front of synchronous requests:

  • Customer triage that feels instant to the caller
  • Moderation checks that run inline, not in a deferred queue
  • Agent loops that route the next action without burning reasoning tokens

Measuring it yourself

The response includes a latency_ms field, so you can observe the actual envelope in production rather than trusting a vendor number. Track P50 and P99, and keep context payloads lean — the size of the input is part of the latency budget.

Back to all guides