The new endpoint lets software branch on an AI picked answer, but a test of the model underneath reports a 99% claimed, 68% actual confidence gap.
OpenAI shipped a new kind of endpoint at DevDay on Sept. 29: a decision primitive that picks one answer from a list you define so software can branch on it. The version in limited preview runs on a model the vendor calls GPT-6 Luna. A reproducible test of that model, published by Anthropic's Alignment team, reports a confidence gap builders will need to know about before they wire agent work to it.
The test ran 3,600 reasoning problems from Anthropic's Hard-Decisions benchmark against Luna, with reasoning disabled, alongside peer models Jev, Kev, and Laya. On a separate log-probability probe of 2,672 answers where the model assigned the answer token at least 99% probability, accuracy came in at about 68.7% on 1,302 open-world answers and 68.1% on 1,370 closed-world answers. The blog also reports that Luna's confidence AUROC, a measure of how well a confidence score discriminates right from wrong, fell to 0.51 at proof depth five, versus 0.84 for Jev. A score of 0.5 is roughly a coin flip.
Translation for the routing decision a builder has to make: when the model says it is 99% sure, it is right about 68% of the time. That is the gap the new endpoint is supposed to close for you, not open up.
Anthropic did not test the Decisions API itself. OpenAI has not published documentation that would let the API be tested directly, so the team ran the benchmark against the underlying model, not the endpoint. The 99%/68% number is a token-probability proxy: it reflects how the model would have ranked its own answer in a regular completion call, not whatever confidence surface the API may eventually expose. It is also a result on synthetic, templated logic problems, not on the kind of work most agents will route. The model used in the test was reasoning-disabled; reasoning-enabled runs may behave differently, and the specialized Decisions API may add its own filtering or reranking on top.
What is not in dispute is the second-order finding. The 0.51 AUROC at depth five means that, on the hardest problems, Luna's confidence score carries almost no discrimination between correct and incorrect answers. If a builder set a threshold on the model's reported confidence to decide whether to escalate, defer, or auto-accept, that threshold would do roughly nothing on the difficult cases, which are the cases the threshold exists for.
The decision endpoint is the layer where an LLM call becomes a routing decision: classify, route, choose the next step. A documented confidence surface, a number a builder can set a threshold on with stated probe method, error bars, and a clear contract for when it is and is not returned, is what makes that switch safe to wire into production. None of that is on the public page for the Decisions API as of the limited preview.
What good looks like is specific. A vendor shipping a decision endpoint should publish the calibration method used to derive the confidence score, the probe set and held-out split it was measured on, expected calibration error at a stated threshold, and the contract for when the endpoint returns a confidence score at all: for every request, on a sample, or only above some temperature floor. A builder reading that page should be able to point to a number and say: above this threshold, the endpoint is right N% of the time, with this error bar, on problems that look like mine.
The Hard-Decisions benchmark, run against the model underneath OpenAI's new endpoint, gives the field a working yardstick. It is not a verdict on the API, because the API was not tested, and reasoning-enabled runs may behave differently. The test a builder should run on any decision endpoint shipping a confidence claim in 2026 is concrete: show the calibration curve, the probe set, and the threshold contract. Without those, a 99% confidence number is not a routing signal.