A test of GPT-6 Luna Decisions found that its confidence scores could not distinguish correct answers from wrong ones on a SHA-256 checking task, while a separate causal-reasoning test found a much more accurate, though narrow, high-confidence subset. The useful lesson for anyone automating decisions is that confidence must be tested against the task, not trusted as a general-purpose safety switch.
Watch Desk analysis
What happened
In a community article on Hugging Face, Stephen Solka describes two 600-case evaluations of OpenRouter’s Decisions API using GPT-6 Luna Decisions. For SHA-256 verification, the model answered “does not match” every time, giving it 50% accuracy across a balanced set of matching and incorrect digests. Its average confidence was 99.8%. Of the 504 predictions at 100% confidence, just 49.6% were correct.
The article’s separate slice of counterfactual causal-reasoning questions produced 413 correct answers among 594 returned responses, or 69.5%. At a 99% confidence threshold, 47 of 49 accepted answers were correct. That is 95.9% accuracy among accepted answers, but they covered only 8.2% of the 600 cases. Read Solka’s evaluation and methods.
Key findings
- Confidence did not rescue the checksum task
A 99% threshold accepted 573 of 600 predictions, with 49.7% accuracy among them. - The causal-reasoning filter found a stronger subset
It accepted 49 cases, got 47 right and left most of the workload unanswered. - A score is not a guarantee
Solka argues that evaluations should measure accuracy against coverage, including the cost of errors and refusals.
Why it matters
A confidence threshold is appealing because it looks like a dial: turn it up, and let the safer answers through. These results show why that assumption needs testing. On the checksum task, high confidence barely filtered anything and did not improve accuracy. On the causal-reasoning slice, the threshold selected a much more accurate group, but only by covering a small fraction of the cases.
The contrast is about these tests, not proof that confidence scores are always useful or useless. The practical question is whether a model’s scores rank answers reliably on the particular task, and whether the remaining error rate and coverage make sense for the action at stake.
Our read
This is a good argument for measuring what a confidence cutoff actually buys before giving it authority over real decisions. If the check is deterministic, use deterministic code; if a model is doing judgement work, test the error-versus-coverage trade-off on fresh examples. A very certain answer is still just an answer with excellent posture.
What to watch
- Whether later evaluations reproduce the results on fresh cases and broader datasets.
- How accuracy changes as coverage rises or falls across different confidence thresholds.
- Whether tests report refusals and the cost of mistakes, not just accuracy among answered questions.
Discussion spark: Would you trust a model’s high-confidence answers for a narrow task if testing showed they were accurate only on a small share of cases, or should it handle the task only when it can cover most of the workload?
Sources and evidence
- Understanding Jev: confidence is only useful if it works (8 October 2026, 22:47 UTC)
Watch Desk is operated by WittyWires as an independent cross-cutting AI news tracker. It does not speak for the organisations or people it covers.