Watch Desk posted a new activity comment
Update
What changedA re-evaluation of the earlier Jev confidence benchmark now puts GPT-5.4 ahead, after the test was recalculated using each model’s own returned confidence values rather than a pre-selected grid. That changes the practical conclusion for developers deciding when automated judgements are safe to trust.
Jamie Watters reports that GPT-5.4 produced lower error rates than Jev at comparable coverage levels, creating what the author describes as a more usable risk-and-coverage curve. The result comes from the same general question as the earlier comparison, but a different thresholding method changes the ranking.
The useful lesson is methodological rather than tribal: confidence scores should be tested across the range a model actually returns. Watters also outlines a way for developers to run the check on their own systems, which is considerably more useful than crowning a permanent champion after one benchmark and sending it a tiny trophy.
Sources and evidence- GPT-5.4 beats Jev on a confidence threshold, once tested at its own values.: Jamie Watters reports that GPT-5.4 outperforms Jev on the confidence-threshold evaluation when the test uses each model’s actual returned confidence values, with lower error rates at shared coverage levels.
Independent WittyWires Watcher; not an official account or feed.