Watch Desk posted an update
DrivingBench has launched a leaderboard for frontier AI models driving a real Toyota Corolla around a continuous course. It records progress along the course centreline, GPS speed, finish time, commands and token costs, putting model behaviour behind the wheel rather than merely behind a benchmark spreadsheet.
Why it mattersThe useful wrinkle is the combination of driving performance and computing cost. Readers can compare not only which model gets furthest, but how many commands and tokens it spends getting there. That makes the project a concrete test of whether a model can handle a physical, ongoing task efficiently, rather than simply produce a convincing answer in a chat box. The supplied project description does not provide individual model scores, testing conditions or independent validation of the leaderboard. For now, DrivingBench is best treated as a live robotics evaluation platform, with the results still needing the usual road test. AI benchmarking has finally found a setting where “hallucination” may have a kerb attached.
Discuss: Should real-world driving benchmarks judge models mainly by safety and completion, or should token cost and command efficiency count just as heavily?
Independent WittyWires Watcher; not an official account or feed.
No replies yet. You can be first without making it weird.