AI data centres may save meaningful energy by changing how they run software, not only by installing newer chips. A University of Michigan researcher told Tom’s Hardware that tests and training systems point to gains worth taking seriously, though software is no substitute for more efficient hardware or new power capacity.
Watch Desk analysis
What happened
Tom’s Hardware reports that ML.Energy tests of Alibaba’s Qwen 3 235B A22B Thinking model found inference in FP8 used a third less energy than bfloat16 on problem-solving tasks. The article also describes Perseus, a training optimiser from researcher Jae-Won Chung, which reportedly cut training energy by up to 30% without reducing throughput or changing hardware.
Other approaches in the report include matching cloud instances to workloads, caching frequently used prompts, sending routine requests to smaller models, and improving batching and compiler efficiency. Scheduling non-urgent jobs for times or places with spare, lower-carbon electricity could also help, though data sovereignty rules and the difficulty of moving large datasets can limit that option. Read the report in Tom’s Hardware.
Key findings
- Lower-precision inference
ML.Energy’s test found FP8 used a third less energy than bfloat16 on the model and tasks tested. - Smarter training schedules
Perseus reportedly cut training energy by up to 30% while maintaining throughput on the reported setup. - Better-matched workloads
Smaller models, cached prompts and better batching can reduce wasted computation, according to the report. - Shift when and where work runs
Flexible batch jobs could move to periods or regions with more available, lower-carbon electricity, where rules and data movement allow. - Efficiency can invite more demand
If saved power is spent generating more tokens, total electricity use may not fall. Efficiency is not a force field against growth.
Why it matters
Power is becoming a constraint on AI infrastructure, and new chips and electricity supply take time to build. Software changes could offer operators another way to get more useful work from equipment they already have. That matters for costs and grid pressure, as well as for how quickly providers can expand capacity.
The numbers here come from specific tests and research described in Tom’s Hardware’s interviews; they are not a guarantee that every model or data centre will see the same gains. And efficiency alone cannot settle the question of how much extra demand AI will create.
Our read
The useful shift is from asking only how powerful a chip is to asking how much useful work each watt buys. Operators should measure their own workloads before treating any one benchmark as a deployment promise. The least glamorous levers, including caching and retiring legacy kit, may be the first ones worth pulling.
What to watch
- Whether the reported energy savings hold across more models and production workloads.
- How operators balance flexible scheduling against latency, data-location rules and network costs.
- Whether efficiency reduces electricity use, or mainly gives AI systems room to do more.
Discussion spark: If more efficient AI makes each task cheaper but encourages far more usage, should operators be judged on energy per task or on total electricity consumed?
Sources and evidence
- Software could be the easiest fix for hyperscalers' AI power squeeze, researchers say (8 October 2026, 14:00 UTC)
Watch Desk is operated by WittyWires as an independent cross-cutting AI news tracker. It does not speak for the organisations or people it covers.