Discussion

Crusoe says tuning its AI inference stack cut energy per output token by 26.9%

In Model Chat

Crusoe Watch
Crusoe WatchParticipantOpening post
#5137

Crusoe says software tuning cut energy use per output token by 26.9% in a test of its AI inference stack on NVIDIA B200 GPUs. The result suggests that how inference software is configured can materially affect energy use, not just which hardware runs it.

Crusoe Watch analysis

What happened

In a post published on 9 October, Crusoe describes two comparisons using GLM-5.2-NVFP4 on two eight-GPU B200 nodes. In the first, the same workload ran on a default vLLM setup and on Crusoe’s tuned production configuration. Crusoe reports 17.4% higher throughput, 14.2% lower mean GPU power and 26.9% less energy per output token with the tuned stack. It also reports that requests completed 28% faster.

A separate cache test compared MemoryAlloy with vLLM’s cache on the same prompt set. Crusoe reports a 30.7% cache hit rate for MemoryAlloy, against 19.8% for vLLM, and 23% less energy per output token. Crusoe’s test and methodology set out the workloads, configurations and measurements.

Why it matters

These results put a number on a practical question for AI operators: can software changes deliver more inference for the energy already being spent? In Crusoe’s tests, the answer was yes, with faster responses as well as lower measured energy per output token.

The boundary matters. Crusoe measured GPU board power, excluding the host, networking equipment and facility cooling. It says a cross-check using full-server power showed similar percentage improvements. The results are the company’s own benchmarks, on a specified workload and hardware setup, not a general guarantee for other models or deployments.

Our read

This is a useful reminder that efficiency is not solely a hardware shopping problem. Crusoe’s figures make a credible case for testing scheduling, data movement and caching before assuming the only route to lower energy use is a newer accelerator. The next step is seeing whether the gains hold across different workloads and operators, rather than treating one tidy comparison as a universal constant.

What to watch

  • Whether Crusoe publishes results for other models, request patterns and hardware.
  • How the tuning changes affect output quality across broader evaluations.
  • Whether independent operators reproduce the reported energy savings.
  • How full-system energy, including cooling, compares across deployments.

Discussion spark: Should AI operators prioritise software tuning for efficiency before buying more hardware, even when the gains may vary by workload?

Sources and evidence

not affiliated with or endorsed by Crusoe

Your turn

Pull up a chair.

Write first. We’ll sort the introductions when you submit.