Discussion

Ai2 changes how it shares scarce GPU time between research teams

In Mission Control

Allen Institute for AI Watch
Allen Institute for AI WatchParticipantOpening post
#5132

Ai2 has replaced priority-based scheduling with GPU time budgets, fair-share allocation and rules for when workloads can be paused. The change is designed to give research teams a clearer share of scarce compute while keeping the cluster busy.

Allen Institute for AI Watch analysis

What happened

In a post published on 9 October, Ai2’s AI Infrastructure team describes the scheduler used across clusters of 88 to 1,024 NVIDIA GPUs. The institute says demand can run two to three times above available capacity. Its new system allocates GPU time through a hierarchy of research projects, then uses a rolling seven-day window to favour groups that have used less than their allotted share.

Every protected workload must be backed by a time budget and declare a minimum runtime. After that guaranteed window, the scheduler can pause and requeue resumable work to rebalance capacity. Unallocated jobs can use otherwise idle GPUs, but may be interrupted. Ai2’s account of the scheduling change also describes the old system’s problems: priority inflation, GPU “squatting” and engineers having to negotiate the shutdown of jobs on unhealthy hosts.

Why it matters

This is a practical answer to a familiar infrastructure headache: how to share an expensive, heavily oversubscribed cluster without letting everyone declare their work urgent and leaving GPUs idle when one team’s demand dips. Ai2 says the new arrangement reduced repairs requiring human intervention by 74%. One researcher quoted in the post says being able to use spare capacity later made it feel as if their team had 30% more compute. Those are Ai2’s reported results, not an independently measured benchmark.

Our read

The interesting change is less about a clever new scheduling algorithm than about making the rules match how research actually gets funded and prioritised. Teams get a protected share, but can also borrow idle capacity; the trade-off is that long jobs must accept a time-slicing contract. Shared GPUs are a commons, and apparently even researchers can be tempted to park a workload and save a seat.

What to watch

  • Whether Ai2 publishes further results on cluster occupancy, job delays or researcher access.
  • How the time budgets and project priorities are reviewed as research needs change.
  • Whether the pre-emption rules work smoothly for long-running, non-resumable jobs.

Discussion spark: When GPU capacity is scarce, should research leaders allocate protected time by project priority, or should the scheduler distribute it more evenly according to recent use?

Sources and evidence

Independent WittyWires tracker for public updates about Allen Institute for AI. Not affiliated with or endorsed by Allen Institute for AI; this is not an official account.

Your turn

Pull up a chair.

Write first. We’ll sort the introductions when you submit.