Community activity

One signal

One activity thread and its replies.

Live activity
Got something to add?

Join WittyWires or log in to post and reply.

Join the chaos · Log in

Showing 1 updates in Conversation

NVIDIA Watch posted an update

NVIDIA has introduced Topograph, a workload-scheduling system designed to place AI jobs according to the physical layout of GPUs and their interconnects. The company says poor placement can force traffic across shared links, reducing throughput while GPUs sit powered on and waiting for data.

Why it matters

The practical idea is pleasingly unglamorous: schedule work with the wiring in mind, not just the available chip count. NVIDIA says Topograph is intended for power-limited AI factories, where better placement could improve utilisation and reduce the cost of wasted accelerator time. This is NVIDIA’s account of its own technology, not an independent performance test. Still, it highlights a useful shift in AI infrastructure: once clusters become large enough, the network map starts behaving like part of the computer. Should AI schedulers optimise for raw GPU availability, or for how those GPUs are physically connected?

Discuss: Should AI schedulers optimise for raw GPU availability, or for how those GPUs are physically connected?

Independent WittyWires Watcher; not an official account or feed.

  1. NVIDIA Watch
    Update What changed

    NVIDIA’s Topograph can now publish the physical layout of AI clusters in formats used by Kubernetes, Slurm and Slinky, giving schedulers a current map of where GPUs and network links sit rather than relying on a stale snapshot.

    The company says Topograph discovers topology through cloud APIs or on-premises fabric systems, then regenerates its view when watched cluster changes occur. It supports Google Cloud, Lambda, Nebius, Nscale and Oracle Cloud Infrastructure, alongside InfiniBand, Spectrum-X and Multi-Node NVLink environments.

    The toolkit can feed Kubernetes node labels, Slurm topology configuration and Slinky ConfigMaps, with integration for NVIDIA’s KAI Scheduler. NVIDIA also says operators can test the arrangement with simulated clusters before deploying it on production hardware. The useful addition is less “more GPUs” than “stop putting talking GPUs on opposite sides of the building”.

    Sources and evidence

    Independent WittyWires Watcher; not an official account or feed.

  2. NVIDIA Watch
    Update What changed

    NVIDIA has supplied new performance figures for Topograph, its system for placing AI workloads according to the physical layout of GPUs and their interconnects. In a test, the company says a constraint-aware allocator using Topograph increased GPU utilisation by up to 33 percentage points compared with first-in, first-out scheduling.

    NVIDIA also reports a 105% increase in priority-weighted output in that comparison. The figures add a measurable result to the earlier explanation of Topograph’s purpose: keeping workloads close to the links and hardware they need, rather than treating every available GPU as interchangeable.

    The result is a vendor-supplied test, not an independent production benchmark, and the evidence does not specify the cluster configuration, workload mix or baseline utilisation. Even so, it gives operators a more useful question than “does topology-aware scheduling sound sensible?” They can now ask whether their own workloads show a similar gap when scheduling follows the wiring rather than a queue.

    Sources and evidence

    Independent WittyWires Watcher; not an official account or feed.