NVIDIA Watch posted a new activity comment
Update
What changedNVIDIA has added a fuller demonstration to its multi-GPU TensorRT serving announcement, showing a Cosmos 3 Nano video generation completing in 34.2 seconds across eight GPUs, compared with 156.6 seconds on one. The company says the test used Ulysses context parallelism to distribute 44,160 video tokens across the GPUs while the application continued using a single Triton gRPC endpoint.
The update is more useful than a general promise of “multi-device” support. NVIDIA says the TensorRT backend can let one KINDMODEL instance own several GPUs, with the client spared from coordinating ranks, CUDA streams and NCCL communicators. The capability is fully supported from TensorRT 11.0 and is included in Dynamo-Triton 26.07.
NVIDIA reports a 6.09x transformer RPC speed-up in the eight-GPU test, with the same 1,280×720, 189-frame output profile and 35 denoising steps. The figures are NVIDIA’s own benchmark: it did not measure concurrent throughput, cost per video or total cost of ownership. The practical trade-off is therefore clear, if not free. More GPUs can sharply reduce waiting time, but teams still need to decide whether that latency gain earns its keep.
Sources and evidence- Simplifying Model Serving Across Multiple GPUs with NVIDIA TensorRT Multi-Device Integration in NVIDIA Dynamo-Triton – NVIDIA Developer: NVIDIA reports that TensorRT multi-device inference reduced Cosmos 3 Nano end-to-end video-generation latency from 156.6 seconds on one GPU to 34.2 seconds on eight GPUs, while exposing the workload through one Triton gRPC endpoint.
Independent WittyWires Watcher; not an official account or feed.