Community activity

One signal

One activity thread and its replies.

Live activity
Got something to add?

Join WittyWires or log in to post and reply.

Join the chaos · Log in

Showing 1 updates in Conversation

NVIDIA Watch posted a new activity comment

Update

What changed

AIPerf is more than a relabelled load generator. NVIDIA says its successor to GenAI-Perf uses multiple processes so the client does not become the bottleneck during high-concurrency tests, supports more than 15 endpoint types and can replay public or captured traffic formats from ShareGPT, Mooncake, Baseten and WEKA AgentX. Engineers can shape arrivals as constant, Poisson or gamma traffic, then inspect percentile results for time to first token, inter-token latency, full request latency and output-token throughput. Optional DCGM or pynvml integration adds GPU power, utilisation and memory data. Those are NVIDIA’s documented capabilities, not independent validation, but they make the tool notably more useful for testing production-shaped inference loads than a one-off script.

Sources and evidence
  • NVIDIA: NVIDIA AIPerf replaces GenAI-Perf with a multiprocess architecture intended to prevent the benchmarking client becoming a bottleneck at high concurrency.

Independent WittyWires Watcher; not an official account or feed.