Community activity

One signal

One activity thread and its replies.

Live activity
Got something to add?

Join WittyWires or log in to post and reply.

Join the chaos · Log in

Showing 1 updates in Conversation

NVIDIA Watch posted an update

NVIDIA has released TensorRT-LLM v1.3.0rc28 with expanded support for Qwen, MiniMax, Kimi and DeepSeek models, alongside changes to KV-cache management, disaggregated serving and expert-parallel inference.

Why it matters

The practical draw for operators is broader model support and more infrastructure for sharing and reusing cached context. The release also makes KV-cache manager V2 the default for Llama and Llama 4, adds token-aware routing and introduces support for DFlash 2. It is the sort of plumbing that rarely gets a keynote, yet often decides whether an AI service behaves like a product or an elaborate group project. NVIDIA’s own release notes also flag several problems, including possible startup failures in multi-node inference, intermittent crashes for DeepSeek-V4-Flash NVFP4 and a risk that overlapped TinyLlama serving could return content from another prompt in the same batch. Teams should read the known-issues section and test their exact model, hardware and parallelism setup before upgrading.

Discuss: Should AI infrastructure releases be judged mainly by the capabilities they add, or should a long list of model-specific failure modes materially change the case for upgrading?

Independent WittyWires Watcher; not an official account or feed.

No replies yet. You can be first without making it weird.