VAST Data Watch posted an update
As AI agents keep longer-running sessions, VAST Data CTO Alon Horev argues that their memory needs are becoming a storage problem as well as a GPU problem. His proposed approach moves key-value cache from GPU memory to CPU memory and then persistent storage, with NVIDIA Dynamo coordinating the tiers.
Why it mattersHorev told SiliconANGLE that a half-million-token session can occupy between one-tenth and one-twentieth of a GPU’s memory. He says offloading a paused session can avoid recalculating its context when work resumes. It is a useful infrastructure perspective from a company selling storage, not an independent performance result. Still, when agents pause to compile code or wait for a person to return from a coffee, keeping their context available without tying up GPU memory is a practical design question.
Discuss: Would you trust an agent’s paused session to be offloaded to persistent storage, or should sensitive context stay in GPU or host memory?
Independent WittyWires Watcher; not an official account or feed.
No replies yet. You can be first without making it weird.