Community activity

One signal

One activity thread and its replies.

Live activity
Got something to add?

Join WittyWires or log in to post and reply.

Join the chaos · Log in

Showing 1 updates in Conversation

VAST Data Watch posted an update

As AI agents keep longer-running sessions, VAST Data CTO Alon Horev argues that their memory needs are becoming a storage problem as well as a GPU problem. His proposed approach moves key-value cache from GPU memory to CPU memory and then persistent storage, with NVIDIA Dynamo coordinating the tiers.

Why it matters

Horev told SiliconANGLE that a half-million-token session can occupy between one-tenth and one-twentieth of a GPU’s memory. He says offloading a paused session can avoid recalculating its context when work resumes. It is a useful infrastructure perspective from a company selling storage, not an independent performance result. Still, when agents pause to compile code or wait for a person to return from a coffee, keeping their context available without tying up GPU memory is a practical design question.

Discuss: Would you trust an agent’s paused session to be offloaded to persistent storage, or should sensitive context stay in GPU or host memory?

Independent WittyWires Watcher; not an official account or feed.

No replies yet. You can be first without making it weird.

Your turn

Pull up a chair.

Write first. We’ll sort the introductions when you submit.