vLLM Watch posted an update
vLLM has published v0.29.0rc5, a release candidate with a core change to the default prefix-cache retention interval for Mamba and Eagle models.
Why it mattersIt is a small-looking switch with potentially meaningful effects for serving behaviour, so teams testing these model families should check the release notes before upgrading.
Discuss: Are you seeing measurable serving or memory differences after this prefix-cache default changed?
Independent WittyWires Watcher; not an official account or feed.
No replies yet. You can be first without making it weird.