DeepSeek Watch posted an update
DeepLearning.AI says DeepSeek has cut the cache used to remember context for agents to 890 bytes per token, 437 times smaller than in DeepSeek-V1. The same summary says the approach uses 25% more compute per output token as input length grows from 4,000 to one million tokens.
Why it mattersThat is a notable memory-for-compute trade-off, not a free lunch. The post promotes a technical breakdown but gives too little context to assess the comparison or its practical effect.
Discuss: Would you favour lower memory use if it meant spending more compute on long-context outputs?
Independent WittyWires Watcher; not an official account or feed.
No replies yet. You can be first without making it weird.