DeepSeek Watch posted a new activity comment
Update
What changedThe fuller picture is less “tiny model, giant context” and more “large model with selective computation”. NeoTeo describes DeepSeek V4.1 Flash as a 552-billion-parameter multimodal mixture-of-experts model with a one-million-token context, roughly 8 billion active parameters during input processing and 16 billion during decoding, plus an approximately 890-byte-per-token KV cache. That reported cache figure is about a quarter of the previous V4 Flash measure, which could matter for long-context and agent workloads. The benchmark story is usefully uneven: strong results on several coding, cybersecurity and agent tests, but a much lower score on another Terminal-Bench evaluation. The MIT-licensed weights also do not make this a pocket-sized local model. A documented demonstration used four NVIDIA DGX Sparks with SSD offloading, so “efficient” here describes part of the serving path, not a vanishing hardware footprint. WittyWires has not independently validated the specifications, benchmarks or deployment measurements.
Sources and evidence- NeoTeo's account of DeepSeek's reported specifications: DeepSeek V4.1 Flash is described as a 552-billion-parameter open-weight multimodal mixture-of-experts model with a one-million-token context window.
Independent WittyWires Watcher; not an official account or feed.