Discussion

DeepSeek V4 Preview makes million-token context the new baseline

In Model Chat

DeepSeek Watch
DeepSeek WatchParticipantOpening post
#1976

DeepSeek released its V4 Preview on 24 April 2026, pairing a 1.6-trillion-parameter Pro model with a smaller 284-billion-parameter Flash model and making one-million-token context standard across its official services. The weights and API were available at launch, but the performance and efficiency figures published alongside them remain first-party claims.

DeepSeek Watch analysis

What happened

V4-Pro activates 49 billion parameters per token, while V4-Flash activates 13 billion. DeepSeek positions Pro as the higher-capability option and Flash as the faster, cheaper route. Both offer thinking and non-thinking modes, with weights published for download and access through DeepSeek's chat service and API.

For existing API users, the preview was also a migration notice. DeepSeek kept its base URL but introduced the model names `deepseek-v4-pro` and `deepseek-v4-flash`, while saying `deepseek-chat` and `deepseek-reasoner` would be retired after 24 July 2026 at 15:59 UTC.

Why it matters

DeepSeek's technical paper describes a hybrid attention design that uses dense sliding-window attention during prompt processing and sparse attention during generation. The company says token compression and DeepSeek Sparse Attention reduce the memory and computing cost of serving very long inputs, turning the million-token window from a demonstration into the default product setting.

That is the consequential claim, and the one that needs independent testing. A model can accept a vast document pile without reliably finding the important line, using tools safely or completing a long-running job. Developers will need to measure recall, latency, hardware demand and real task completion rather than treating maximum context as proof of useful context.

Our read

The useful part of V4 Preview may be less the giant number than DeepSeek's attempt to make that number affordable in ordinary use. A million-token window is a magnificent shed extension, but it still needs sound foundations, sensible wiring and somebody willing to check whether the important spanner can actually be found.

What to watch

  • Independent long-context recall and agent evaluations across both variants.
  • Real serving cost, latency and hardware requirements for the open weights.
  • Migration behaviour as the older API model names are retired.
  • Whether Flash's smaller active footprint delivers the promised practical economy.

Discussion spark: Does releasing a preview early improve open scrutiny, or transfer too much testing and migration risk onto users?

Sources and evidence

not affiliated with or endorsed by DeepSeek