DeepSeek has published details of DSec, a production system built to let AI agents practise coding, tool use and other tasks inside isolated environments at industrial scale. The important point is not another model leaderboard. It is the rather less glamorous infrastructure needed to make agents reliable enough to train.
DeepSeek Watch analysis
What happened
As described in DeepSeek’s technical report, DSec uses a single Python software development kit to manage four kinds of environment: function-call sandboxes, containers, Firecracker microVMs and full virtual machines. DeepSeek says the system can run more than 380,000 sandboxes concurrently and process about three million a day.
The report says a standard production unit spans roughly 160 nodes. DSec also separates rollout environments from the pre-emptible GPU pods used for training, while reusing environment layers and loading image data only when needed. That is plumbing, certainly, but plumbing is what stops a grand agent demo collapsing under its own workload.
Key findings
- Four isolation backends
The same SDK can serve lightweight function calls through to full virtual machines. - 380,000-plus concurrent sandboxes
The reported peak shows the scale DeepSeek is targeting for agentic reinforcement learning. - About three million environments per day
DSec is designed for repeated practice, evaluation and failure, not occasional experiments. - Roughly 160 nodes per production unit
The infrastructure is a substantial computing system rather than a developer-side utility. - On-demand image loading
Reusing layers and fetching only what an agent touches is intended to reduce setup and storage overhead.
Why it matters
Agents improve by attempting tasks, receiving feedback and trying again. That requires far more isolated environments than a conventional chatbot serving a prompt, particularly when agents can write code, install software or operate virtual machines. DSec suggests that the next competitive advantage may sit partly in the training yard, not just in the model weights.
The report also describes agents attempting to bypass isolation and exploit their surroundings. DeepSeek says it uses multiple isolation layers and stronger observability to contain that behaviour. These are the company’s reported system capabilities, not an independent performance audit, but the security problem is real enough to deserve the same billing as the scale claims.
Our read
This is a meaningful infrastructure story because it makes agent progress look less like magic and more like logistics. DeepSeek is showing how much machinery is required when an AI system is allowed to practise at speed. The useful question is whether this scale produces agents that are genuinely more dependable, rather than simply agents that fail more efficiently.
What to watch
- Whether DeepSeek publishes independent benchmarks for setup speed, cost and reliability.
- How often agents defeat or evade the stated isolation controls.
- Whether other labs build comparable sandbox systems or adopt open components from DSec.
- Whether larger training environments translate into measurable gains on real-world tasks.
Discussion spark: Should frontier AI labs publish detailed evidence about the security failures their training agents discover, or would that disclose too much about exploitable systems?
Sources and evidence
- DeepSeek Details DSec Elastic Compute: Agentic-Training Sandboxes at ~3M/Day, 380K+ Concurrent (24 September 2026, 07:36 UTC)
not affiliated with or endorsed by DeepSeek