DeepSeek releases a new paper, publicly unveiling the "headquarters" of V4.1 Agent training for the first time, with Liang Wenfeng as the signatory.
DeepSeek and Tsinghua University have released their latest technical report *DeepSeek Elastic Compute (DSec): A Sandbox Infrastructure for Effective Agentic Training at Scale*, publicly unveiling for the first time DSec, the "industrial foundation" that underpins DeepSeek's training from V3.2 to V4.1.
The paper's author team has more than 130 members, with Liang Wenfeng ranked last on the list. This shows that the infrastructure for agentic training has been treated as a core engineering priority by DeepSeek.
DSec is a production-grade self-developed elastic computing sandbox platform of DeepSeek, an exclusive internal infrastructure that was previously introduced in the DeepSeek-V4 report. It is responsible for providing a secure and stable sandbox runtime environment for large-scale training and evaluation of large model agents.
The paper mentions that one production unit consists of approximately 160 CPU nodes, 30,000 CPU cores and 250TB of memory, hosting petabyte-level images. In terms of capability, DSec can serve about 3 million sandboxes per single day, with a peak concurrent volume exceeding 380,000, and a creation speed of over 5000 sandboxes per second. A single training task can pull up to 32,000 sandboxes at one time at most.
This paper not only provides system-level architecture solutions, but also records a large number of dramatic gameplays where models inside the sandbox "attempt to escape, peek at answers, and crash the system".
If the past Large Model Scaling refers to stuffing more parameters, data and Tokens into GPU clusters, what DSec reveals is the greatly underestimated expansion in the Agent era: when models enter the real environment, the number of repeated task execution and trial-and-error learning is also being scaled.
Model-generated code is not the end point, and the real bottleneck lies in the physical infrastructure that carries the safe, high-speed and concurrent operation of these codes.
Failure of Traditional Cloud Architecture
To understand DSec, we must first see the essential fundamental differences in underlying logic between Agentic Reinforcement Learning (Agentic RL) and traditional large model training.
In the past, when training language models, the interaction logic was closer to "answering questions": input prompts or math problems to the model, the model generates a derivation or answer, and the system calculates rewards according to rules and backpropagates to update parameters.
However, agents oriented to real Software Engineering (SWE) or Computer-Use operations are completely different. After receiving a task, the agent needs to retrieve code repositories, read context, configure dependent environments, modify code logic, execute tests in the terminal, and continue debugging according to error feedback. In this process, the model is not statically outputting text, but actually operating a computer.
The cumulative distribution (CDF) of single-task sandbox creation volume, showing the transient pulse characteristics of thousands or even tens of thousands of sandbox concurrent requests for a single task in production
In a single Rollout of reinforcement learning, if thousands of agents explore simultaneously, the system must instantly deliver tens of thousands of isolated execution environments. More critically, these environments must be stateful and highly consistent. The commands executed by the agent at the 10th step must be strictly based on the modified files, compiled binaries and started service processes from the previous 9 steps.
However, traditional container orchestration represented by Kubernetes and stateless Serverless are built on the underlying assumptions of smooth scaling and stateless microservices. Facing this high-frequency pulse, long-resident and extremely heterogeneous computing workload, they directly encounter three structural mismatches:
The sandbox lifecycle distribution curve intuitively presents the long-residence phenomenon where long-tail sandboxes survive for more than 3 hours
Transient Pulse Concurrency: A single task can apply for tens of thousands of sandboxes within seconds (peak value reaches 32,000). Full pull and decompression will instantly cause severe write I/O congestion, greatly prolonging startup latency and forcing GPU clusters to idle;
Massive Heterogeneity and Extremely Low Reuse: Facing environment data exceeding 130TB, the median fanout of image reuse is only 1 to 3, and conventional local caching mechanisms completely fail;
Long Residence and Computing Power Tide: Sandboxes survive for tens of minutes or even hours, but CPU utilization is less than 5% for 90% of the time, resulting in extreme waste of exclusive resources, while rude overcommitment easily causes memory overflow and hyper-threading contention.
Traditional cloud computing solutions cannot take into account high concurrency, strong state and low cost at the same time. DSec is exactly the "specialized digital training ground" built by DeepSeek to solve these contradictions.
Tradeoffs of Four Types of Backends
Since general container clusters cannot be directly applied, the most intuitive idea may be to build an "ultimate sandbox with the best performance and most complete functions". But DeepSeek's engineering practice proves that a single virtualized runtime cannot simultaneously balance security isolation and resource overhead.
Different tasks have natural gaps in environmental requirements: running an operator script only takes a few milliseconds, system attack and defense requires strict hardware-level virtualization, while running the Android emulator relies on a complete operating system and graphics drivers. If heavy virtual machines are all used, hardware costs will quickly get out of control; if all light containers are used, security and environment compatibility cannot meet the requirements.
Therefore, DSec abandons the fantasy of a universal sandbox, and divides the environment into four types of echelon layouts:
DSec overall architecture topology diagram, showing the layered interaction of libdsec, cluster management layer (IAM, Placement Engine), single-machine runtime (Edge, Aether, Chronus) and 3FS
FnCall (Function Call): Oriented to short-term stateless tasks (such as algorithm problem judging or GPU operator benchmark test), running based on resident preheating pool to eliminate cold start latency, supporting exclusive or shared GPU;
Container (Standard Container): Mainly carrying software engineering and tool interaction, a single machine can deploy 3,200 instances at high density, balancing lightweight features and basic Linux toolchain compatibility;
MicroVM (Lightweight Virtual Machine): Builds independent kernel based on Firecracker, provides hardware-level isolation for security attack and defense and adversarial tasks, and prevents cross-tenant escape;
Full VM (Full-Featured Virtual Machine): Based on QEMU and connected to virtualized GPU drivers, it is dedicated to advanced environments that require complete commercial operating systems, graphical interfaces, Android emulators or game rendering.
Comparison of four types of sandbox backends (FnCall, Container, MicroVM, Full VM) in performance, overhead, isolation level and scenario adaptation
The essence of this design is: the environment does not pursue unification, and what is unified is the control interface for accessing the environment. The platform shields environmental differences from the upper layer through a unified Python SDK (libdsec). Algorithm R&D personnel only need to declare the calculation type and resource upper limit required by the task in the request, and the scheduling system will automatically route the task to a suitable computing carrier.
Computing Power Tide Overcommitment
Establishing four types of sandbox backends is only to build "rooms" for carrying tasks. The real problems that push the system to the limit follow: if calculated according to conventional resource quotas, 160 physical machines cannot accommodate 380,000 concurrent sandboxes at the same time at all.
To make a single computing node stably carry up to 3,200 containers or 800 MicroVMs, the core is to grasp the special operation rule of agent computing: sandboxes do not consume CPU non-stop all the time.
In actual interaction, when the model is thinking, the sandbox waits silently; after the model issues instructions, the CPU usage instantly spikes; after execution is completed, the computing power quickly drops back to a low level. Data shows that 90% of sandboxes have an average CPU utilization of less than 5% of the applied quota during their lifecycle. Since computing power is highly pulsed and tidal, the system has an objective basis for implementing extreme resource Overcommitment.
Curve of CPU pulse and memory resident during the sandbox lifecycle
Extremely sparse distribution feature of actual CPU/memory usage
However, high-magnification overcommitment easily causes system-level stampede. DSec completes three key closures at the kernel layer:
- Read-only memory penetration sharing: Enable Virtio-pmem in MicroVM with DAX technology, directly map read-only disk content to host memory, bypass the guest's own page cache, make hundreds of virtual machines on a single machine completely share the same physical memory of the host, and the peak memory of the host drops by 40.2%;
Comparison curve of combined memory optimization measures
Cold memory dynamic eviction: For writable disks, use the Linux kernel DAMON mechanism to periodically sample memory activity, actively eliminate long-term idle cold pages, and the virtual balloon device returns physical pages to the host, further reducing long-term memory consumption by 21.2%;
Hyper-threading hard isolation: To prevent background loads from interfering with tasks with strict response latency (such as game-playing agents with limited step frequency), the system grants scheduling priority to key tasks while enabling kernel Core Scheduling, strictly forbidding low-priority tasks from occupying the twin hyper-threads of the same physical core, and strictly suppressing the increase of long-tail latency under high load from 45.2% to 17.3%.
It is through eliminating repeated memory occupation and suppressing computing contention that high-density overcommitment has turned from a paper concept into a production-available reality.
Dual Decoupling of Storage and Training
In addition, the system must also complete complete loosening in data transmission and collaborative architecture.
The first is to say goodbye to full-volume image copying and realize streaming on-demand loading. Facing the massive heterogeneous environment of over 130TB, monitoring shows that the image data actually accessed by sandboxes only accounts for 4.2% to 13.3% of the total volume.
DSec splits the image into three immutable layers: basic system, task code area and toolchain. Combined with the EROFS compressed file system and self-developed 3FS distributed storage, it achieves "the background concurrently retrieves the corresponding part only when the agent execution involves it", and runtime writes are strictly stored on the local disk.
This design directly eliminates I/O congestion during concurrent startup, reduces disk writes by 57%, and shortens task duration by more than 40%.
Performance comparison between EROFS on-demand streaming pull and traditional Docker full pull in terms of concurrent container number and disk write volume
The second is to decouple the agent exploration loop from the volatile GPU training cluster. In the architecture evolution of DeepSeek-V4.1, the team completely migrated the Agent Loop that drives multiple rounds of interaction out of the GPU cluster, and entrusted it to DSec's independent containers for hosting.
In large-scale training, GPU jobs frequently encounter scheduling preemption. If the two are bound, the cost of restoring state by replaying logs after interruption is extremely heavy. After decoupling, when GPU preemption occurs, the sandbox only needs to sleep in place and release memory; after computing power is restored, it wakes up transparently in situ to continue advancing long-term tasks.