Training AI agents is a sandbox problem, not a GPU problem, and DeepSeek just published the math behind 50x oversubscription on the same hardware.
DeepSeek, the Chinese AI lab behind the R1 reasoning model, published a 130-author engineering paper on the sandbox infrastructure that runs all of its AI agent training. In production data drawn from the paper, the system packs 50 times more training sandboxes onto the same hardware, because roughly 90 percent of those sandboxes use less than 5 percent of the CPU they request.
The paper, posted to arXiv on September 19, 2026, details DSec (DeepSeek Elastic Compute), the system that supports every stage of the lab's agent pipeline, from data preprocessing through evaluation of the V3.2 through V4.1 model family. A single shard of DSec runs about 160 servers, 30,000 CPU cores, and 250 terabytes of memory. Across that shard, the system serves roughly 3 million sandbox sessions a day, peaks above 380,000 concurrent, and creates more than 5,000 sandboxes per second at full load.
The headcount is downstream. The cost compression comes from four layers of engineering.
The bottleneck for training AI agents (software that reads code, runs tests, and edits files on its own) is not GPU silicon. It is the isolated sandbox each agent needs to work safely. A single training run can spin up hundreds of thousands at once, and most never touch the bulk of the file system or compute they request. DSec treats that mismatch as the design problem.
First, the file system. Agents only access 4.2 to 13.3 percent of a full base image during runtime, so DSec loads images on demand from 3FS, DeepSeek's distributed storage layer, instead of pulling the whole image upfront. Eager pulls of multi-gigabyte images that used to take more than an hour now complete in about 35 minutes, a 1.71x speedup with 57 percent less disk write. For tar.gz workspaces, an EROFS read-only overlay mount cuts setup from 79 minutes to 45 minutes, with disk write down to roughly one-fifth of the baseline.
Then the memory layer. virtio-pmem, a paravirtualized persistent-memory interface, lets multiple sandboxes share the same page cache in main memory via DAX (direct access). In DeepSeek's measurements, that change alone cuts host peak memory by about 40 percent. A second pass using DAMON, a Linux kernel memory monitor, plus a balloon driver that reclaims unused pages from running guests, adds another 21 percent on top.
Finally, the scheduler. Because most sandboxes idle, DSec oversubscribes CPU aggressively, packing many more workloads onto each node than its rated capacity, on the bet that they will not all peak at once. Production oversubscription runs above 50x. A latency-aware scheduler then shields the small fraction of jobs that actually need real-time response. For those workloads, the share of tasks seeing meaningful delay drops from 45.2 percent to 17.3 percent, even when peer tasks are already using half the node's CPU.
The system unifies four backends behind one interface: FnCall for plain function-calling agents, Container for standard tool-using workflows, MicroVM for hard isolation, and Full VM for GUI, render, and Android workloads. In a single recent week, the Container backend alone served 11,266 distinct base images and 102,171 workspaces, plus hundreds of toolkits.
DeepSeek's stated next step is 100 to 1000 times more environment types and counts. The lab is now publicly hiring senior roles on the elastic compute team, a recruiting move, not a neutral job post, and a confirmation that the paper was a deliberate signal about how the team plans to compete.
Two honest caveats. The 50x oversubscription figure is self-reported, drawn from a single lab's own infrastructure and an arXiv preprint, not an audited benchmark. The paper does not characterize the tail latency or failure modes of workloads that the scheduler mistakenly assumes will stay idle. The number is a working hypothesis about how much headroom is left on existing hardware for training agents, not a verdict.
The team says the next infrastructure milestone is the 1000x environment target. The math behind that ramp is now public.