Open code from Ai2 keeps model 'experts' (specialized sub networks) resident on each GPU and routes data to them; Ai2's benchmarks show parameters scaling 10× with under 5% throughput loss, on NVIDIA B300 GPU clusters.
Ai2 released open training code this week designed to keep trillion-parameter sparse AI within reach of labs that do not own the largest GPU clusters. The system, called Olmo-core 3, changes how a mixture-of-experts model is split across many chips. Mixture-of-experts, an architecture that activates only a small subset of its parameters for any given input, is how most frontier labs now control cost at scale. By Ai2's own benchmarks, total parameters can grow from 4.6 billion to 47 billion while throughput drops by less than 5%.
Earlier distributed training at Ai2, and at many peer labs, repeatedly gathered a model's full set of "expert" weights onto whichever GPU was about to process the next microbatch of text. The model weights moved to the data. Olmo-core 3 inverts that: each GPU keeps its assigned experts resident in memory, and the data flows to whichever experts need to handle it. Expert parallelism and pipeline parallelism are layered on top, along with a distributed optimizer that handles gradients across the cluster. As Ai2 describes it, the swap is "from per-microbatch FSDP weight gathering to a DDP-based system that keeps experts resident and routes data to them."
Mixture-of-experts, shortened to MoE, has always carried a coordination cost. The experts have to talk to each other, and that communication gets expensive as the cluster grows. Olmo-core 3 is a bet that most of that cost can be absorbed by data movement, not weight movement. Ai2 says increasing experts from 8 to 128, with four selected per token (around 3.2 billion active parameters), kept throughput within 5% of the smaller configuration.
The benchmarks are Ai2's own. On a preliminary comparison using eight NVIDIA B300 GPUs, the 47-billion-parameter MoE in Olmo-core 3 reached 52,000 tokens per second per GPU, versus 19,400 for Ai2's earlier implementation, roughly 2.7× faster. The comparator is not Megatron-Core, the Nvidia-developed MoE training system used at some peer labs, so the number does not generalize. A controlled four-B300 test with uniformly distributed expert workloads reported about 21% higher throughput using MXFP8 (a low-precision format that packs eight bits per parameter) selectively versus standard BF16, with peak active memory falling from 103 GiB to 95 GiB. Both are single-vendor numbers.
Ai2 reports a 1.2-trillion-total-parameter configuration with 58.36 billion active parameters running on 512 B300 GPUs, hitting a highest observed throughput of 858 TFLOP/s/GPU. The lab separately ran a 2.38-trillion-parameter short-capacity test using DeepEP v2, the latest version of a low-latency expert-communication library. Both used random routing rather than trained-model evaluation, so they show what the system can carry, not what the resulting model can do. Ai2 does not publish a dollar training cost.
"Open" here does not mean "free." The public repository ships source and PyPI installation, optional kernel dependencies, and Docker images, but is tested in-house on H100 clusters and warns that different hardware or CUDA/driver setups may need adaptation. Running Olmo-core 3's trillion-parameter configurations still requires real GPU capital: a B300 cluster is not a hobbyist rig, and the trillion-parameter benchmark is a vendor claim, not an independently replicated result. MoE routing and inter-GPU communication remain an open research problem; Ai2's design is one approach, not a solved one.
Olmo-core 3 is one node in a broader open MoE training landscape, not a singular arrival. Megatron-Core, DeepSpeed-MoE, and a handful of community efforts are chasing the same coordination-cost bottleneck. What Ai2 is releasing is a non-profit lab's bet that the bottleneck can be moved, the recipe can be open, and the floor for who can train a trillion-parameter model can be lowered, provided the lab already has a serious cluster to stand on.
The next test is whether independent runs reproduce the 1.2T throughput number and what a real trained model, not a random-routed capacity test, looks like on top of the new stack. Ai2 says detailed ablations will follow in a technical report. Until then, the headline is that the move from weight-gathering to data-routing is real, the speedup is real on Ai2's own GPUs, and the trillion-parameter claim is a ceiling, not a floor.