A Google Cloud team rebuilt the Allen Institute's open Olmo 3 on Google's TPU chips. Held out tests matched the original, and exposed a data loader bug that made the new stack look better than it was.
An open 7-billion-parameter language model from the Allen Institute for AI (Ai2) was rebuilt on a different company's AI chips and software stack, and matched the original on held-out tests the new team had never trained on. The result, published September 24, 2026 by Google's MaxText team, is a working demonstration that frontier-scale open training can move between chip families and frameworks without changing what the model learned.
The model in question is Olmo 3, a 7B-parameter language model from Ai2 whose training data, code, configurations, checkpoints, and evaluation logs are all public. The Google Cloud MaxText team set out to reproduce both Olmo 3's stage-1 pre-training and its stage-2 mid-training anneal on Google's TPU chips, using MaxText, an open-source training framework built on JAX, rather than the PyTorch stack the original recipe used.
The first hurdle was making sure the port was actually the same model in a different costume. The team converted the Olmo 3 reference checkpoint from PyTorch to JAX, then ran both checkpoints on the same inputs and compared their next-token probability distributions. A divergence of about 1.5 × 10⁻³ in KL (a standard measure of how much two probability distributions differ; values near zero mean the models behave the same) is roughly the noise floor for "same model, different framework." At full 8,192-token context in bfloat16, the two checkpoints agreed on the top-predicted next token 98.75% of the time. That is the floor under everything that follows: the comparison is between two genuinely equivalent models.
A reproduced training run is only as good as its evaluation, and the team ran the same held-out benchmarks the original Olmo 3 release used. Early results showed MaxText beating the reference. That looked like a win for the new stack until closer inspection revealed what the win actually was. A data-loader bug had let MaxText's evaluation set leak into the training set. The "improvement" was memorization of answers the model had already seen. The held-out eval caught the leak; without it, the new stack would have shipped a falsely flattering number. The fix was a loader change, not a model change, and the corrected run matched the reference.
Reliability over weeks of training is the other half of "portable." Olmo 3 stage-2 training spans roughly 1.4 million steps; the team needed to know a saved checkpoint could be resumed without drift. Their A/B test answered that with no ambiguity: a checkpoint saved at step N, reloaded into a fresh run, and compared against a run that never stopped, agreed to Δ = 0.000 on logged loss and perplexity (the standard measure of how surprised a model is by held-out text; lower is better) at every step. When a host actually failed during stage-2, the team retrained the last 127 steps from the last good checkpoint. The retrained run matched the uninterrupted one to the same six-decimal precision. Bit-exact resume is what makes multi-week jobs survivable, and the team has published the numbers behind that claim.
The reproduction also exercised two kinds of elasticity that production training demands. At step ~1.05 million the cluster lost three quarters of its capacity. Rather than stop, the team resumed on the remaining one quarter of the slice with no recipe change, holding global batch size constant. Per-device throughput stayed within 1% in both directions, a result the post describes as approximately 100% strong scaling: each device kept doing roughly the same amount of useful work even as the total job shrank. Stage-2 also swapped chip generation mid-recipe, from Ironwood to v5p, by changing the device type in the launcher. The job kept running and reached 57.4% Model FLOPs Utilization, or MFU, the fraction of theoretical chip compute actually used for training, on v5p, against 44.5% MFU on Ironwood in stage-1 after SparseCore collective offload, rematerialization tuning, and optimal sharding. The team describes the stage-1 figure as roughly a third of the compute budget bought back.
The architecture port itself is a short list of 7B-scale choices that no longer need a footnote: reordered-norm blocks, QK-norm, and a 3:1 ratio of sliding-window to global attention layers. The MaxText team reimplemented each in JAX, then ran the logit-parity check to confirm the port.
The post is not a cross-vendor benchmark. It is a single team's reproduction of one open model on one chip family, with the same vendor's framework. Ai2's open release is what made the reproduction possible; the held-out evaluation gate, the bug catch, and the bit-exact resume are what made it trustworthy. The next test is whether a different lab can run the same public recipe and the same run script on a third stack and get the same answer.