A 300 million parameter transformer trained only on randomly generated sequences can compress real Wikipedia in six languages when given new data.
A 300-million-parameter byte-level transformer was pretrained on nothing but randomly generated sequences and then handed frozen weights. Given up to a million bytes of real Wikipedia in any of six languages, it learned to predict the next byte well enough to cut its compression loss by more than two-thirds, according to a new preprint that lands the prior-fitted network recipe on natural language for the first time.
The model, called PFLM, is small by today's standards and openly downloadable. The interesting part is what it never saw during training: any human text. Every training sequence is drawn fresh from a synthetic prior built around a recurrent structural causal model: a generator that produces plausible-looking character streams without ever being wired to a real corpus. The model is shown each sequence once, so to do well on a new sequence at test time it has to infer the rules of that sequence from context alone, with its weights frozen.
That is the bet. If the training distribution is the right shape, the network can become a general in-context inference engine over latent languages, instead of a memorization engine over a fixed corpus. The authors, working under the cbl repository, reuse the prior-fitted network (PFN) idea that powered TabPFN for small tabular data and apply it to byte sequences. The shift from tables to text is what makes the paper more than a controlled demo: it tests whether PFN-style inductive bias survives the jump to a domain where humans already have strong intuitions about what "real language" looks like.
The headline result is the bits-per-byte curve. On Wikipedia in English, Chinese, Hindi, Arabic, Japanese, and Korean, the model's next-byte prediction loss falls from the uniform baseline of eight bits per byte to between 0.9 and 2.4 bits per byte at one million bytes of context, the paper reports. That is far worse than what classical language models trained on trillions of tokens achieve on the same text, and the authors say so plainly. The point is not parity with GPT-class systems. The point is that a 300M model with no exposure to real text can compress any of six human languages at all, just by reading them at test time.
The same in-context machinery also picks up structure the model was never trained on. Given enough numerals in its prompt, PFLM learns to count, to compare magnitudes, and to perform approximate addition. It also predicts deterministic sequences, including the prime numbers and the Kolakoski sequence, entirely from context, the preprint and the model card report. These are not a calculator or a programming language; they are evidence that the synthetic prior prepares the network to recognize algorithmic patterns it has never been shown, as long as the patterns show up in the prompt.
The honest ceiling is in the abstract. PFLM is acknowledged to be far worse on text than classical language models trained on trillions of tokens, and at test time it sees at most a million bytes of a single language. That is enough to make the experiment interesting, and not enough to make it a product. The result is a proof of behavior, not a recipe for shipping a chat model: the right training distribution can teach a small network to do in-context inference over languages, but the ceiling on real text still belongs to large models trained on the open web.
The next test is whether the same prior-fitting recipe compresses languages outside the original six, and whether prior-fitted-network researchers read the result as a meaningful extension of TabPFN or a controlled demo. The artifacts are public: the paper, the code, and the weights are all downloadable now.