The first fight over AI's economics is being misread. Distillation, training a smaller model to mimic a larger one's outputs, has been framed as a frontier-killer: copy the teacher, undercut the teacher, win the market. LatePost's editors argue the opposite on episode 179 of their podcast, built from past reporting and conversations with practitioners at several Chinese AI labs. The algorithm is the easy part. The moat is operational.
Building a student that actually approximates a frontier teacher requires account provisioning at scale, real user questions, and a data pipeline that survives both. That plumbing is what most labs cannot replicate quickly, which is why so few discuss the practice on the record, even as nearly everyone runs it. The technique is commoditized. The capacity is not.
Agent trajectory distillation, which expands the distillable surface from single answers to multi-step reasoning and tool-use traces, has raised the operational bar with every generation of frontier model. ByteDance founder Zhang Yiming reportedly told his team the opposite of the conventional playbook: skip distillation, do not chase Coding benchmark hype. His stated reasoning, that a student can at most approach a teacher and rarely truly surpass one, is the structural argument that survives.
Anthropic and Google Cloud's threat teams now treat distillation as an exfiltration risk measurable at the API level, which lets defenders price the gap between student and teacher in security logs rather than leaderboards. LatePost's frame, that frontier advantage is operational know-how rather than the pedagogy, holds against that evidence. When a lab claims parity with a frontier model, the question is which data infrastructure produced the student, not which teacher it copied.
Reported by Ava for Type0, from 179: 蒸馏风暴:一场无人公开谈论的技术竞赛. Read the original: podcast.latepost.com