The standard way to score a model for censorship is to ask about foreign-interest topics: Tibet, Taiwan, Xinjiang, the Dalai Lama. The model that fails those is the one the field learns to flag. That battery has a structural blind spot: it cannot see a list.
CTGT's LineageEval probe on Ox Alpha, the model that surfaced on OpenRouter on August 20, found exactly that. On Xinjiang and Taiwan, Ox Alpha answers in detail and reads to the standard suite as identical to GPT-OSS-120B. On seven domestic topics, including Xi Jinping personally, it refuses in Chinese state voice and is statistically indistinguishable from the most censored model CTGT has tested.
Most responses are short and permitted; a small set is refused hard. CTGT's matched-pair data makes the concentration legible: seven probe pairs contribute essentially the entire censorship score, while the rest contribute nothing. Averaged, the model looks fine. Examined at the right granularity, it is one of the most censored models in production.
The mechanism travels. A foreign-interest audit will wave any narrow-list censoring model through. A model with distributed political shading will fail foreign-interest audits, even when nothing is actually refused. The same averaged score can hide either shape, and the field has been scoring the wrong one. The question worth asking of the next benchmark is not whether the model answered Tibet. It is whether the battery tested domestic politics at all.
Reported by Sky for Type0, from Behaviorally Fingerprinting Ox Alpha's Provenance and Censorship. Read the original: ctgt.ai