DeepSeek, Qwen, Kimi, and GLM diverged on the iterated prisoner's dilemma, a standard repeated cooperation game, more than the Chinese and Western ecosystems diverged on the same measure.
A new pre-registered study put four leading Chinese AI systems through a standard repeated-cooperation test, and the four labs diverged from each other on the same measure more than the Chinese and Western ecosystems diverged from each other. The authors argue the result rejects "Chinese AI" as a bloc.
The study ran an evolutionary Iterated Prisoner's Dilemma tournament on four models the authors call frontier-tier: DeepSeek V4 Pro, Qwen3-Max, Kimi K2.5, and GLM-5.1. The Iterated Prisoner's Dilemma is the textbook cooperation test: two players choose each round whether to cooperate or defect, and the payoffs reward mutual cooperation over mutual defection only when neither side tries to exploit the other. The "evolutionary" version runs many tournaments in parallel, lets successful strategies spread, and asks which behavior takes over the population.
To control for an obvious confound in prior work, that each lab's model might be better or worse at writing the code that plays the game, the authors fixed the strategy-writing model to GPT-5.4 Mini across all four labs. The labs' agents wrote the cooperation strategies. The converter that translated strategies into code stayed the same.
The proportion of population runs that ended in an aggressive equilibrium ranged from roughly 1% at Qwen3-Max to 9% at DeepSeek V4 Pro, an 8-percentage-point range across the four Chinese labs. Four of the six pairwise comparisons survived a Holm-Bonferroni correction for multiple testing (a standard statistical adjustment that protects against false positives when running many comparisons at once). The within-China spread exceeded the cross-ecosystem mean gap the paper reports between Chinese and Western frontier models on the same measure. The paper's hypothesis H6, that the four Chinese labs are not monolithic in their cooperative disposition, is supported.
The other leg, the cooperative-bias claim, is qualified. The paper's H5 asks whether the cooperative bias reported for Western frontier models in prior work (9 of 12 lab-prompt combinations) holds for Chinese models. In the main run, the Chinese count was 6 of 12, dragged down by near-ties the authors label Cooperative-Neutral. Under a pre-registered alternate-converter robustness check, the Chinese count rises to 9 of 12. The authors describe the result as "consistent but qualified."
Two pieces of provenance. The 9/12 Western figure comes from Phase 1 of the same project, archived on Zenodo; the present study does not produce it. The 5.0% vs 5.0% Chinese-vs-Western mean aggressive-equilibrium proportion that the paper reports as its within-paper cross-ecosystem headline is an internal number on the IPD/Moran protocol, not an external benchmark readers can cross-check against public leaderboards.
The pre-registration is public on GitHub, the full replication package is released, and the full HTML of the paper is available for cross-check. The work is an arXiv preprint. The hypotheses and the design are ante-hoc and the code is open, but the findings are not peer-reviewed.
The paper's claim is internal to the IPD/Moran setup and does not extend to multi-agent deployments, real-world negotiation, or model behavior outside the lab. The cross-ecosystem comparison runs four Chinese labs against a Western set the authors describe as frontier-tier in the same framing, not against a public benchmark. Readers should treat the China-vs-West gap as the authors' within-paper number, not as a scorecard.
The next check is independent reproduction: whether other groups reproduce the 8-point spread and the 6/12-vs-9/12 qualification under the same fixed-converter protocol. The bloc claim, "Chinese AI does X," does not survive the data as written.