AI alignment research builds refusal, filtering, and 'preference shaping' tools — the training tricks that teach models which answers to prefer.
The training tricks that make an AI model refuse to write a phishing email are the same tricks an authoritarian deployer can use to make that model refuse to mention a protest. A position paper accepted as an ICML 2026 Oral makes the structural argument that alignment research is dual-use by construction. The refusal, filtering, and preference-shaping tools built to keep models from causing harm are also a turnkey censorship engine for the actors who control the deployment stack.
The paper, titled 'Position: The Alignment Community is Unintentionally Building a Censor's Toolkit', is a position piece rather than an empirical study. The authors map current alignment techniques to two tracks: the possibility of misuse, and actual cases of misuse. Their core claim is structural. The pursuit of a 'perfectly aligned' model hands malicious actors an ever-improving tool for informational dominance, because the very mechanisms that constrain model output can be re-pointed at any output the operator wants suppressed.
They name three amplifiers. The first is speed: user adoption of AI as an information provider is still moving fast, and assistants are replacing search as the first stop for many queries. The second is economic concentration: the companies that train frontier models sit on capital and compute that no individual researcher, regulator, or end user can match. The third is political: the paper's authors flag a global drift toward authoritarian governance, which raises the demand side for any tool that can shape what an AI system is willing to say.
The paper's most useful contribution is the technique-by-technique mapping. Refusal training, in which a model is taught to decline certain categories of request, maps to topic suppression: an operator who controls the training pipeline can move any subject from 'harmful' to 'sensitive' to 'off-limits' by retraining or fine-tuning. Reinforcement learning from human feedback, originally designed to make models more helpful and less toxic, maps to narrative control: the same preference signal that pulls a model away from rude answers can be tuned to pull it toward a preferred framing of contested events. Output filters and content classifiers, layered on top of model outputs, map to keyword blacklists: the same classifier that screens for hate speech can be re-pointed at dissident vocabulary with a new label set.
They call for the alignment community to take dual-use seriously as a design constraint, not a post-hoc concern, and they propose a menu: transparency reports on how refusal categories are constructed, third-party auditability of alignment data, re-deployment guardrails that lock a fine-tuned model against silent re-tasking toward suppression, and an explicit refusal to publish techniques whose only safe counterpart is the operator's good faith. The constructive ask is to design for the worst-case deployer, not the best-case one.
ICML 2026 runs the oral session this month, according to the conference schedule, which puts the dual-use thesis in front of the main ML research audience at the moment AI assistants are absorbing the search-and-answer role that used to belong to a list of links. Once a model is the main doorway to information, its refusal set is its editorial line, and editorial lines are easier to audit before they harden than after.
The paper treats the alignment community as the right place to design the dual-use guardrails, not as the source of the censorship problem. The argument is that the people who build refusal and preference-shaping mechanisms are the ones with the technical knowledge to make those mechanisms auditable, and that waiting for outside regulators after the suppression patterns harden leaves the field reacting to harm it could have designed against. The full preprint, the HTML version, and the authors' project page carry the case-mapping in detail.