A new arXiv paper finds that panels of language models can flip into shared bias past a tipping point, and that mixing different kinds of AI models flattens the curve.
AI committees are starting to make real calls: picking investments, grading other models, weighing contested answers. A new arXiv paper warns that once those agents agree with each other past a threshold, the panel can lock into a shared bias instead of canceling out its individual blind spots.
The paper, Emergence of Biased Consensus in Multi-Agent LLM Debates, recasts the problem in physics language. Once conformity in a panel surpasses a critical point, the group tips into a collective bias. The authors identify sampling temperature, the knob that controls how varied an LLM's outputs are, as a key driver of that flip. Mix models from different families and the transition smooths out instead of snapping.
A homogeneous panel of the same model, prompted the same way, can entrench that model's blind spots. Panel diversity, not head count, is the safety dial.
The paper's experiments are controlled and the abstract does not expose a specific temperature or conformity threshold. Its claims to generalize to investment calls and LLM-as-judge evaluation come with no public benchmark numbers. For now, the paper describes a phase transition, not a number operators can tune to.