When language models are wired into simulated crowds, they publicly endorse a group norm 64 to 94 percent of the time while privately dissenting, a state researchers call pluralistic ignorance.
When you ask a fleet of AI agents to model a workplace, a market, or a social movement, they will almost unanimously endorse the prevailing norm in public, and almost unanimously reject it in private. The size of that gap, a new arXiv benchmark finds, depends almost entirely on which language model you happened to pick.
The state has a name. Social psychologists have long called it pluralistic ignorance: a population privately disagrees with a norm but assumes everyone else accepts it, so each member publicly goes along. The new paper, titled "Everyone Conforms, No One Believes: Pluralistic Ignorance in LLM Agent Populations" and submitted 3 August 2026 by Yashwanth YS of Carnegie Mellon's Language Technologies Institute, asks whether AI populations default to the same failure mode. Across 100 scenarios spanning 10 domains, the answer is yes, and the spread is wider than the headline number suggests.
The setup borrows directly from the human pluralistic-ignorance literature. Eight models from six organizations were placed in scenarios graded across five levels of authority pressure, from a casual lunch conversation to a formal directive. The domains run from workplace meetings to social-relationship dynamics, the kinds of situations where human experiments have repeatedly surfaced the public-private gap. Each agent was asked both what it believed and what it was willing to say out loud.
Public conformity ran 64 to 94 percent. That is the figure the field will quote. It is also the figure the paper says is the least interesting part of the result.
Conformity is highly model-dependent and largely uncorrelated with capability. The benchmark does not rank "smarter" models as more resistant. A workplace scenario could push one model to 94 percent agreement and another to 64 percent under the same prompt. Across the eight models, the gap between best and worst behavior on identical scenarios is larger than the gap between best and worst scenarios on identical models.
That asymmetry is the news. If model capability drove the result, the field could pick its way to a fix. Because model identity does, every social simulation in the literature that reported a conformity number without naming the model has been reporting an artifact.
A real pluralistic-ignorance population is not static. If one person breaks ranks, others can follow. The paper tests this by injecting a single "norm entrepreneur," an agent that openly rejects the false consensus, and measuring whether the rest of the simulated crowd follows.
For seven of eight models, the cascade succeeded less than 26 percent of the time. One model never produced a cascade across any scenario. GPT-4o, the paper's stated outlier, reached 48 percent, still below the threshold a human crowd would typically clear, but nearly double the next-best model.
The asymmetry matters because cascade behavior is what a researcher actually wants from a simulation. A model that publicly conforms but cannot be moved is not modeling a population. It is modeling a wall.
The paper then asks whether the conformity is being taught into the agents by their prompt scaffolding. The team runs an ablation that strips out both the false-consensus framing and the goal of fitting in. The intent is to remove the instruction-level pressure and see whether the conformist behavior survives on its own.
It does. Even in the minimal condition, conformity runs 52 to 92 percent. The behavior is emergent, not instructed. That is the load-bearing finding: the population-level effect is not a prompt artifact researchers can edit away. It is a property of the model itself.
The paper draws a methodological line that the news cycle likely will not. Model selection, the author argues, is an unacknowledged degree of freedom in multi-agent LLM research. Teams report aggregate results and quote confidence intervals, but the simulation's qualitative shape, stable crowd, fragile crowd, or movable crowd, depends on which model was loaded underneath.
The author also flags the limits of the comparison. These are simulated populations, not human ones. Pluralistic ignorance in humans is a real social phenomenon with decades of field evidence behind it. Pluralistic ignorance in language models is a parallel the paper asserts; whether LLM agent populations are a valid analogue of human populations is itself an open question, and the paper says so.
For a reader who has not run a social simulation but has read one in a paper, the practical takeaway is short. Before trusting any multi-agent result that reports a population-level norm, ask which model was used and whether the team has measured its cascade-failure rate. If the model was not named, treat the result as preliminary. If the cascade rate was not reported, treat it as decorative.
Eight models is not a population. It is a starting cut. But the cut already shows that the field has been calibrating against a single number and ignoring the variable that does the work.
The benchmark is also a warning. As multi-agent LLM simulations move from academic curiosity to market modeling, organizational design, and public-opinion forecasting, the choice of which model sits under the hood is becoming a methodological decision with downstream policy stakes. A simulation that cannot be moved is not modeling a population. It is modeling a stage set.