A preprint on film trivia finds that AI access cuts the 'I don't know' rate even when the model is wrong and accuracy is rewarded.
A confident bot doesn't have to be right to change what you say next. It just has to be there. A new preprint from researchers in Italy and France finds that when people are allowed to consult a large language model on trivia, the rate at which they say "I don't know" collapses, even when the model is wrong and participants are paid for accuracy.
The team, led by Valerio Capraro of the University of Milano-Bicocca with Chiara Marcoccia of the École Normale Supérieure and Walter Quattrociocchi of Sapienza University of Rome, ran a between-subjects experiment on visual-detail film trivia, the kind of questions where a confident guess is usually worse than a clean admission. Sample items: what color is the uniform in Bend It Like Beckham, and what vehicle does Monica drive in Like a Cat on a Highway? The full study is posted on arXiv as 2607.13562.
One group answered unaided. The other could consult an LLM. The baseline mattered: without AI, 44% of participants said "I don't know" on items where a wrong model couldn't have helped them anyway. That 44% gives the AI group a yardstick and shows how often people will admit ignorance when the alternative is guessing.
The primary model was Step 3.5 Flash, picked because it is usually wrong on these items. That choice is the experiment's hinge. If the team had used a frontier model that nailed the trivia, the drop in "I don't know" answers could be read as sensible delegation. With a bot that fails the questions, any change in participants' answers is the suppression of uncertainty, not better information. Confidence rose. Accuracy fell.
Frontier models were tested too, including GPT-5.5, Claude Sonnet 4.6, and Gemini 3.5 Flash. They missed the Monica vehicle question and often got other visual details right. The suppression of "I don't know" held in those runs as well, which the authors read as the same mechanism: any plausible-looking answer preempts the reflex to admit ignorance, whether or not the answer is correct. The variable the design isolates is the presence of a chat box, not the quality of its output.
Capraro, in comments to The Register, put the stakes in plain terms: "I don't know" represents recognition of the limits of one's own knowledge, and easy AI answers may interfere with that capacity. A bot in the room doesn't have to be reliable to change the way a person closes a sentence.
The paper is a preprint on arXiv, not yet peer-reviewed. The task domain is film visual trivia, which is narrow: generalizing the result to high-stakes professional work (medical, legal, coding) is an open question. Whether the effect holds with stronger incentives, in-domain experts, or longer workflows is left for later work.
The practical test is portable. The next time a workflow or school policy frames AI use as a neutral productivity add-on, ask what the "I don't know" rate looks like in that workflow and whether the design preserves a clean path to it. If it does not, the study suggests the workflow is also designing for overconfident answers.