A peer reviewed study of 1,682 US adults found that AI written short stories rated higher than matched human ones on quality and 'absorbing', and that readers could not tell them apart.
In a controlled experiment with 1,682 US adults, short stories generated by ChatGPT 4.0 outscored human-written stories from a literary journal on both quality and how "absorbing" they were. The same readers, when shown a human story and a matched AI story on the same theme, picked which was which at roughly a coin flip. The result sits inside a 2×2 design that crossed true story origin with the label participants were shown, and the paper's own data carries a finding the wire has skipped: every story in the study scored higher when participants were told a human wrote it, regardless of true origin.
The study, by Sydney Sears and Deena Skolnick Weisberg at Villanova University, was published August 4 in the peer-reviewed journal [Judgment and Decision Making](https://doi.org/10.1017/jdm.2026.10042). Six short stories (three from a literary journal, three generated by ChatGPT 4.0 and theme-matched to the human set) were rated on a -3 to +3 scale. AI-written stories averaged 1.54 on quality against 0.97 for the human ones; on a separate "absorbing" scale, AI stories averaged 1.42 against 1.00. In two follow-up tests with 905 adults, the share who correctly identified the AI story in a paired comparison was 39% in the first round and 52% in the second, a result the researchers describe as "essentially chance" and for which they offer no real explanation.
When participants were told a story was written by a human, they rated it higher regardless of its true origin. The same was true for stories actually written by humans. The "I can spot AI writing" instinct is responding to authorship labels, not to prose craft.
Senior author Weisberg said the finding "reveals a bias towards narratives written by real people." The Guardian quotes her as saying that "AI systems can already generate short stories that are seen as being at least as good as, if not better than, human-written stories," and that "we should update our views of AI's abilities accordingly." Those are the two halves of the finding: blind preference and labeling bias.
The New Scientist quotes a source who argues that models shaped by reinforcement learning from human feedback are explicitly optimized against human preference, producing a "milquetoast everyman" output that "appeals to everybody." Humans, by contrast, remain "weird and idiosyncratic." The styles the models pick up (em dashes, the "it's not just X, it's Y" sentence shape) are the readable AI tells, and they double as the patterns readers rate as pleasing.
A separate preprint from Haverals and colleagues, "Everyone prefers human writers, including AI," reports the opposite bias. In a 14×14 matrix of AI evaluators and creators, humans showed a 13.7-percentage-point bias toward calling a story human-written; AI models showed a 34.3-point bias in the same direction, a 2.5x stronger effect. The Haverals result is not a direct contradiction; it measures attribution bias, not story quality. A working paper from IZA finds a different split: 650 US participants told a story was AI-generated rated it more harshly, but spent the same time and money finishing it, and roughly 40% said afterward they would have paid less for the same story labeled AI. Revealed preference and stated preference diverge.
The Guardian quotes a source who called AI-generated fiction "predictive code based on a massive stolen database of actual writing, physically based in vast, expensive and destructive datacentres with catastrophic repercussions for the surrounding communities." It is, the source said, an "existential threat" rather than a "tool." The score a reader gives a story is not the only thing the story touches, and the numbers do not answer that question.
The study is one US sample: 1,682 adults, six short stories, ChatGPT 4.0, literary fiction. It says nothing about novels, journalism, screenwriting, or other languages. Within those limits, the blind preference for AI prose is real, the labeling preference for human authorship is also real, and the two together describe a reader whose judgment is shaped by what they think they are reading more than by the prose underneath. The next replication round will tell which of those two effects generalizes first.