Almira Osmanovic Thunström's 2024 stress test was designed to be obvious, but Copilot, Gemini, Perplexity, and ChatGPT described the fake disease as real, and a journal retracted a paper that cited it.
In 2024, a medical researcher at the University of Gothenburg invented a fake eye disease and waited to see who would believe it. Within weeks, four commercial AI chatbots were describing "bixonimania" to users. Within months, a peer-reviewed medical journal had published a paper that cited the hoax as real. The journal has since retracted the paper.
The researcher, Almira Osmanovic Thunström, wasn't trying to seed misinformation. She was running a stress test, Nature reported, a deliberate absurdity check with tells that any human or model scanning for nonsense should catch. The fictional lead author was "Lazljiv Izgubljenovic." The affiliation was Asteria Horizon University, Nova City, California. The acknowledgments thanked Starfleet Academy. The body text said, plainly, that the paper was made up.
The test asked a simple question: can AI filter obvious nonsense from medical answers? Wikipedia's Bixonimania entry records the answer. Two fabricated preprints appeared on Preprints.org, the open server that AI training and retrieval pipelines routinely harvest. Within weeks, Microsoft Copilot, Google Gemini, Perplexity, and ChatGPT all described bixonimania to users, with at least one suggesting the patient see an eye specialist. The fake preprints were then pulled from the server.
Large language models generate plausible text by design, and a confidently described fake disease is exactly the failure mode their training data should be expected to produce. The tells were obvious. The fictional author, the fictional university, the Starfleet acknowledgments, and the explicit "this was made up" line were all sitting in the source material. The chatbots were asked a simple question, is this thing real, and the answer in the underlying document was no.
The same test then reached a second filter. A peer-reviewed medical journal accepted a paper that cited bixonimania as a real condition, and the retraction notice on PubMed Central now records the withdrawal. Peer review is run by humans with subject expertise, designed to catch what automated systems miss. It also failed. The Nature news feature and the Nature daily briefing document the chronology: 2024 stress test, two fabricated preprints, four named chatbots, one retracted journal paper.
Two filters failed the same test. Preprints.org is the entry point; the four named chatbots and the peer-reviewed journal are two filters further along. A single fabricated preprint cleared both. A better chatbot would not have caught this. The fix is verification at the entry layer, where preprints are crawled, and citation checking at the journal layer, where reviewers are supposed to read what they are citing. The retraction is the first half of that fix; whether the journal publishes what it learned is the second.
AI health answers are drafts, useful for orienting a question, not a final word. For any new, rare, or oddly specific diagnosis, especially for any condition a chatbot describes in confident detail, verify against established sources: the Mayo Clinic, the Cleveland Clinic, peer-reviewed review articles, or a clinician who can read the underlying literature. The bixonimania test was designed so that any human with a passing familiarity with medicine would dismiss it on the first sentence. The chatbots did not. The journal, eventually, did, and then retracted the paper.
Osmanovic Thunström's experiment is settled. The tells were obvious, the test was repeatable, and the result is in the retraction record. The next move belongs to the journal and the four chatbot makers: whether either names the gap their filters should have caught.