24 PhD reviewers scored essentially a coin flip on 104 biomedical images, half AI forged, and the best commercial AI detector managed only moderate accuracy — the fix is a third layer: provenance, watermarking, or forensic review.
A diagnostic study published in the Journal of Medical Internet Research put 24 PhD-level reviewers in front of 104 biomedical images, half of them AI-forged to a high-fidelity standard. The reviewers' mean accuracy was 50.5% (SD 6.9%), essentially a coin flip. The best commercial AI detector tested on the same set reached an AUC of 0.790 (95% CI 0.695 to 0.885). That is moderate performance, not decisive. Peer review and off-the-shelf AI detection, the two safeguards research integrity currently leans on, both failed the same task.
The task is narrower than the headline suggests, and the narrowness is the point. The test set was built from two specific image types central to preclinical biology. A western blot is a standard lab assay that visualizes a target protein as a band on a membrane. A subcutaneous xenograft tumor image is a photograph of a human tumor grown under the skin of a mouse, used to test whether a candidate therapy shrinks a real, vascularized cancer. These are the figures a biologist looks at to decide whether a drug works, and they are the figures a journal editor looks at to decide whether the paper is sound. The forgeries were produced with GPT Images 2.0, according to the study's diagnostic protocol, and were designed to be conclusion-oriented: a band dimmed, a tumor shrunk, a result tilted from null to positive. They were not sloppy. That is the threat class the paper isolates.
Legacy image-integrity checks were built for a weaker adversary. Duplication detection and pixel-level forensic review catch copy-paste and splicing, the kind of manipulation that image-checking services and the Retraction Watch office have spent a decade refining. A conclusion-oriented forgery that is statistically and visually indistinguishable from a real figure is a different problem. It does not need to copy a band from a neighbor lane. It needs to produce a new band that is convincing. The 50.5% reviewer accuracy and the 0.790 AUC say the existing tools are calibrated for the first threat, not the second.
The honest size of the finding is worth holding. This is one diagnostic study of 104 images, drawn from two image types, against a single generative model, with no claim that the same numbers would replicate across journals, modalities, or vendors. Independent replication is the open question, and the paper does not foreclose it. The point is not that biomedicine is overrun with AI-forged papers. The point is that the two safeguards a non-specialist assumes are in place were tested on a worst-case, modern forgery and did not catch it. The diagnostic study does not audit the published literature, and that distinction is what keeps the claim falsifiable rather than vague.
The fix the paper actually points at is a third layer. Peer review was the first. Commercial AI detection was the second. The third is image provenance metadata attached at acquisition, cryptographic watermarking for high-stakes modalities, and dedicated forensic review that is paid for and trained for this specific threat class. None of these are science fiction. C2PA-style provenance is shipping in consumer cameras and some lab instruments. Watermarking is a known engineering problem with known protocols. Forensic review as a separate, funded lane is the closest to a policy change rather than a technical one, and it is the one a single institution or journal could deploy this year.
The JMIR diagnostic study, authored by Shuhang Luo, Runhua Tang, Ziyin Chen, Jianye Wang, Ming Liu, Jianfeng Wang, and Li Ma, is the first controlled quantification of the failure mode rather than a literature audit of fraud. The 50.5% reviewer accuracy and the 0.790 detector AUC are the specific, falsifiable numbers. The paper's argument: the two existing safeguards are calibrated for an older threat, and the third layer, including provenance, watermarking, and dedicated forensic review, is the work the field now has to do.