OpenEvidence and peer clinical AI chatbots now sit beside most US physicians at the point of care. The deeper risk for medical students is not losing clinical judgment but never building it.
A medical student sits down with a case at 2 a.m. She is asked to list every plausible diagnosis, and the act of trying, of getting most of them, missing one, and being told about the miss, is the training. OpenEvidence, the clinical AI chatbot now used by roughly two-thirds of US doctors, hands her a near-complete answer in seconds, and the gap the training was designed to surface never opens.
Forty years of cognitive research converge on the same lesson: skills are built by the struggle to recall them under uncertainty, not by exposure to correct answers. Robert and Elizabeth Bjork call this the principle of desirable difficulties, and they are specific that the difficulty has to be productive and felt in the moment, which is what a missing diagnosis is and what a confident chatbot answer is not. (Bjork & Bjork 2011)
A recent Nature Medicine study reports that general-purpose large language models now outperform specialized clinical AI tools on standard medical benchmarks, which is why trainees can lean on them. (Nature Medicine) According to company figures reported by NBC News, the chatbot was used by roughly 650,000 US physicians across 27 million clinical encounters in April 2026, and outside coverage of the same data puts the share of US doctors actively using it near two-thirds. (NBC News) (ai2.work) The scale is large enough that the first cohort of medical students who trained with it as a default tool is now entering residency.
This is where Simar Bajaj and Joseph Sakran, in a Guardian op-ed, introduce a distinction the rest of the coverage tends to skip. "Deskilling" is the familiar problem: a pilot who stops flying manual approaches loses a skill that existed and can be rebuilt with practice. "Never-skilling" is the one they coin: a trainee who never builds the skill in the first place, because the formative window when it was supposed to form was filled by a tool that already had the answer. The aviation analogy is borrowed from an existing regulatory document. The FAA's SAFO 17007 explicitly warns that pilots who rely on automation can lose manual flight skills, and the comparison holds because it is the same shape of risk applied to a different cockpit. (FAA SAFO 17007)
The op-ed names a problem; the question is what training programs are doing about it. Two accreditation bodies are now moving. The Accreditation Council for Graduate Medical Education has begun updating its Common Program Requirements to ask programs how they supervise AI use in formative clinical reasoning, and the Liaison Committee on Medical Education is signalling parallel changes in pre-clerkship curricula. The redesigns share a shape: AI is allowed at the bedside, sometimes required for documentation, but the moment a trainee is being assessed on judgment, the screen is set aside and the question is asked again out loud, by hand, with a faculty member in the room. Faculty are being asked to do something they did not have to do a generation ago, which is to be the deliberately slower alternative to the tool the trainee already trusts.
No one has measured whether this protects the skill. There is no published data on board scores or patient outcomes for residents who trained with default AI access versus those who did not, because the first generation of those residents has not yet taken the boards. The next eighteen months are the window in which that data will start to arrive, and the formative rules written in 2026 are the ones it will test. The counterargument is real and worth naming: this generation of doctors may simply have a different skill profile, not a worse one, with the chatbot doing what UpToDate and pocket references used to do. The falsifier is whether board scores and patient outcomes hold steady. They have not been measured yet, and the policy window is open precisely because they have not.
The next doctor a reader sees is being formed right now, in residency programs and medical schools, in the small interaction where a chatbot offered the answer before the doctor-in-training had to reach for it. The rules written this year are the ones that decide what that doctor actually knows.