A phrase is circulating in medical education circles this year that deserves more attention: never-skilling. Not deskilling — that implies a doctor who once reasoned clearly and has since gone soft. Never-skilling is worse. It describes a trainee who reaches attending level without ever sitting with real diagnostic uncertainty long enough to resolve it alone. Deskilling is recoverable. Never-skilling might not be.
OpenEvidence, an AI system built specifically for clinicians, is forcing this distinction into the open. Doctors at hospitals like South Shore Health and Northwestern Medicine now pull it up mid-shift to check a drug interaction or confirm a lab threshold. The answer lands in seconds, citations attached, tone indistinguishable from a seasoned colleague. For a physician with fifteen years of pattern recognition already banked, that’s a real productivity gain. For a third-year medical student who hasn’t built that pattern recognition yet, it’s closer to a shortcut around the one exercise meant to build it.
The exercise that used to hurt
Here’s what clinical training looked like before these tools existed. A student gets a case, generates a differential diagnosis, and misses something. The attending points it out. It stings — and that sting is the mechanism. Discomfort encodes the list into memory in a way that survives the next similar case, three months later, when nobody’s there to fill the gap.
Today’s student has a faster option. They open the app, get the complete list back in seconds, and walk onto the wards sounding like someone who already knew it. No missed diagnosis, no correction, no sting — and, less obviously, no encoding. The performance of competence and the acquisition of competence have quietly come apart. Nothing in the interaction warns the student this happened.
Why “just use it responsibly” won’t work
It’s tempting to treat this as a discipline problem: tell trainees to lean on AI less, expect self-restraint to close the gap. That framing misreads the incentive structure trainees actually live inside. Most already sense the risk. What keeps them reaching for the tool is closer to a coordination failure than a willpower failure. If every other resident on the team uses AI to arrive at rounds looking sharp, the one person who reasons it out manually looks — and feels — behind. Nobody wants to be the single trainee who skipped the fastest path to sounding prepared.
This isn’t unique to medicine. The same pattern is showing up across wellness culture right now, where people hand a chatbot the judgment call instead of building it themselves. The same substitution shows up across knowledge work broadly: reasoning that gets outsourced consistently enough doesn’t stay dormant, it weakens. Medicine just raises the stakes, and physician AI use has roughly doubled since 2023, with the adoption curve still climbing. That’s exactly why a “use it wisely” message can’t fix this. Only a structural change in the training loop can.
The fix already being tested: reason first, ask second
The most concrete proposal doesn’t ask trainees to avoid AI. It asks them to sequence it. Before consulting any tool, a trainee writes a brief commitment: a leading diagnosis, the dangerous alternatives to rule out, and the next step they’d take. Only after that unaided first pass does the AI get consulted — now a check on the reasoning, not a substitute for it. A resident admitting a patient overnight might write a short “pre-AI assessment” right after the history and physical. When new information changes the picture mid-case, the team pauses to reason through the shift before anyone opens the app.
This kind of imposed friction has a name in learning science: a desirable difficulty. UCLA researchers Elizabeth and Robert Bjork coined the term for a step that slows performance in the moment but strengthens what actually sticks. Used this way, AI stops being a shortcut and becomes something closer to a tutor — a tool that shows a trainee, after their own attempt, exactly what they missed and where they overweighted a familiar pattern.
Borrowing a habit from aviation, and going one step further
Medicine isn’t the first high-stakes field to face this. Aviation solved a similar problem decades ago, and the answer wasn’t banning autopilot. Regulators required pilots to keep hand-flying, specifically so the skill wouldn’t atrophy behind the automation. Medical training is starting to borrow that logic: scheduled no-AI cases, graded purely on unaided reasoning, designed to surface drift before it becomes a blind spot.
But medicine may need a second layer aviation doesn’t. It needs to train doctors to actively interrogate the machine, not just periodically work without it. One promising version looks like a simulator drill built from real cases: an AI-generated assessment with a subtle error slipped in. The debrief isn’t just about whether the trainee caught the mistake. It covers when they started trusting the output, when suspicion should have kicked in, and what tipped them off. Correct AI answers get mixed in too, so the goal isn’t reflexive distrust. It’s calibrated judgment — knowing when to lean on the tool and when to override it.
What’s actually being protected here
None of this is nostalgia for a harder, more punishing version of medical school. The struggle being defended isn’t hazing. It’s the specific cognitive work of learning to notice when a familiar pattern shouldn’t be trusted: the pneumonia that turns out to be heart failure, the heart failure that first looked exactly like pneumonia. That instinct doesn’t come from reading a correct answer. It comes from being wrong, alone, often enough to recognize the shape of your own blind spots.
AI can reason cleanly over the facts it’s handed. It can’t tell you which facts were missing from the chart, or which detail the patient mentioned in passing that actually mattered. Physicians still earn their place in that gap — but only if training keeps producing people who can stand in it. Right now, the tools are moving faster than the training loop built to absorb them. This year’s medical schools are quietly wrestling with a real question. It isn’t whether AI belongs in the room. It’s whether anyone will still know how to catch it when it’s confidently, fluently wrong.
Related: AI Face Is Changing Plastic Surgery — But Can Surgeons Actually Create It?
