Keep your hand in

AI is taking over the everyday work that kept doctors' skills sharp. The skills it cannot replace will now have to be practiced on purpose.

A metronome whose pendulum arm is an amber reflex hammer, swinging against a blue background.

When endoscopy centers in Poland began using an AI that flags polyps, their experienced endoscopists got worse at finding precancerous growths on their own. In the three months after the software arrived, the share of patients in whom they found an adenoma during colonoscopies done without it fell from 28 percent to 22 percent, compared with the three months before. The study was observational and cannot prove that the machine caused the decline. It is still the plainest warning yet of a problem aviation met long ago: a skill handed over to a machine begins to fade.

Most of the argument about AI and doctors concerns which jobs it will take. For a practicing physician the more useful question is which skills to keep, and how. The skills worth having are the ones machines cannot supply, together with the competence to catch a machine when it is wrong. All of them decay without use, and AI is quietly taking away much of the everyday use that kept them sharp. Doctors should practice them on purpose, and employers should give them the time.

The line between what machines can and cannot do in medicine runs through the hands. Anthropic's economic index finds that AI covers few of a radiologist's listed tasks "because AI can't do the hands-on or administrative work in their job profiles," even though it does well at the reading and reporting that fill the working day. A company that has logged a million virtual hospital medicine visits says its remote hospitalists handle the full range of inpatient duties except hands-on procedures such as central lines and intubation, and robots that intubate are still mostly in testing. On the screen side of the line the machines are formidable: in a trial published in March, an AI working alone was right on almost nine in ten written cases, against three in four for physicians' own answers using conventional resources.

What the machines leave to doctors comes down to four skills. The first is the hands: the examination, the procedure, the ultrasound probe. The second is presence. Evaluators preferred a chatbot's written answers to medical questions posted online over doctors' 79 percent of the time and rated them empathetic far more often, so the durable skill lies in the room: sitting with a frightened patient, delivering bad news and being trusted with the decision that follows. The third is judgment that someone must answer for. The federation of America's state medical boards holds that "the physician is ultimately responsible for the use of AI," and that failing to apply human judgment to its output breaches a doctor's professional duties. The fourth is the one most often forgotten: the ability to tell when the machine is wrong.

That last skill depends on a doctor's own competence. In a multicenter study published in September, junior clinicians reviewing GPT-4o's output on simulated cases caught fewer than one in six of its hallucinations; the authors noted that the "clinician-in-the-loop" safeguard assumes "that clinicians can reliably identify and correct these hallucinations." When researchers planted wrong suggestions in software presented as AI, inexperienced radiologists reading mammograms chose the right category less than one time in five, against about four times in five when the suggestion was right, and even veterans were misled, though less often. Doctors can check a machine only with skills of their own, and trainees who lean on one from the start risk what educators have begun to call never-skilling: not acquiring those skills at all.

Skills like these fade without use. The American Heart Association's guidelines note that CPR skills "often show decay by as early as 3 months" after training. Volume seems to matter in the same way: in a large observational study of elderly patients in American hospitals, those treated by older hospitalists were more likely to die within 30 days, except when the older doctor saw a high volume of patients. Nor can training be relied on to supply the skills in the first place: in a multi-site survey of internal-medicine residents, only 30 percent felt confident doing bedside procedures. AI caused none of this, but it adds to it, because every task a machine takes over is practice a doctor no longer gets.

The strongest objection is that this is nostalgia. Nobody asks an accountant to do long division, and if machines are right more often than doctors, practicing what they do better looks like a waste of scarce time. The objection fails on two counts. No machine is close to taking over the examination or the procedure. And for the rest, machines fail in the moments that matter most, on the rare case, the confident error and the system that goes down, which is exactly when a doctor must take over and when a doctor out of practice is least able to. Pilots were warned of exactly this in 2013, when America's air-safety regulator told airlines that "continuous use of autoflight systems could lead to degradation of the pilot's ability to quickly recover the aircraft from an undesired state," and encouraged them to promote manual flying when appropriate.

Building the skills takes practice with feedback, and the evidence that practice pays is good. When residents trained on a simulator before placing central lines in patients, bloodstream infections in their intensive-care unit fell by more than four-fifths. Skill is also visible to others: when peers rated videos of surgeons at work, those judged least skilled had more complications and higher mortality, which is an argument for being watched and coached. Even the hard conversation can be taught: in a randomized trial, oncology clinicians given training, a structured guide and reminders for serious-illness conversations documented them for more patients, earlier, and more often recorded what mattered to the patient.

Keeping the skills is mostly a matter of dose and timing. The heart association recommends following a course with brief booster sessions, weekly or monthly, rather than relying on the course alone, and cites a trial in which nurses given more frequent boosters had better CPR skills a year later. The same logic should apply well beyond resuscitation. Doctors should set a yearly floor for each procedure they mean to keep, log every case and teach the skill to others. They should take regular stretches without AI for the reads and diagnoses they want to own, as pilots fly by hand. And they should commit to their own assessment before looking at a model's: in the March trial, physicians who did so were about as accurate as those shown the AI's answer first, and had done the thinking themselves.

Employers have the bigger job, and many are heading the wrong way: a fifth of physicians told Doximity they already face higher expectations for productivity because of AI. Hospitals should spend some of the time AI saves on practice instead, through protected hours, simulation, booster sessions and procedures shared among doctors rather than routed to a single service. Medical staffs at accredited hospitals renew privileges every two or three years and must review each doctor's performance data at least once a year in between; both reviews should ask for evidence that doctors still do what their privileges allow. Residency programs should have trainees commit to a diagnosis before they consult a model, and test them without one.

The Polish endoscopists had each done more than 2,000 colonoscopies before the study began. A few months of letting software share the looking seems to have been enough to blunt what those years had built. Every doctor who now works beside a capable machine faces the same slow drift, and the remedy is the one pilots were given: keep your hand in, deliberately and often, so that the skill is still there on the day the machine is wrong.

  • In an observational study at four endoscopy centers in Poland, 19 endoscopists who had each performed more than 2,000 colonoscopies had an adenoma detection rate of 28.4 percent (226 of 795) in standard colonoscopies during the three months before an AI polyp-detection system was introduced, and 22.4 percent (145 of 648) in standard colonoscopies during the three months after, an absolute difference of 6.0 percentage points. The rate in AI-assisted colonoscopies was 25.3 percent. The authors concluded that continuous exposure to AI "might reduce" the detection rate of standard colonoscopy. The study was published in The Lancet Gastroenterology & Hepatology.
  • Junior clinicians reviewing GPT-4o output in simulated clinical scenarios identified 15.8 percent of its hallucinations, and 13.1 percent of the clinicians identified none, according to a multicenter cross-sectional study published September 24 in npj Digital Medicine. Detection did not improve as clinical risk rose, and most of the variability came from differences between clinicians rather than between scenarios. The authors called for structured human-AI workflows and tiered clinical certification pathways.
  • In a study of 27 radiologists reading mammograms with BI-RADS categories presented as AI suggestions, published in Radiology in 2023, inexperienced readers chose the correct category almost 80 percent of the time when the suggestion was correct and less than 20 percent of the time when it was incorrect. Very experienced readers, with more than 15 years of experience on average, fell from 82 percent to 45.5 percent, according to the Radiological Society of North America.
    Sources: RSNA, Radiology
  • In a randomized trial of 70 clinicians (39 residents and 31 attendings) published March 18 in npj Digital Medicine, diagnostic accuracy on clinical vignettes was 85 percent when the AI's opinion came first and 82 percent when clinicians gave their own assessment first, against 75 percent for the initial answers that clinicians in the second-opinion arm gave using conventional resources; in both workflows an AI-generated synthesis then integrated the two views. The AI alone scored 87 percent in the results (the abstract gives 90 percent). The authors reported that performance was comparable across workflows and not statistically different from the AI alone.
  • The American Heart Association's 2025 guidelines on resuscitation education say CPR skills acquired in basic life support training "often show decay by as early as 3 months." They recommend booster sessions when training uses a massed approach, such as a single intensive course (Class 1), say a spaced approach is reasonable in its place (Class 2a), and describe booster training as brief weekly or monthly sessions, citing a randomized trial in which nurses given more frequent booster training showed a dose-dependent improvement in CPR skills at one year.
  • In an observational study of 736,537 admissions of elderly patients managed by 18,854 hospitalists, published in The BMJ in 2017, adjusted 30-day mortality was 10.8 percent for patients of hospitalists younger than 40, 11.1 percent for those aged 40 to 49, 11.3 percent for 50 to 59 and 12.1 percent for 60 and older. Among physicians with a high volume of patients, there was no association between physician age and patient mortality.
    Source: The BMJ
  • In a study published in Archives of Internal Medicine in 2009, catheter-related bloodstream infections in a medical intensive care unit were 0.50 per 1,000 catheter-days after simulator-trained residents entered the unit, against 3.20 per 1,000 catheter-days in the same unit before the intervention. The reduction came even though proven preventive strategies were already in place, according to AHRQ's Patient Safety Network.
  • Doximity's 2026 physician compensation report drew on more than 250,000 survey responses collected over seven years, including nearly 23,000 U.S. physicians surveyed last year. It found that more than 65 percent of physicians use AI daily or weekly for clinical or administrative work, that a fifth face higher expectations for productivity because of it and that 42 percent are concerned they will in the future. More than 40 percent said doctors should capture most of the savings if AI helps them work faster, Healthcare Dive reported.