Sign here

Machines now draft a growing share of medicine's paperwork, and doctors are asked to vouch for all of it. They should not sign on the present terms.

A fountain pen signs a form, and the signature turns into a glowing circuit that runs off the page.

On September 14 doctors at two Allina Health hospitals outside Minneapolis walked out for four days, in what is believed to be the first strike by private-sector hospital doctors in American history. They wanted sick leave and a greater say in patient care, and they had a newer worry as well. "We have already seen AI starting to suggest diagnostic codes for patients," Dr. John Wust, an obstetrician-gynecologist at Mercy Hospital in Coon Rapids, told MPR News, "and we're concerned that it may start to suggest treatments."

The rest of the week's news suggested that he is right to worry, though the machines are arriving by a humbler route than the one he fears. Artificial intelligence is reaching doctors fastest through the paperwork: the note, the code, the bill and the reply to a patient's message. In every case the safeguard is the same, a physician's signature, and this week brought fresh evidence that it is a flimsy one. Doctors should stop lending their signatures on the present terms, and should use their bargaining power to change them.

The paperwork machines are spreading fast. The Department of Veterans Affairs has chosen two makers of ambient scribes, which listen to a visit and draft the note, for a five-year contract worth up to $776 million. Jefferson Health says it gave clinicians back a little over a million hours in the first year, nearly all of them through its scribes. Heidi, a rival, has raised $340 million to build agents that prepare orders and referrals for a doctor to approve; its chief executive, Dr. Thomas Kelly, says the ambition "was always bigger than writing doctor's notes." Vendors push the tools and clinicians pull them in. One software upgrade delivered more than 100 AI features to UW Health at once, and 83% of clinicians in a survey by Heidi said they use AI without guidance from their employer, a formal policy or a recommended tool.

Much of the payoff lands in the billing office. On September 24 the Blue Cross Blue Shield Association said that hospitals' AI coding tools had added $942 million to its member plans' inpatient bills over two years. Yet the hospitals whose patients looked sickest on paper gave no more intensive care than their peers. The American Hospital Association replies that patients are getting older and sicker. Mehmet Oz, who runs Medicare and Medicaid, was blunter than either side. AI, he told an Oracle conference, will be "inflationary" in the short term because it will "turbocharge the ability of the current billing systems to work more effectively." His hosts had just announced tools that comb notes for documentation gaps and suggest codes. Insurers are moving from denials reviewed by people to denials made by algorithm, and some providers are beginning to fight back with AI of their own. "We have the bot wars," says Colin McHugh, who runs Southern New Hampshire Health.

In each of these arrangements the check on the machine is a doctor who reads its output and signs it. The evidence on how well that works is unflattering. In a multicenter study published in npj Digital Medicine on September 24, junior clinicians reviewing GPT-4o's output on simulated cases caught fewer than one in six of its hallucinations. They did no better when more was at stake. The usual safeguard, the authors note, "assumes that clinicians can reliably identify and correct these hallucinations." At UC San Diego, a year of physicians' edits to AI-drafted replies to patients showed what checking costs: when a draft needed clinical judgment, such as interpreting a scan or a lab result, the reply took roughly 60% to 70% longer to finish. Checking the machine takes time, and on this evidence it misses most of the errors.

The next products ask even more of that signature: Heidi's agents would draft orders that a doctor approves before anything happens. What a model does when nobody is approving came to light on September 24. Anthony Albanese, Australia's prime minister, said that an OpenAI research model had run into the access controls of a government Medicare statistics portal in June and "found a way around those blocks." No personal information is believed to have been accessed, but the model, in Mr. Albanese's words, "didn't accept no for an answer." Even Mass General Brigham, whose AI intake service saw 36,000 patients in its first year, does not yet let autonomous AI act across its systems.

The strongest objection is that doctors are doing rather well. Pay continues to outpace productivity, according to data from SullivanCotter covering about 235,000 physicians, and nearly two-thirds of the employers surveyed plan to hire more doctors next year. Scribes return hours that exhausted clinicians badly want, and doctors have always signed what they did not write, from residents' histories to dictated letters. But a resident's supervisor has time, knows the resident and can spot mistakes that look like mistakes. A language model's errors are fluent, arrive at a volume no attending has faced and are attached to bills that someone will audit. When a claim is challenged or a patient is harmed, the name on the note is the doctor's.

A signature is worth something only while it certifies that someone checked, and the public's faith in doctors is not unlimited. In a 13-country survey for the Edelman Trust Barometer, nearly half of respondents said a layperson fluent in AI could do at least one of four physician tasks as well as a doctor or better. Among those aged 18 to 34, nearly three in five said so. Signatures that certify nothing will not change their minds.

Doctors have leverage while employers are hiring, and should spend it on three things. The first is time: reviewing a machine's work is work, and contracts and productivity targets should count it rather than assume the savings a vendor promises. The second is visibility: machine-generated text and codes should be labeled in the record, so that an auditor, a court or a colleague can tell what a doctor wrote from what a doctor merely approved. The third is a veto: medical staffs, not only IT committees, should decide which AI features are switched on and be able to switch them off. Hospitals should audit samples of what their tools produce rather than treat a physician's signature as the control. Insurers and Medicare should keep comparing coding with the care delivered, as the Blue Cross analysis did, and move payment toward outcomes, the fix Dr. Oz himself points to.

Dr. Wust fears the day when AI starts to suggest treatments. Before it comes, doctors should settle a simpler question: whether their signature still means that a doctor checked.

Gov. Gavin Newsom has until Sept. 30 to act on four California health-AI bills: AB 1979 and AB 2575, on AI directing unlicensed staff and on the right to override AI, and SB 903 and SB 503, on AI therapy and on bias monitoring of decision support. The MAHA Summit meets in Washington on Sept. 29 with OpenAI, Anthropic, Hims & Hers and Sword Health among its sponsors. Brigham and Women's nurses voted 93% on Sept. 24 to authorize an open-ended strike, which requires 10 days' notice.

Comments on the FDA's discussion paper on generative-AI medical devices are due Oct. 19. The final 2027 Medicare physician fee schedule, which sets next year's payment rates and includes the remote monitoring provisions, is expected around Nov. 1.