Evidence
23 items from the daily postsWhat the studies and the professional societies say, and what nobody has shown yet. Each item links to the day it ran and to its original source.
- Evidence Sept 26, 2026University of Oxford researchers distilled a transformer neural network into SIMPLE-HF, an 11-variable score built from age, body mass index, comorbidities and medications, using UK primary care records for 373,389 adults with heart failure, according to the paper published Sept. 26 in npj Digital Medicine. For 12-month all-cause mortality the score reached a C-index of 0.801, compared with 0.735 for an electronic record version of the MAGGIC score, the authors reported. At a set risk threshold it identified 177 true events per 1,000 people screened, compared with 89 for MAGGIC. The score uses point-of-care variables and does not require echocardiography, the authors said, and they called for external validation and prospective clinical evaluation before use.
- Evidence Sept 29, 2026Researchers in the University of Pennsylvania's radiology department described Percival, a model trained on paired CT images and radiology reports, in npj Digital Medicine on Sept. 29. The model learned from more than 400,000 CT-report pairs and was evaluated on more than 20,000 held-out participants, the authors reported. Its internal representations tracked demographic, physiological and laboratory differences, and training on reports as well as images captured clinical information that image-only comparison models did not, in disease classification and longitudinal risk assessment, according to the paper.
- Evidence Sept 29, 2026Kobe University urologists analyzed instrument motion recorded by the hinotori surgical robot during the bladder-urethra reconnection step of 98 robot-assisted radical prostatectomies by 16 surgeons, according to a study published Sept. 29 in npj Digital Medicine. A random forest classifier distinguished certified proctors from other surgeons with an area under the curve of 0.892, and a three-feature model generalized to surgeons it had not seen, the authors reported. Proctors finished the step in 45.5 minutes on average, compared with 67.8 minutes. The authors described the work as exploratory, citing a single platform, a single surgical step, a 1-hertz sampling rate and only five anastomotic leaks for outcome analysis.
- Evidence Sept 26, 2026University of Alabama at Birmingham researchers reviewed health record data for 5,280 patients with hypertension and 88 primary care clinicians at a large academic health system in the Deep South, according to a paper published Sept. 26 in npj Digital Medicine. Remote patient monitoring reached 5.2% of eligible patients, and 62.5% of primary care clinicians made at least one referral, the authors reported. Referrals were concentrated among a small number of clinicians, and the authors said more research is needed on how clinicians and patients decide to use the service.
- Evidence Sept 24, 2026Ipsos surveyed 254 U.S. patients and 355 clinicians, 203 physicians and 152 nurses, for Wolters Kluwer Health from March 11 to 14, the company said in a Sept. 24 release. Among patients, 54% used AI to research medication side effects, 52% used it to look up a diagnosis and 26% said they sought medical care sooner after AI research. Some 74% worried about the privacy of their health information, including 83% of rural patients, 89% called human expert validation of AI answers important and 81% expected their health system to formally approve AI tools used in clinical decisions. Confidence in AI reliability was 81% among men and 66% among women.
- Evidence Sept 26, 2026Researchers in Shanghai randomized 268 men with newly diagnosed prostate cancer scheduled for radical prostatectomy; 138 received personalized answers from a locally deployed large language model before the routine conversation with their physician, who did not know which patients had received them. Mean GAD-7 anxiety scores were 3.2 in the AI group and 5.7 in controls, and physician workload on the NASA-TLX scale was 39.9 compared with 56.8, according to the phase 2 trial in npj Digital Medicine. The authors also reported better patient satisfaction and illness perception in the AI group. The paper appeared Sept. 26 as an early online version.
- Evidence Sept 25, 2026LEME, a suite of open-weight models tuned on 211,149 curated examples and 29,747 preference-labeled examples, outperformed all seven baseline models tested, according to the paper in npj Digital Medicine by researchers from Yale, the National University of Singapore, Massachusetts Eye and Ear and Stanford. The models extracted visual acuity from notes 14.1% better than Llama 3 70B, and attending ophthalmologists rated their answers to patient questions above expert answers on completeness, the authors said. The models, datasets and code are public.
- Evidence Sept 25, 2026Artera said a commercial analysis of about 20,000 patients showed its multimodal AI risk estimates tracked observed rates of metastasis and prostate cancer death, and that analyses of NRG/RTOG trial data found no algorithmic bias across racial and age subgroups. Other abstracts from Australia, India and U.S. cooperative-group trials examined using the score to guide androgen deprivation therapy and whole-pelvis radiation in intermediate- and high-risk disease, the company said in a Sept. 25 release for the American Society for Radiation Oncology annual meeting. The figures come from company-announced meeting abstracts.
- Evidence Sept 24, 2026A model that combines deep learning features from multiphase CT with clinical variables predicted organ failure in acute pancreatitis with an area under the curve of 0.89 in validation and 0.81 in an external test set, compared with 0.68 to 0.74 for the modified CT Severity Index and 0.67 to 0.71 for clinical models, researchers reported in the Journal of Pancreatology. The study covered 2,746 patients treated from 2011 to 2024. The model's negative predictive value was 97.2%, and it flagged 55% of cases at least three hours before organ failure was clinically evident, with a median lead of 3.5 hours, according to Radiology Business. The analysis was retrospective.
- Evidence Sept 24, 2026The cross-sectional study, published Sept. 24 in npj Digital Medicine by Yunyun Zhang and colleagues, asked junior clinicians at several centers to review GPT-4o output in simulated clinical scenarios. Detection did not improve as the clinical risk of a scenario rose, and most of the variation came from differences between clinicians rather than between scenarios, the authors said. They wrote that current safeguards rely on a "clinician-in-the-loop" model that "assumes that clinicians can reliably identify and correct these hallucinations," and called for structured human-AI workflows and tiered clinical certification. The study used simulated cases, not live patient care.
- Evidence Sept 22, 2026Xu and colleagues at Zhongshan Hospital in Shanghai reported in npj Digital Medicine on FARUSS, a six-axis robotic arm that positions the probe and scans the thyroid on its own using deep learning, feeding an AI platform that detects nodules and assigns TI-RADS categories. The team tested it prospectively against conventional scans in 262 participants at three Chinese institutions from March 2024 to August 2025: intraclass correlations for thyroid measurements were 0.76 to 0.80, nodule identification matched 85.6% of the time, and TI-RADS agreement after a sonologist revised the AI's call reached kappa 0.75 to 0.88. The robot was slower to scan, 202 seconds against 147, and the AI read faster than a human, 183 seconds against 254. In the authors' triage scheme, the system avoided 74.8% of on-site examinations that turned out not to be needed and missed 19.2% of the fine-needle aspiration recommendations a human would have made. Three authors work for the technology companies involved, and the study was funded by the Chinese government.
- Evidence Sept 22, 2026Zhou, Guo, Wang and colleagues trained a two-stage model, one stage to locate the esophagus and one to segment lesions and score malignancy, on 6,813 patients from two centers, then validated it on 80,612 patients across 12 centers in China, the Czech Republic and Australia. On an external test set of 11,466 patients from eight centers, the model found esophageal cancer with 90.0% sensitivity and 98.5% specificity and high-grade precancer with 52.5% sensitivity; on 1,607 low-dose chest CTs it ran at 88.4% sensitivity and 99.0% specificity. In a reader study it outperformed all 17 radiologists, exceeding their average by 19.1 points of sensitivity and 18.4 points of specificity, and as an assistant it raised the radiologists' cancer sensitivity from 71.9% to 85.7%. Run prospectively on hospital scans from January to April 2025, followed through July 2026, and on 10,959 consecutive asymptomatic lung screening participants, the model's flags had a positive predictive value of 42.2%, with specificity of 99.94% in the screening cohort. The authors listed the study's limits: nearly all patients were Chinese, a population in which squamous cell cancer dominates, with only preliminary validation abroad; performance was worse in women because the training set skewed male; follow-up was less than two years with imperfect compliance; and the reader study lacked a full crossover design. In the U.S. the common tumor is adenocarcinoma of the distal esophagus, which the model has not been shown to detect. The paper appeared one day after Alibaba's abdominal CT model was released free of charge.From the Sept 22, 2026 post · Source: Nature Medicine
- Evidence Sept 16, 2026The 49% figure comes from a special analysis of the 2026 Edelman Trust Barometer written with the Yale School of Public Health and released this month, which Becker's Hospital Review reported on Monday. Edelman surveyed 12,998 people in 13 countries, about 1,000 per country including 995 Americans, between Feb. 28 and March 11. Asked which tasks a person with no medical training but skilled with AI could do as well as a trained professional, 26% chose deciding whether someone needs care, 19% chose determining treatment or medication, 19% chose performing basic procedures and 16% chose diagnosing illness. Agreement ran 59% among ages 18 to 34, 52% among ages 35 to 54 and 35% among those over 55, and 56% among the university-educated against 42% among those without a university education. The same analysis found that only 51% of people who receive conflicting recommendations always follow their doctor's. Edelman is a public relations firm and the barometer is its product; the analysis was written with Yale and its method is published. The question measured what respondents believe an AI-skilled layperson could do, not what such a person can do.
- Evidence Sept 18, 2026The model was trained on more than 400,000 contrast-enhanced abdominal CT exams paired with their reports, about 15 million anatomy-tagged image-text pairs, according to the GitHub page, which links to the Science paper. The South China Morning Post reported Thursday that the model was tested on nearly 40,000 real-world exams and recorded a mean AUC of 0.913 across 146 findings in 18 organs, cancers included. The test set came from eight outside centers, the model outperformed 23 of 26 radiologists, and radiologists using it as an assistant gained about 10 points of sensitivity and read more than 30% faster, TechTimes reported; the code is released under the Apache 2.0 license and the weights are licensed for research, not commercial use. The Science paper is behind a paywall, and the figures here come from the company's page and the press reports. The model was validated on Chinese patients only and on contrast-enhanced abdominal CT only, with no prospective trial and no regulatory clearance in any country. Under a final order published Thursday in the Federal Register, every radiology detection, diagnosis and triage algorithm in the U.S. still requires its own 510(k) clearance from the FDA, open weights or not.
- Evidence Sept 19, 2026Emergency physicians and a bioinformatician at UT Southwestern built pairs of patients from two academic emergency departments, matched on Emergency Severity Index, in which one patient deteriorated within six hours and the other did not, then asked Gemma, Qwen and DeepSeek which to see first, before and after inserting stigmatizing or neutral versions of 18 attributes. Stigmatizing language about frequent emergency department use pushed the deteriorating patient down the list in every model and dataset, by as much as 9.4%, and neutral phrasing scored 1.4 to 7.1 points higher, according to the paper in npj Digital Medicine. Psychiatric history moved rankings by up to 9.5% but reached significance in only one model, and race, language and insurance showed no consistent harm. The authors concluded that the wording of the input is a safety-critical design choice. Thursday's post covered a model that scored identical presentations as less severe when the patient was female.
- Evidence Sept 21, 2026Radiologists at St. Jude reviewed every chest radiograph foundation model published through March 2025, 41 in all, and scored each paper against CLAIM, the reporting checklist for medical imaging AI. The brief communication in npj Digital Medicine found transformer architectures and language-model components becoming dominant, model sharing still uncommon and one variable that tracked with quality: publications with a physician among the authors adhered to CLAIM significantly better and discussed fairness more often. The paper is short, and its authors report the link as a correlation only.
- Evidence Sept 17, 2026The study, published in NEJM AI, examined a year of Epic-integrated draft replies to patient portal messages across specialties at UC San Diego Health. Researchers sorted every physician edit into a 15-category taxonomy built with language models and expert review and used response time as the measure of workload. Scheduling changes were the most common edit, appearing in 38.5% of edited messages, followed by lifestyle and non-drug advice in 18.2% and added empathy in 16.1%; stopping or tapering a medication was the rarest edit at 2.4%. Clinical edits added the most time per message: interpreting a radiology result added 70.1% to the time spent on a message, clarifying or ruling out a diagnosis added 63.9% and interpreting laboratory results added 60.8%. The authors concluded that the burden is not uniform, with judgment edits costing the most per message and high-volume administrative edits costing the most across the system, according to reports in Becker's Physician Leadership and Healthcare Innovation; the paper is behind NEJM AI's paywall, and the figures come from those two reports.
- Evidence Sept 11, 2026Kohn and colleagues ran Epic's one-year mortality model, which the paper describes as "perhaps the most widely available commercial mortality risk model," on 154,063 encounters at Trinity Health and 133,043 at Kaiser Permanente Southern California from 2022 through 2023. Discrimination was moderate to high, with a C statistic of 0.76 at Trinity and 0.81 at Kaiser, but calibration was poor in both systems, and the scaled Brier score at Trinity was essentially zero, meaning the probabilities the model reported there were about as useful as the base rate. Performance was weakest in patients 75 and older, where the C statistic fell to 0.69 and 0.73, and it degraded in liver disease, kidney disease and heart failure. Dr. James Deardorff, a UCSF geriatrician, wrote in STAT on Friday that the same prediction is harmless when it prompts a goals-of-care conversation and dangerous when it feeds a transplant or resource decision. Many hospitals have the score switched on.
- Evidence Sept 16, 2026Guerra-Adames and colleagues at Bordeaux University Hospital trained a Mistral model to reproduce nurses' triage scores, then fed it pairs of cases that differed only in the patient's sex, they reported in npj Digital Medicine. Female versions came out 1.1% lower in predicted severity in Bordeaux and 2.2% lower on the U.S. MIMIC-IV records; retraining on sex-neutralized text made the gap disappear, and the pattern shifted with the sex of the triage nurse. The authors described the result as a probe of what the charts record rather than proof of bedside undertriage, and as a hypothesis generator rather than a verdict.
- Evidence Sept 17, 2026A group from the Harvard T.H. Chan School of Public Health and the University of Pennsylvania published the rapid review in npj Digital Medicine on Sept. 17, covering 229 randomized trials from 29 systematic reviews in nutrition, maternal health, mental health and sleep from 2013 to 2025. Developers took part in 73% of the trials. Developer-involved trials had 2.47 times the odds of being preregistered and, when weighted by sample size, higher odds of reporting a statistically significant result, an odds ratio of 1.23 with a confidence interval of 1.16 to 1.31. The authors called for independent efficacy trials, more transparent reporting and regulatory oversight. The domains studied were behavioral rather than diagnostic; the review did not cover ambient scribes or diagnostic agents.
- Evidence Sept 15, 2026Kather's group tested a fully autonomous agent that interviews the patient, examines, orders tests and diagnoses on 551 MIMIC-IV cases across seven acute conditions (appendicitis, cholecystitis, diverticulitis, pancreatitis, pneumonia, pulmonary embolism and urinary infection), 2,400 abdominal cases and 990 published multispecialty cases. Open-weight models hosted on premises (Qwen, GLM and GPT-OSS), which keep patient data inside the hospital, reached 90.0% on the seven-disease task and 83.8% on the four-disease task, near a cloud GPT-5.2 baseline, the authors said. Agreement across five repeated runs separated right from wrong answers with an AUC of 0.86, and at a consistency threshold of 0.90 the agent retained 49.4% of cases at 98.9% accuracy; blinded physician review of 181 cases agreed with the automated grading 92.3% of the time. The authors listed the limitations: retrospective simulation only, text only, one institution's data, lower accuracy in older patients and roughly five times the compute. Under that design, the remaining cases, those on which the agent's repeated runs disagreed, are routed to a person.From the Sept 16, 2026 post · Source: Nature Medicine
- Evidence Jun 17, 2026Google's AMIE trial in Nature was a randomized, blinded virtual OSCE with 100 multi-visit cases and 21 primary care physicians; the model's management plans were rated appropriate 95% to 98% of the time versus 72% to 81% for the physicians. The same paper describes the system as "not ready for real-world translation," citing a text-only format with patient actors and no pharmacist or order entry. In April, Science published a Harvard study in which an OpenAI model matched attending physicians on emergency department triage and admission decisions using real records. Stanford researchers have repeatedly found that a physician using the model does no better than the model alone, and that the gap closes only when the physician commits to an assessment first and then compares.
- Evidence Sept 11, 2026Becker's Hospital Review asked chief information officers and chief medical information officers at Cornell, Baptist Health, HSS, Premier Health, Christ Hospital, North Country Healthcare and Denver Health what the ADVOCATE agents of the Advanced Research Projects Agency for Health (ARPA-H) would need. Dr. Curtis Cole of Cornell said the agents will not succeed if the companies that build and oversee them have no liability. Joy Oh of Christ Hospital asked who is responsible when a supervisory agent misses an inappropriate recommendation, and Dr. Daniel Kortsch of Denver Health said the technology has been tested on curated cases, not safety-net populations. The consensus was narrow use cases, named accountability and escalation paths, not broad autonomy.