Key Points
- Hospitals are deploying AI rapidly, but evidence that systems improve patient outcomes remains limited.
- Sepsis alerts show how strong laboratory performance can falter amid complex workflows and clinician fatigue.
- Weak validation and uneven oversight could amplify errors, bias and disparities between healthcare institutions.
The latest
Artificial intelligence is spreading through hospitals, clinics and research faster than evidence of its real-world clinical value, sharpening scrutiny of whether accurate models improve outcomes after deployment. Kayla Secrest saw the gap during residency at a Michigan hospital, where a sepsis-detection system scanned records every 15 minutes but produced so many alerts that clinicians began tuning them out. She later learned the tool had been deployed across hundreds of hospitals without extensive testing in real clinical environments.
Details
- Clinical efficacy: Yale researcher Jess Morley said developers often prioritise statistical validation, while healthcare needs evidence that a system changes patient outcomes. Eric Topol of Scripps Research cited stronger evidence in imaging, including a 2024 study in which AI-assisted gastroenterologists detected substantially more polyps during colonoscopies.
- Reported efficiencies: OpenAI said AdventHealth cut time spent on certain administrative tasks by 80% using ChatGPT for Healthcare. An independent review of its work with Kenya-based Penda Health reported a 16% reduction in diagnostic errors among clinicians using the AI assistant, while other systems are being tested for monitoring, risk prediction and drug development.
- Transparency limits: University of Michigan clinician-researcher Andrew Wong said efficiency, workflow and savings claims can be difficult to assess without details on training, datasets, development and independent validation. Utrecht researchers also found “spin” in model papers: 20 of 21 abstracts recommending daily clinical use lacked external validation.
- Regulatory strain: A Stanford report found the US Food and Drug Administration authorised 258 AI medical devices in 2025. Most required no new clinical trials, and 2.4% had the type of trial evidence generally required for drugs. Generative models can also be updated, retrained or replaced, complicating conventional testing and post-market oversight.
- Uneven capacity: Paige Nong’s 2024 research found only slightly more than half of surveyed hospitals had assessed AI tools for potential bias before resource allocation or patient care. Nong said understaffed rural hospitals may lack capacity for such reviews, raising the risk that adoption reinforces institutional disparities.
- Embedded bias: Ziad Obermeyer of the University of California, Berkeley, studied a resource-allocation algorithm that used previous healthcare spending as a proxy for future need. Because spending can reflect access rather than illness, it risked overlooking patients who historically received less care, including some ethnic minorities and non-English speakers.
Between the lines
The central divide is between prediction and consequence: laboratory accuracy measures whether a model identifies a condition, while clinical efficacy asks whether its use produces fewer errors, better outcomes or measurable efficiencies amid actual hospital workflows.
What’s next
UK regulators will decide whether to accept a national commission’s proposal for conditional authorisation until AI systems demonstrate safety and effectiveness, alongside continuous monitoring. US regulators are examining approvals based on evidence generated through real-world use.