
Medical AI in Clinical Practice: What the Evidence Actually Shows in 2026
Introduction
Medical artificial intelligence has passed the point where the interesting question is whether it works. Thousands of tools are cleared for market in the United States, most large health systems are running at least one in production, and four out of five surveyed US physicians now report using some form of AI professionally. The harder and more useful question — the one that determines whether a purchase improves care or quietly degrades it — is which medical AI works, for whom, under what conditions, and with what evidence behind the claim.
That question has become genuinely answerable in the last two years. Randomised trials in cancer screening have reported outcome data rather than accuracy proxies. Multi-site studies of AI documentation tools have measured real time savings against matched controls. Regulators in the US, EU and UK have published lifecycle frameworks that make explicit what evidence they expect. And a body of critical research has quantified how much of the deployed estate rests on thin validation.
This article synthesises that picture. It is written for professionals who must make or defend decisions: clinicians deciding whether to trust an output, executives approving a contract, regulatory staff preparing a submission, developers designing an evidence strategy. It covers what medical AI is, what the strongest evidence currently supports, how the three major regulatory regimes now work and when their deadlines bite, how money flows, where the failure modes are, and what a defensible deployment process looks like. Contested and preliminary findings are labelled as such throughout.
Executive summary
- The US Food and Drug Administration's public list of AI-enabled medical devices recorded 1,451 authorisations through the end of December 2025, of which 1,104 (76%) are in radiology. The FDA states the list is not a comprehensive inventory.
- No generative-AI-enabled medical device had been authorised by the FDA for any clinical purpose as of mid-2026. Every cleared device on the list uses non-generative techniques.
- The strongest outcome evidence in the field comes from cancer screening. The Swedish MASAI trial, published in The Lancet on 29 January 2026, randomised over 105,000 women and reported a non-inferior interval-cancer rate with AI support, alongside higher sensitivity and unchanged specificity.
- The strongest operational evidence concerns documentation. A five-site study of 8,581 ambulatory clinicians found ambient AI scribes saved roughly 16 minutes of documentation time per eight hours of patient care — meaningful but far below vendor-adjacent claims.
- Evidence quality across the cleared estate is uneven. One cross-sectional analysis of 903 devices cleared through August 2024 found only about 56% had publicly available clinical performance data; another found devices lacking clinical validation were disproportionately represented among recalls.
- Regulatory deadlines moved in 2026. The EU's Digital Omnibus, given final Council approval on 29 June 2026, pushed high-risk obligations for AI embedded in medical devices from 2 August 2027 to 2 August 2028.
- Access to a good model does not automatically improve clinician performance. In a randomised trial, physicians given an LLM scored 76% on diagnostic reasoning versus 74% for controls — a non-significant difference — while the model alone scored 92%. Interaction design, not raw capability, was the binding constraint.
- Reimbursement remains the weakest link. There is no unified US payment pathway; roughly 26 CPT codes covered clinical AI as of January 2026, with most tools absorbed into existing bundled payments.
- Industry self-governance has partially stalled. The Coalition for Health AI's proposed network of independent assurance labs did not materialise, though its model card, registry and governance playbooks did.
- The practical bottleneck is local: site-specific validation, integration, monitoring for drift, and clear accountability — not model architecture.
What "medical AI" actually means
The term covers at least four technically distinct families that carry different risks, evidence requirements and regulatory treatment. Conflating them is the single most common source of confused procurement decisions.
Table 1 — Four families of medical AI
| Family | Typical function | Representative examples | Regulatory status | Dominant risk |
|---|---|---|---|---|
| Perceptual / diagnostic | Detect, segment or classify findings in images and signals | Mammography reading support, stroke triage, retinal screening, fracture detection | Usually a regulated device; most FDA clearances sit here | Distribution shift across scanners, sites and populations |
| Predictive / prognostic | Estimate future clinical events from structured records | Deterioration and sepsis scores, readmission risk, no-show prediction | Often deployed inside EHRs under clinical-decision-support exemptions | Silent performance decay; label leakage; poor calibration |
| Generative / language | Draft, summarise, translate or converse | Ambient documentation, discharge-summary drafting, patient messaging support, coding assistance | Largely outside device regulation when used administratively | Hallucination, omission, sycophancy, automation bias |
| Operational / administrative | Optimise scheduling, staffing, supply, revenue cycle | Theatre scheduling, prior-authorisation triage, denial management | Generally unregulated as devices | Fairness and access effects; opaque payer use |
A fifth category — agentic systems that plan and execute multi-step clinical or administrative tasks — is emerging in research settings. It is not yet a mature deployment category, and readers should treat vendor claims in this space as forward-looking rather than validated.
How the field arrived here
Computer-aided detection in medicine is not new; the FDA's device list reaches back to a cervical-smear rescreening tool authorised in 1995, and computer-aided detection for mammography dates to the late 1990s. Those early systems used hand-engineered features and had a mixed real-world record.
The inflection came after 2016, when deep convolutional networks reached expert-comparable accuracy on narrow image tasks. Authorisations accelerated sharply from that point. A second inflection arrived in 2022–2023 with large language and multimodal models, which shifted attention from perception to language, documentation and reasoning — domains where the FDA's device framework applies less cleanly and where health systems have consequently moved fastest with least oversight.
The current period is best understood as a consolidation phase: capability is no longer the limiting factor for most narrow tasks, and the field's centre of gravity has moved to evidence generation, governance, integration and payment.
What the evidence actually supports
Cancer screening: the strongest outcome evidence
The Mammography Screening with Artificial Intelligence (MASAI) trial, run within Sweden's national screening programme by researchers at Lund University, is the field's most consequential result to date. It randomised more than 105,000 women to AI-supported screening or standard double reading.
The trial reported in stages. Interim safety results in The Lancet Oncology (2023) showed a 44% reduction in screen-reading workload. A screening-accuracy analysis in The Lancet Digital Health (2025) reported a 29% increase in cancer detection with no increase in false positives. The definitive interval-cancer analysis, published in The Lancet on 29 January 2026, found 82 interval cancers in the AI arm versus 93 in the control arm — a non-inferior result, with higher sensitivity, equivalent specificity, and fewer interval cancers carrying unfavourable characteristics.
Why this matters: interval cancers — those diagnosed between screening rounds — are the standard test of whether extra detection reflects genuine earlier diagnosis or merely overdiagnosis. MASAI is the first randomised trial to answer that question for AI-assisted screening.
Important limitations. MASAI evaluated one commercial system, within one national programme, using a specific triage-plus-detection-support workflow, in a population screened by an established double-reading standard. It does not license the conclusion that any mammography AI improves any screening programme. Supporting real-world evidence from a nationwide German implementation (Nature Medicine, 2025) and a prospective Korean cohort (Nature Communications, 2025) points in a consistent direction, but these are not randomised.
Documentation: modest, measured, real
Ambient AI scribes — tools that listen to a consultation and draft the note — became the fastest-adopted clinical AI category, driven by documentation burden rather than diagnostic ambition.
The best-controlled evidence comes from the Ambient Clinical Documentation Collaborative, published in JAMA on 1 April 2026. It tracked 8,581 ambulatory clinicians across Mass General Brigham, Emory Healthcare, UCSF, Yale New Haven Health and UC Davis between June 2023 and August 2025, comparing 1,809 adopters with 6,772 non-adopters. Adoption was associated with roughly 13 fewer minutes of total EHR time and 16 fewer minutes of documentation time per eight hours of patient care, plus about half an additional visit per week.
Wellbeing findings have generally been more favourable than time findings. A multicentre quality-improvement study across six health systems (JAMA Network Open, October 2025) reported burnout prevalence falling from 51.9% to 38.8% after 30 days of scribe use among 263 clinicians.
How to read this honestly. Time savings are real but modest, and they are averages that conceal wide variation — benefits concentrate among clinicians who received structured coaching. Wellbeing improvements measured in uncontrolled or self-selected cohorts are vulnerable to expectancy effects. The direction of evidence is positive; the magnitude is smaller than promotional material typically implies.
Generative clinical reasoning: capability without demonstrated clinical benefit
Large language models score extremely well on medical examinations and vignettes. That has not yet translated into demonstrated improvement in physician performance.
The most-cited randomised evidence is Goh and colleagues' trial (JAMA Network Open, October 2024). Fifty physicians were randomised to a leading LLM plus conventional resources, or conventional resources alone. Median diagnostic reasoning scores were 76% versus 74% — an adjusted difference of 2 percentage points (95% CI, −4 to 8; P = .60). The LLM run alone scored 92%, significantly above the control group.
That gap is the central finding of the last two years: the model was better than the physicians and better than the physician–model pair. Subsequent work has explored why. A randomised trial in Pakistan found that physicians given a 20-hour AI-literacy curriculum before LLM access did show substantial gains, while a companion trial in the same programme demonstrated automation bias — when deliberately erroneous suggestions were introduced, trained physicians still absorbed some of the errors.
The defensible conclusion. Model capability, clinician training and interface design are separate variables, and the evidence to date suggests the second and third are currently binding. Claims that deploying a frontier model will improve diagnosis are not supported by trial evidence; claims that it could, given deliberate interaction design and training, are reasonable hypotheses.
Table 2 — Evidence maturity by application
| Application | Best available evidence | Maturity | What is still missing |
|---|---|---|---|
| Mammography screening support | Randomised trial with interval-cancer outcomes (MASAI, 2026) | Established for the studied workflow | Generalisation across vendors, populations, programmes |
| Diabetic retinopathy screening | Prospective validation; established reimbursement codes | Established in defined settings | Long-term outcome and equity data |
| Stroke / large-vessel-occlusion triage | Real-world time-to-treatment studies | Adopted; outcome evidence developing | Randomised outcome data |
| Ambient documentation | Large multi-site controlled study (2026) | Operationally proven, modest effect | Note-quality, safety and downstream-error data |
| Sepsis / deterioration prediction | Mixed; some widely used models failed external validation | Contested | Prospective outcome trials; site-level recalibration standards |
| Generative diagnostic assistance | RCTs show no significant physician-level gain to date | Experimental | Interaction design; trials with clinical endpoints |
| Autonomous generative clinical agents | Research-stage only | Speculative | Essentially everything |
The regulatory picture
United States
The FDA regulates AI tools that meet the statutory definition of a medical device. Its public AI-Enabled Medical Device List recorded 1,451 authorisations through 31 December 2025, dominated by radiology. The agency emphasises the list is compiled from terminology in public authorisation summaries and is not an exhaustive inventory.
Three documents define current expectations:
- Good Machine Learning Practice for Medical Device Development: Guiding Principles (FDA, Health Canada and MHRA, October 2021), adopted in final form by the International Medical Device Regulators Forum in January 2025.
- Marketing Submission Recommendations for a Predetermined Change Control Plan for AI-Enabled Device Software Functions — final guidance, December 2024. A PCCP lets a manufacturer pre-specify permitted model modifications and their validation, so routine updates do not each require a new submission.
- Artificial Intelligence-Enabled Device Software Functions: Lifecycle Management and Marketing Submission Recommendations — draft guidance, January 2025, organised around submission sections rather than software-lifecycle phases. It appeared on the FDA's fiscal-year 2026 "B" list for finalisation; as of this writing it remains in draft.
Two structural gaps deserve attention. First, most authorisations proceed via 510(k), which requires substantial equivalence to a predicate device rather than demonstration of improved clinical outcomes. Second, no generative-AI device has been authorised for any clinical purpose. The FDA's Digital Health Advisory Committee met on 6 November 2025 specifically on generative-AI digital mental-health devices, where members flagged bias, hallucination and sycophancy as novel risks requiring novel trial designs.
European Union
The EU AI Act entered into force on 1 August 2024. Its application timetable changed materially in 2026. The Digital Omnibus simplification package — proposed 19 November 2025, agreed in trilogue on 6–7 May 2026, endorsed by Parliament on 16 June 2026 and given final Council approval on 29 June 2026 — deferred high-risk obligations on fixed dates:
- 2 December 2027 for stand-alone Annex III high-risk systems.
- 2 August 2028 for high-risk AI embedded in products covered by EU harmonisation legislation under Annex I — which is where medical devices and IVDs sit.
- 2 August 2026 remains active for Article 50 transparency obligations, with synthetic-content marking for systems already on the market due by 2 December 2026.
The stated rationale was that harmonised standards and Commission guidance were not ready. The obligations themselves were not weakened. Medical AI in Europe therefore faces a dual regime: the Medical Device Regulation for safety and performance, and the AI Act layered on top from 2028.
United Kingdom
The MHRA has taken a sandbox-led approach. The AI Airlock, launched in spring 2024, ran a four-product pilot to April 2025 and a second phase with seven innovators across three regulatory challenges, completing in May 2026. Phase 2 examined large language models in clinical decision support, synthetic data for radiology validation, and real-time post-market surveillance. In April 2026 the Department of Health and Social Care committed £3.6 million over three years (2026–2029), moving the programme beyond annual pilot cycles, with a Phase 3 in design and findings feeding the MHRA's National Commission into the Regulation of AI in Healthcare.
Table 3 — Comparing the three regimes
| Dimension | United States (FDA) | European Union | United Kingdom (MHRA) |
|---|---|---|---|
| Primary instrument | FD&C Act device pathways (510(k), De Novo, PMA) | MDR/IVDR plus AI Act | UK MDR, under reform |
| Handling of model updates | PCCP (final guidance, Dec 2024) | Substantial-modification rules; AI Act change management | Under exploration via AI Airlock |
| Generative AI | No device authorised to date | Covered by AI Act; sectoral rules unsettled | Tested in Airlock cohorts |
| Key near-term date | Finalisation of Jan 2025 lifecycle draft guidance | 2 Aug 2028 for device-embedded high-risk AI | Phase 3 Airlock; Commission recommendations |
| Characteristic strength | Volume and speed of authorisation | Breadth and horizontal coverage | Iterative, evidence-generating |
| Characteristic weakness | 510(k) rarely requires outcome evidence | Complexity and standards lag | Smaller market pull |
Figure descriptions for the OneWise design team
Figure 1 — The Medical AI Lifecycle: From Clearance to Continuous Assurance
Purpose. Show that regulatory authorisation is one gate among many, and that most deployment risk sits after purchase.
Structure. A horizontal five-stage pipeline occupying the upper two-thirds, with a feedback loop beneath.
Stages (left to right), each a rounded rectangle with a numbered badge:
- Development — labels: training data provenance, intended use definition, subgroup representation
- Regulatory authorisation — labels: 510(k) / De Novo / CE mark, PCCP, GMLP alignment
- Local validation — labels: site-specific test set, subgroup performance, calibration check, threshold setting
- Integration & deployment — labels: EHR workflow placement, alert design, clinician training, escalation path
- Monitoring — labels: drift detection, override rates, incident reporting, scheduled revalidation
Arrows. Solid right-facing arrows connect stages 1→5. A thick dashed arrow returns from stage 5 to stage 3, annotated "Revalidate on data shift, model update, or population change." A second dashed arrow returns from 5 to 1, annotated "Feed real-world performance to developer."
Visual hierarchy. Stages 3 and 5 rendered in the accent colour and slightly larger, with a subtle band beneath the whole pipeline labelled "Accountability sits with the deploying organisation from Stage 3 onward."
Caption. "Regulatory clearance establishes that a tool performed adequately on the manufacturer's evidence. It does not establish that the tool performs adequately in your setting. Stages 3 and 5 are where deploying organisations carry the residual risk."
Figure 2 — Four Failure Modes of Deployed Clinical AI
Purpose. Give clinical and governance teams a shared vocabulary for how AI tools fail after go-live.
Structure. A 2×2 grid of equal quadrants. Vertical axis: "Failure is visible ↔ Failure is silent." Horizontal axis: "Cause is technical ↔ Cause is human-system."
Quadrant contents:
- Visible / technical — Distribution shift: new scanner, new patient mix, degraded input quality. Signal: sudden change in output distribution.
- Silent / technical — Performance drift: gradual decay as clinical practice, coding or case mix evolves. Signal: none without active monitoring.
- Visible / human-system — Alert fatigue: high false-positive volume drives systematic dismissal. Signal: rising override rate.
- Silent / human-system — Automation bias: clinicians accept incorrect outputs they would otherwise have caught. Signal: none without deliberate audit.
Annotation. A thin arrow along the bottom labelled "Detection cost rises left to right and top to bottom." A footnote box: "The two silent quadrants account for most unmeasured harm."
Caption. "Only one of the four common failure modes announces itself. Monitoring plans built solely around technical uptime will miss the other three."
Risks, unresolved questions and honest uncertainty
Evidence thinness is measurable, not rhetorical. A cross-sectional analysis of 903 devices cleared through August 2024 found roughly 56% had publicly available clinical performance data. A study of the 168 machine-learning-enabled Class II devices authorised during 2024 found only 29.2% reported both sensitivity and specificity, 15.5% provided demographic data on evaluation populations, and 16.7% included a PCCP. A JAMA Health Forum analysis published 22 August 2025 examined 950 devices and linked 60 of them to 182 recall events, with about 43% of recalls occurring within a year of authorisation — and found devices lacking clinical validation over-represented among recalls.
External validation frequently disappoints. The best-documented example remains a widely deployed proprietary sepsis prediction model that performed substantially worse in independent evaluation than in vendor materials. The generalisable lesson is not about one vendor: it is that performance claims generated on development data have limited predictive value for a new site.
Bias is structural, not incidental. The canonical demonstration — Obermeyer and colleagues in Science (2019) — showed that an algorithm allocating care management used healthcare cost as a proxy for health need, systematically under-referring Black patients because less was historically spent on them. The mechanism, a plausible-seeming proxy label encoding structural inequity, recurs across applications.
Generative-specific risks lack mature mitigations. Hallucination, omission, and sycophancy behave differently from classifier errors: they are fluent, contextual and hard to detect by sampling. Emerging research on clinician editing of AI-drafted notes suggests systematic changes to hedging and uncertainty language, with implications for downstream interpretation that are not yet well characterised. This is an active research area, and readers should treat current findings as preliminary.
Governance infrastructure is incomplete. The Coalition for Health AI proposed a national network of independent assurance labs to test models before hospital deployment. Reporting in February 2026 found those labs never materialised, leaving a gap the organisation had been widely expected to fill. CHAI did deliver a model card template, a public governance registry launched in February 2025, joint guidance with the Joint Commission, and a set of governance playbooks across eight domains released on 27 May 2026. The gap between voluntary frameworks and enforceable assurance remains open.
The candid summary from the field's own leadership. The JAMA Summit Report on Artificial Intelligence (JAMA, 11 November 2025), authored by more than fifty clinicians, researchers, policymakers and industry figures, concluded that health AI is being adopted faster than evidence about its consequences is being generated, and called for an ecosystem capable of producing rapid, robust and generalisable knowledge about real-world effects.
The economics: where medical AI stalls
Clinical validity does not produce revenue. Payment does, and payment is the field's least developed layer.
In the US there is no unified reimbursement pathway for AI. Tools reach payment through fragmentary routes: the New Technology Add-on Payment under the inpatient system, the Transitional Coverage for Emerging Technologies pathway for breakthrough devices, Category I and III CPT codes, and local coverage determinations. An analysis by the Bipartisan Policy Center found roughly 26 CPT codes covering clinical AI solutions as of January 2026. Most AI is otherwise absorbed into existing bundled payments, meaning the hospital bears the cost and the payer captures any savings — a structural disincentive.
There has been movement. Category I codes now exist for AI-assisted retinal imaging analysis and certain cardiac imaging interpretation, and the 2026 Hospital Outpatient Prospective Payment System final rule established national payment for AI-assisted cardiac analysis. But these remain exceptions.
The practical consequence. Tools that reduce cost within a single organisation's budget (documentation, scheduling, revenue cycle) scale quickly because the buyer captures the benefit. Tools that improve outcomes across a longer time horizon scale slowly. This is an economic fact about payment design, not a statement about clinical value — and it explains most of the observed adoption pattern.
Latest developments
Dated items relevant to decisions being made now. Company statements and preliminary research are labelled as such.
- 29 June 2026 — EU Digital Omnibus given final Council approval. High-risk obligations for AI embedded in medical devices deferred from 2 August 2027 to 2 August 2028; stand-alone Annex III systems to 2 December 2027. Article 50 transparency duties remain on the original 2 August 2026 track. Established regulatory fact.
- 27 May 2026 — CHAI publishes governance playbooks across eight domains, developed with over 150 health-AI leaders. Voluntary framework, not binding.
- 1 April 2026 — Largest controlled study of ambient AI scribes published in JAMA: 8,581 clinicians, five academic systems, ~16 minutes of documentation time saved per eight hours of care. Peer-reviewed observational evidence with matched controls.
- April 2026 — MHRA secures £3.6 million over three years for the AI Airlock, following Phase 2 completion in May 2026. Established.
- March 2026 — FDA reportedly granted breakthrough device designation to a patient-facing clinical generative-AI application, per a Congressional Research Service report. Breakthrough designation is not marketing authorisation. Single authoritative source; treat as reported rather than fully corroborated.
- 12 March 2026 — AMA publishes its 2026 Physician Survey on Augmented Intelligence: 81% of 1,692 surveyed physicians report professional AI use, up from 66% (2024) and 38% (2023); average 2.3 use cases per physician; 85% want to be consulted on adoption decisions; 92% want more training. Self-reported survey data.
- February 2026 — Reporting that CHAI's assurance-lab network never launched. Investigative journalism; CHAI has continued other workstreams.
- 29 January 2026 — MASAI final results published in The Lancet, reporting non-inferior interval-cancer rates with AI-supported screening. Randomised trial; strongest outcome evidence in the field.
- 11 November 2025 — JAMA Summit Report on AI published, calling for faster and more robust real-world evidence generation. Expert consensus statement, not primary research.
- 6 November 2025 — FDA Digital Health Advisory Committee convened on generative-AI digital mental-health devices. Advisory discussion; no binding outcome.
Practical takeaways
For clinicians. Ask what population and equipment a tool was validated on, and whether that resembles your patients. Treat a confident generative output as a hypothesis requiring the same scrutiny you would apply to a colleague's suggestion — the automation-bias evidence indicates trained physicians still absorb model errors. Record your reasoning when you override a tool; override data is among the most useful monitoring signals a health system has.
For health-system executives. Budget for local validation and ongoing monitoring as a permanent operating cost, not a one-off implementation line. Require vendors to supply subgroup performance and a documented monitoring plan before contracting. Assume the burden of proof sits with you after purchase, because regulatory clearance does not transfer to your setting. Fund clinician coaching — the multi-site scribe evidence shows benefits concentrate among coached users.
For regulatory and quality professionals. If you market in the EU, work backwards from 2 August 2028 and treat the deferral as engineering runway rather than reprieve; harmonised standards are still in development. In the US, build a PCCP into the submission strategy early. Align documentation to GMLP principles, which now have IMDRF standing.
For developers. Design the evidence strategy alongside the product. Prospective evaluation with clinical endpoints is expensive and slow, but the recall data indicates that clearance without clinical validation correlates with post-market trouble. Report sensitivity, specificity and subgroup performance transparently; the current baseline is low enough that doing so is a genuine differentiator.
For policy analysts. The binding constraints are payment design and post-market surveillance infrastructure, not premarket rules. A framework that authorises tools efficiently but cannot detect silent degradation after deployment addresses the smaller half of the problem.
Common myths and mistakes
| Common claim | What the evidence indicates |
|---|---|
| "FDA clearance means the tool improves outcomes." | Most clearances proceed via 510(k), which requires substantial equivalence to a predicate — not outcome improvement. |
| "AI outperforms doctors, so it will improve care." | In the best-known RCT the model outscored physicians and physician–model pairs. Capability did not transfer to practice. |
| "Bigger models will solve clinical AI." | Site-level validation, workflow integration, monitoring and payment are the current constraints. None are model-scale problems. |
| "Once validated, a model stays valid." | Performance drifts silently as equipment, coding, case mix and practice change. Scheduled revalidation is mandatory, not optional. |
| "Bias is fixed by removing race from the inputs." | The canonical failure used cost as a proxy for need. Bias enters through label choice and structure, not just explicit variables. |
| "Ambient scribes eliminate documentation work." | Controlled data show roughly 13–16 minutes saved per eight clinical hours, with wide variation and continued clinician review. |
| "The EU deferral means Europe deregulated AI." | The obligations are unchanged; only the dates moved, because standards were not ready. |
Key insights
- Medical AI's centre of gravity has shifted from capability to evidence, governance and payment.
- The deployed estate is far narrower than the discourse implies. Radiology accounts for 1,104 of 1,451 FDA authorisations — 76% — meaning most clinical specialties have very few cleared tools and almost no specialty-specific outcome evidence.
- Cancer screening holds the strongest outcome evidence, following the randomised MASAI results published in January 2026.
- Documentation tools deliver modest but real, measured operational gains — roughly a quarter-hour per eight clinical hours.
- Generative diagnostic assistance has not yet demonstrated physician-level benefit in randomised trials, despite strong standalone model performance.
- Automation bias is empirically demonstrated, and AI-literacy training reduces but does not eliminate it.
- A large share of the cleared estate lacks public clinical performance data, and unvalidated devices are over-represented in recalls.
- Transparency reporting is the exception, not the norm. Among the 168 machine-learning-enabled Class II devices authorised in 2024, only 29.2% reported both sensitivity and specificity, 15.5% disclosed demographic data on evaluation populations, and 16.7% included a change-control plan — a low enough baseline that disclosure functions as a genuine commercial differentiator.
- No generative-AI medical device has been authorised by the FDA for any clinical purpose as of mid-2026.
- The EU's high-risk deadline for device-embedded AI is now 2 August 2028, following the Digital Omnibus approved on 29 June 2026.
- There is no independent assurance infrastructure at national scale. The Coalition for Health AI's proposed network of assurance labs did not launch, leaving pre-deployment testing a local responsibility that most organisations are not resourced to discharge.
- Reimbursement is the strongest determinant of adoption pattern, favouring cost-reducing over outcome-improving tools.
- Post-deployment accountability sits with the deploying organisation, and most unmeasured harm arises from failure modes that produce no visible signal.
Frequently asked questions
What is medical AI? Software that uses machine learning to perform clinically relevant tasks — detecting findings in images, predicting patient events, drafting documentation, or supporting decisions. It spans regulated medical devices and unregulated administrative tools.
How many AI medical devices has the FDA authorised? The FDA's public list recorded 1,451 authorisations through 31 December 2025, with 1,104 (76%) in radiology. The agency notes the list is not a comprehensive inventory of all AI-enabled devices.
Has the FDA approved any generative AI or LLM-based medical device? No. As of mid-2026 the FDA had not authorised a generative-AI-enabled device for any clinical purpose. Generative tools used for administrative documentation typically fall outside the device definition.
Does AI improve cancer screening? The MASAI randomised trial, published in The Lancet in January 2026, found AI-supported mammography screening produced a non-inferior interval-cancer rate with higher sensitivity, unchanged specificity, and substantially reduced reading workload. This applies to the specific system and workflow studied.
Do AI scribes reduce physician burnout? Evidence points that way, with caveats. A six-system study found burnout prevalence fell from 51.9% to 38.8% after 30 days of use, though this was a quality-improvement design without a control group. Controlled time data show smaller gains than uncontrolled studies suggest.
Do LLMs make doctors better at diagnosis? Not demonstrably, on current trial evidence. A randomised trial found physicians with LLM access scored 76% versus 74% for controls — a non-significant difference — while the model alone scored 92%.
When does the EU AI Act apply to medical devices? High-risk obligations for AI embedded in medical devices apply from 2 August 2028, following the Digital Omnibus approved by the Council on 29 June 2026. Article 50 transparency obligations began on 2 August 2026.
What is a Predetermined Change Control Plan? A PCCP is a pre-authorised specification of how a manufacturer may modify an AI model after market entry, including the validation each change requires. It lets routine updates proceed without a new marketing submission. FDA final guidance issued December 2024.
What is Good Machine Learning Practice? A set of ten guiding principles published jointly by the FDA, Health Canada and the MHRA in October 2021, covering multidisciplinary development, data quality, training/test independence, human-AI integration and transparency. The IMDRF adopted them in final form in January 2025.
Why do AI models fail when moved between hospitals? Because performance depends on data distribution. Different scanners, laboratories, coding practices, documentation habits and patient populations shift the input distribution away from the training data. This is why local validation is necessary regardless of clearance status.
What is automation bias in clinical AI? The tendency to accept an automated recommendation that would have been rejected if it came from another source. Randomised evidence shows it persists even among clinicians who have completed AI-literacy training.
How is medical AI reimbursed in the United States? There is no unified pathway. Tools reach payment through NTAP, the TCET pathway, Category I or III CPT codes, or local coverage determinations. Roughly 26 CPT codes covered clinical AI as of January 2026; most tools are absorbed into existing bundled payments.
Who is liable when clinical AI causes harm? Liability allocation is unsettled and jurisdiction-dependent, which is why it consistently appears as a top clinician concern. In practice, clinicians retain responsibility for care decisions and deploying organisations retain responsibility for the safety of tools they implement. This is a general description, not legal advice.
What should a health system check before buying a clinical AI tool? Intended use and population; validation data including subgroup performance; whether validation was prospective; the monitoring and revalidation plan; integration and alert design; escalation pathway; update and PCCP arrangements; and contractual allocation of responsibility for post-deployment performance.
Are AI assurance labs available to test models independently? Not at national scale. The Coalition for Health AI's proposed network of independent assurance labs did not launch, though the organisation published model cards, a public registry and governance playbooks. Independent testing remains largely a local responsibility.
Will AI replace radiologists? There is no evidence supporting this. Radiology holds the most cleared AI devices, yet the trial evidence supports AI as reading support within an existing workflow rather than replacement, and demand for imaging interpretation has continued to grow.
Glossary
510(k) — US premarket notification pathway requiring demonstration that a device is substantially equivalent to a legally marketed predicate.
Agentic AI — Systems that plan and execute multi-step tasks with limited step-by-step human direction. Research-stage in clinical medicine.
Ambient documentation — Software that captures a clinical conversation and drafts a structured note.
Automation bias — Undue acceptance of automated output over independent judgement.
Calibration — The correspondence between a model's predicted probabilities and observed event frequencies. A model can discriminate well yet be badly calibrated.
CONSORT-AI / SPIRIT-AI — Reporting guidelines extending trial reporting and protocol standards to interventions involving AI (2020).
De Novo — US pathway for novel low-to-moderate-risk devices without a suitable predicate.
Distribution shift — Divergence between deployment-time input data and training data, a principal cause of performance loss.
Drift — Gradual degradation of model performance over time as the underlying data-generating process changes.
Foundation model — A large model pretrained on broad data and adapted to many downstream tasks.
GMLP — Good Machine Learning Practice; ten guiding principles from the FDA, Health Canada and MHRA (2021), adopted by the IMDRF (2025).
Hallucination — Fluent but unfounded generated content.
IMDRF — International Medical Device Regulators Forum, a voluntary group of medical-device regulators pursuing harmonisation.
Interval cancer — A cancer diagnosed between scheduled screening rounds after a negative screen; a standard measure of screening-programme sensitivity.
MDR — EU Medical Device Regulation (2017/745).
Model card — A structured disclosure document describing a model's intended use, performance, evaluation populations and known limitations.
NTAP — New Technology Add-on Payment; US Medicare mechanism providing payment above standard inpatient DRG rates for qualifying new technologies.
PCCP — Predetermined Change Control Plan; a pre-authorised specification of permitted post-market model modifications and their validation.
SaMD — Software as a Medical Device; software intended for a medical purpose that is not part of a hardware device.
Sycophancy — A model's tendency to align with a user's stated view at the expense of accuracy.
TCET — Transitional Coverage for Emerging Technologies; a CMS pathway coordinating coverage review with FDA premarket review for breakthrough devices.
TRIPOD+AI — Reporting guideline for studies developing or validating clinical prediction models, including those using machine learning (2024).
References
Academic papers
Angus, D. C., Khera, R., Lieu, T., Liu, V., Ahmad, F. S., Anderson, B., … Weinstein, J. (2025). AI, health, and health care today and tomorrow: The JAMA Summit report on artificial intelligence. JAMA, 334(18), 1650–1664. https://doi.org/10.1001/jama.2025.18490
Chang, Y.-W., Ryu, J. K., An, J. K., et al. (2025). Artificial intelligence for breast cancer screening in mammography (AI-STREAM): Preliminary analysis of a prospective multicenter cohort study. Nature Communications, 16, 2248.
Chen, R. J., Ding, T., Lu, M. Y., Williamson, D. F. K., Jaume, G., Song, A. H., … Mahmood, F. (2024). Towards a general-purpose foundation model for computational pathology. Nature Medicine, 30(3), 850–862.
Cruz Rivera, S., Liu, X., Chan, A.-W., Denniston, A. K., & Calvert, M. J. (2020). Guidelines for clinical trial protocols for interventions involving artificial intelligence: The SPIRIT-AI extension. Nature Medicine, 26(9), 1351–1363.
Dai, T., et al. (2025, August 22). Benefit-risk reporting for FDA-cleared artificial intelligence and machine learning devices. JAMA Health Forum. https://doi.org/10.1001/jamahealthforum.2025.3351
Eisemann, N., Bunk, S., Mukama, T., et al. (2025). Nationwide real-world implementation of AI for cancer detection in population-based mammography screening. Nature Medicine, 31, 917–924.
Goh, E., Gallo, R., Hom, J., et al. (2024). Large language model influence on diagnostic reasoning: A randomized clinical trial. JAMA Network Open, 7(10), e2440969. https://doi.org/10.1001/jamanetworkopen.2024.40969
Hernström, V., et al. (2025). Screening performance and characteristics of breast cancer detected in the Mammography Screening with Artificial Intelligence trial (MASAI): A randomised, controlled, parallel-group, non-inferiority, single-blinded, screening accuracy study. The Lancet Digital Health, 7, e175–e183.
Lång, K., et al. (2026). Interval cancer, sensitivity, and specificity comparing AI-supported mammography screening with standard double reading without AI in the MASAI study. The Lancet, 407(10527). https://doi.org/10.1016/S0140-6736(25)02464-X
Liu, X., Cruz Rivera, S., Moher, D., Calvert, M. J., & Denniston, A. K. (2020). Reporting guidelines for clinical trial reports for interventions involving artificial intelligence: The CONSORT-AI extension. Nature Medicine, 26(9), 1364–1374.
Moor, M., Banerjee, O., Abad, Z. S. H., et al. (2023). Foundation models for generalist medical artificial intelligence. Nature, 616, 259–265.
Obermeyer, Z., Powers, B., Vogeli, C., & Mullainathan, S. (2019). Dissecting racial bias in an algorithm used to manage the health of populations. Science, 366(6464), 447–453.
Olson, K. D., Meeker, D., Troup, M., et al. (2025). Use of ambient AI scribes to reduce administrative burden and professional burnout. JAMA Network Open, 8(10), e2534976. https://doi.org/10.1001/jamanetworkopen.2025.34976
Qazi, I. A., Ali, A., Khawaja, A. U., et al. (2026). Large language model diagnostic assistance for physicians in a lower-middle-income country: A randomized controlled trial. Nature Health. https://doi.org/10.1038/s44360-025-00007-8
Wong, A., Otles, E., Donnelly, J. P., et al. (2021). External validation of a widely implemented proprietary sepsis prediction model in hospitalized patients. JAMA Internal Medicine, 181(8), 1065–1070.
Official documentation and standards
International Medical Device Regulators Forum. (2025). Good machine learning practice for medical device development: Guiding principles (final).
Medicines and Healthcare products Regulatory Agency. (2026). AI Airlock: The regulatory sandbox for AIaMD. GOV.UK. https://www.gov.uk/government/collections/ai-airlock-the-regulatory-sandbox-for-aiamd
Medicines and Healthcare products Regulatory Agency. (2026, April). MHRA expands AI Airlock programme with a £3.6 million funding boost over three years. GOV.UK.
U.S. Food and Drug Administration. (2024, December). Marketing submission recommendations for a predetermined change control plan for artificial intelligence-enabled device software functions: Guidance for industry and FDA staff.
U.S. Food and Drug Administration. (2025, January 6). Artificial intelligence-enabled device software functions: Lifecycle management and marketing submission recommendations (draft guidance).
U.S. Food and Drug Administration. (2025, November 6). Digital Health Advisory Committee meeting: Generative artificial intelligence-enabled digital mental health medical devices — executive summary and brief summary.
U.S. Food and Drug Administration. (2026). Artificial intelligence-enabled medical device list. https://www.fda.gov/medical-devices/software-medical-device-samd/artificial-intelligence-enabled-medical-devices
Government and intergovernmental sources
Congressional Research Service. (2026). FDA regulation of AI-enabled devices (IF13245). https://www.congress.gov/crs-product/IF13245
Council of the European Union. (2026, June 29). Digital Omnibus: Final adoption of the AI Act simplification package.
European Parliament and Council of the European Union. (2024). Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act). Official Journal of the European Union.
World Health Organization. (2024). Ethics and governance of artificial intelligence for health: Guidance on large multi-modal models. Geneva: WHO. https://www.who.int/publications/i/item/9789240084759
Industry reports and professional bodies
American Medical Association. (2026, March). 2026 physician survey on augmented intelligence. AMA Center for Digital Health and AI. https://www.ama-assn.org/practice-management/digital-health/physician-survey-augmented-intelligence
Bipartisan Policy Center. (2026, March). Paying for AI in U.S. health care. https://bipartisanpolicy.org/issue-brief/paying-for-ai-in-u-s-health-care/
Coalition for Health AI. (2026, May 27). CHAI releases comprehensive governance playbooks to streamline AI implementation for health systems. https://www.chai.org/news/
Joint Commission & Coalition for Health AI. (2025). Responsible use of AI in healthcare (RUAIH) guidance.
Note on sourcing: figures cited for the FDA device list, the AMA survey, the EU legislative timetable and published trials are drawn from the primary sources above. Where a claim rests on a single report — notably the March 2026 breakthrough designation reported by the Congressional Research Service — this is stated in the text.
One Tech & AI · Friday, July 31, 2026 · 34 min read
Early Disease Detection: AI analyzes medical images, lab results, and patient data to detect diseases such as cancer, heart conditions, and neurological disorders earlier, enabling faster and more accurate diagnoses.
Personalized Patient Care: By evaluating a patient's medical history, genetics, and lifestyle, AI helps healthcare professionals create customized treatment plans that improve outcomes and reduce unnecessary interventions.
Smarter Healthcare Systems: AI automates administrative tasks, supports clinical decision-making, predicts patient risks, and optimizes hospital operations, allowing healthcare providers to deliver more efficient and cost-effective care.
Conclusion
Medical AI in 2026 is neither the transformation its promoters describe nor the disappointment its critics anticipated. It is a technology with one genuinely strong outcome result in cancer screening, one solid operational result in clinical documentation, a large cleared estate of narrow imaging tools whose real-world value varies widely, and a fast-moving generative frontier whose clinical benefit remains unproven despite impressive standalone capability.
What has changed most is the quality of the questions being asked. Five years ago the field argued about accuracy metrics. It now argues about interval cancers, matched controls, subgroup reporting, drift detection and payment design — which is the argument a maturing clinical technology is supposed to have.
Three uncertainties deserve to be stated plainly. We do not know how well the MASAI result generalises beyond its system, workflow and population. We do not know how to make clinician–model collaboration reliably better than either party alone, though evidence increasingly suggests training and interface design matter more than model capability. And we do not have working post-market surveillance infrastructure for tools that can degrade without producing any visible signal.
The most consequential insight from the last two years is also the least technological. In the trial that has defined the field's self-understanding, the model outperformed the physicians who were using it. That result is not an argument for autonomy; the model was tested on curated vignettes, not on patients with incomplete histories and competing priorities. It is an argument that the value of medical AI is realised or lost in the space between a capable system and the clinical work it is meant to support — and that space is designed, trained for, measured and governed by people, not by models.
TOPIC
Health Tech