The Local Validation Gap: Why FDA-Cleared Clinical AI Still Has to Prove Itself in Your Hospital

The Local Validation Gap: Why FDA-Cleared Clinical AI Still Has to Prove Itself in Your Hospital

Clinical artificial intelligence has crossed the line from pilot to plumbing. In the American Medical Association's 2026 Physician Survey on Augmented Intelligence, fielded between 15 January and 2 February 2026 among 1,692 physicians, 81% reported using AI in practice — more than double the 38% recorded in the same survey series in 2023 (American Medical Association [AMA], 2026). Regulatory throughput has kept pace: the US Food and Drug Administration's public list of AI-enabled medical devices records more than 1,400 marketing authorisations, roughly three-quarters of them in radiology (FDA, 2026a; MedTech Dive, 2026).

Yet the question that most often decides whether a deployment helps or harms patients is not answered by any of those numbers. It is a local one: does this tool work here, on our patients, in our workflow, at the threshold we have chosen? Two regulatory shifts in 2026 have made that question sharper. In January, the FDA narrowed the set of clinical decision support (CDS) software it intends to regulate as devices. In May and June, the Joint Commission and the Coalition for Health AI (CHAI) published governance playbooks and launched a voluntary certification that puts the burden of validation and monitoring squarely on the healthcare organisation.

This article is for the people who now carry that burden: clinical leaders, informaticists, quality and safety teams, procurement and compliance functions, and the executives who sign the contract. It explains what regulatory authorisation does and does not establish, what local validation involves in practice, what the peer-reviewed evidence shows across three very different deployment archetypes, and how to monitor systems after go-live. It draws only on published research, official guidance and documented implementations; where evidence is thin or contested, it says so.


Executive summary

  • Clearance is a market-entry decision, not a clinical benefit claim. Of 903 AI-enabled devices authorised through August 2024, 55.9% had publicly reported clinical performance studies; 28.7% reported results by sex and 23.2% by age (Windecker et al., 2025).
  • Most clinical AI reaches market via the 510(k) pathway, which requires substantial equivalence to a predicate device rather than prospective clinical testing.
  • Post-market signals matter. A cross-sectional study of 950 authorised devices found 60 associated with 182 recall events, with about 43% of recalls occurring within a year of authorisation (Lee et al., 2025).
  • Local performance can collapse. The Epic Sepsis Model, deployed at hundreds of US hospitals, achieved an area under the curve (AUC) of 0.63 and 33% sensitivity in external validation at Michigan Medicine — far below the developer-reported 0.76–0.83 (Wong et al., 2021).
  • The regulatory perimeter narrowed in 2026. The FDA's revised CDS guidance (issued 6 January 2026, reissued 29 January 2026) extends enforcement discretion to some single-recommendation CDS, meaning fewer tools carry an FDA marker for buyers to rely on.
  • Accreditation is filling the gap. The Joint Commission and CHAI issued responsible-use guidance on 17 September 2025; CHAI released governance playbooks on 27 May 2026; the Joint Commission launched a voluntary Responsible Use of AI in Healthcare certification in late May 2026.
  • Evidence quality varies enormously by archetype — strong for autonomous diabetic retinopathy screening and computer-aided polyp detection, mixed and early for ambient documentation.
  • Unintended effects are now documented, not hypothetical. An observational study of 1,443 unassisted colonoscopies found adenoma detection fell from 28.4% to 22.4% after endoscopists became habituated to AI assistance (Budzyń et al., 2025).
  • EU timelines have moved. Under the Digital Omnibus on AI agreed in 2026, high-risk obligations for AI embedded in regulated medical devices shift from 2 August 2027 to 2 August 2028.

Why authorisation is not proof of benefit

Regulators answer a bounded question: is there reasonable assurance of safety and effectiveness for a stated intended use? For most clinical AI, that question is answered through the 510(k) route by demonstrating substantial equivalence to an existing predicate device. A scoping review of FDA-regulated radiology AI found that of 717 devices with available submission documentation, 5% underwent prospective testing and 8% included a human-in-the-loop evaluation (Sivakumar et al., 2025). Substantial equivalence is a reasonable regulatory instrument for stable hardware. It is a weaker instrument for a statistical model whose behaviour depends on the population, the scanner, the coding practices and the alert thresholds of the site where it runs.

The transparency picture is similarly uneven. Across 903 authorised devices, fewer than a third of clinical performance studies reported sex-stratified results and fewer than a quarter reported age-stratified results (Windecker et al., 2025), which means a purchaser usually cannot tell from public documents whether performance holds for the subgroups they serve.

None of this makes cleared devices unsafe. It means the clearance answers a different question from the one a chief medical information officer needs answered.

Callout — The buyer's shift. Where FDA clearance was informally used as a free quality proxy during procurement, the narrowing of device oversight in January 2026 removes that proxy for a growing category of tools. Diligence moves to the purchaser.


What changed in 2025–2026

Table 1. Governance and regulatory milestones shaping clinical AI deployment

DateDevelopmentStatusPractical effect for health systems
Dec 2024FDA final guidance on Predetermined Change Control Plans (PCCPs)FinalManufacturers can pre-authorise defined model updates; buyers should ask what the PCCP permits
6 Jan 2025FDA draft guidance, AI-Enabled Device Software Functions: Lifecycle Management and Marketing Submission RecommendationsDraft; on CDRH's FY2026 agendaSignals expectations on performance monitoring and transparency; not binding
17 Sep 2025Joint Commission & CHAI, Responsible Use of AI in HealthcareNon-binding guidanceSeven elements including governance, local validation and ongoing monitoring
6 Jan 2026 (reissued 29 Jan 2026)FDA final CDS guidanceFinal, non-bindingEnforcement discretion for some single-recommendation CDS; fewer tools regulated as devices
27 May 2026CHAI governance playbooks (eight elements)Published, openBaseline controls mapped to certification
Late May 2026Joint Commission Responsible Use of AI in Healthcare certificationVoluntary, liveCertifies organisations — not products — across five standard areas
2026 (agreement May; Council adoption June)EU Digital Omnibus on AIAdopted; verify Official Journal statusHigh-risk obligations for AI in regulated devices move to 2 Aug 2028; Annex III systems to 2 Dec 2027

Table compiled by OneWise from primary regulatory and standards sources listed in the references.

Two features of this landscape deserve emphasis. First, the Joint Commission certification explicitly does not validate individual products; it assesses whether an organisation has governance, data management, bias reduction, monitoring and training in place. Second, the EU timeline shift buys manufacturers preparation time but does not weaken the substantive requirements on data governance, logging and human oversight that arrive with them.


What local validation actually involves

Local validation is often reduced to "we tested it on our data." Done properly it is a structured, multi-week exercise with a defined stopping rule.

Figure 1 (for OneWise design team). Title: From procurement to sunset: the clinical AI deployment lifecycle Purpose: Show that validation is a loop, not a gate, and that monitoring feeds back into threshold setting. Components (left to right, six stages in rounded rectangles): (1) Intended-use definition and risk classification → (2) Evidence review and vendor due diligence → (3) Silent (shadow) evaluation on local data → (4) Threshold and workflow design → (5) Staged clinical go-live → (6) Continuous monitoring. Feedback arrows: a bold curved arrow from stage 6 back to stage 4 labelled "recalibrate / retune thresholds"; a dashed arrow from stage 6 back to stage 1 labelled "re-approve or retire." Side annotations (small boxes attached above the relevant stage): stage 3 — "discrimination, calibration, subgroup performance, alert burden"; stage 5 — "one unit or service line first"; stage 6 — "drift, override rates, downstream outcomes." Visual hierarchy: stages in a single horizontal band; feedback arrows below the band; annotations above. Caption: Validation is continuous. Performance measured before go-live expires as populations, documentation practices and upstream systems change.

The core technical steps are:

  • Discrimination and calibration on local data. AUC alone is insufficient; a model can rank patients acceptably while producing badly miscalibrated probabilities that make any fixed threshold misleading.
  • Subgroup analysis. Performance by age, sex, race and ethnicity where recorded, language, insurance status and care setting — because the published evidence usually will not supply it.
  • Alert burden modelling. Convert sensitivity and positive predictive value into the number of alerts a nurse will see per shift, and the number needed to evaluate to find one true case.
  • Threshold selection as a clinical decision. The vendor's default threshold encodes an assumption about the relative cost of false positives and false negatives at another institution.
  • Workflow and human-factors testing. Where does the output appear, who acts on it, what does the clinician do when they disagree, and is that disagreement recorded?

The canonical cautionary case remains sepsis prediction. External validation across 38,455 hospitalisations at Michigan Medicine found the widely deployed Epic Sepsis Model achieved an AUC of 0.63, with 33% sensitivity and 12% positive predictive value at the developer-suggested threshold (Wong et al., 2021). A later external validation in two county emergency departments in Texas, covering 145,885 encounters in 2023, reported 14.7% sensitivity and 7.6% positive predictive value within a six-hour window (Ostermayer et al., 2024) — a reminder that performance is setting-specific, not merely vendor-specific.

The follow-up is instructive in a different way. A multicentre prospective validation of the second-generation model across four large US health systems, covering 227,091 inpatient encounters and published on 27 February 2026, found materially improved discrimination — encounter-level AUROC between 0.82 and 0.92 — but high variability between institutions, positive predictive values between 0.13 and 0.26, and a correspondingly high alert burden (Wong et al., 2026). The model got better; the case for validating it locally, tuning thresholds and designing alert-silencing strategies did not go away. That is the pattern professionals should expect from this technology generally: improvement in average performance does not remove site-level variance.


Three archetypes, three evidence profiles

Treating "clinical AI" as one category is the most common analytical error in procurement. The risk profile, evidence base and monitoring requirements differ fundamentally.

Table 2. Deployment archetypes compared

DimensionAutonomous diagnostic (e.g. diabetic retinopathy screening)Assistive detection (e.g. computer-aided polyp detection)Ambient documentation (AI scribes)
Clinician roleNot required for the diagnostic outputClinician decides; AI promptsClinician reviews and signs the note
Typical regulatory statusClass II device, De Novo/510(k)Class II deviceFrequently outside device regulation
Strongest evidenceProspective pivotal trial: 87.2% sensitivity, 90.7% specificity (Abràmoff et al., 2018)Meta-analysis of randomised trials: ~8 percentage-point absolute gain in adenoma detection (Soleymanjahi et al., 2024, as summarised in Ahmad, 2025)One randomised trial; mixed results (Lukac et al., 2025)
Documented downsideUngradable images; access dependent on camera and staffingDeskilling signal in unassisted procedures (Budzyń et al., 2025)"Occasional" clinically significant inaccuracies reported by users
ReimbursementCPT 92229; Medicare national payment ~US$40–47 (2022–2024)Bundled into the procedureNone specific; justified on time and retention
Monitoring priorityImage quality and referral follow-throughDetection rates with and without assistanceNote accuracy audits and edit rates

Table compiled by OneWise; figures as reported in the cited sources.

Autonomous diagnosis is the most mature archetype. The pivotal trial of the system now marketed as LumineticsCore enrolled 900 participants in primary care and reported 87.2% sensitivity and 90.7% specificity against a reading-centre reference standard (Abràmoff et al., 2018). Reimbursement followed: CPT code 92229 for point-of-care autonomous retinal analysis carried a Medicare national payment of roughly US$47 in 2022, US$46 in 2023 and US$40 in 2024 (Teng et al., 2026). The instructive lesson is that a reimbursement code did not by itself produce adoption; sites still had to solve camera placement, operator training, image quality and referral follow-up.

Assistive detection shows how a well-evidenced intervention can produce a second-order effect. Randomised evidence supports computer-aided detection in colonoscopy. But an observational study across four Polish centres found that among 19 experienced endoscopists, the adenoma detection rate in unassisted colonoscopies fell from 28.4% before routine AI exposure to 22.4% afterwards — a six percentage-point absolute reduction (Budzyń et al., 2025). The design was observational and cannot establish causation; secular trends and case-mix shifts are plausible alternatives. Even so, it is the first real-world clinical signal of deskilling, and it argues for measuring unassisted performance rather than assuming it is preserved.

Ambient documentation is where enthusiasm currently outruns evidence. In a pragmatic three-arm randomised trial at a large Californian academic system, 238 outpatient physicians across 14 specialties were assigned to one of two ambient scribe products or usual care between 4 November 2024 and 3 January 2025. Only one of the two products produced a statistically significant reduction in time-in-note (−9.5%; 95% CI −17.2% to −1.8%; p = 0.02); both were associated with improvements in burnout-related survey measures, and users reported clinically significant inaccuracies "occasionally" (Lukac et al., 2025). Uptake was itself a limiting factor: the tools were used in roughly a third of eligible visits. The authors call for confirmation in larger multicentre trials — an appropriate reading of secondary endpoints in a single-site study.


After go-live: monitoring, drift and human factors

A model that passed validation in March can fail quietly in September. Upstream changes — a new laboratory analyser, a revised documentation template, a coding policy change, a shift in case mix — alter the input distribution without any change to the model.

Practical monitoring rests on four measurement families:

  1. Input monitoring: distribution of key features versus the validation period; missingness rates; unexpected new value codes.
  2. Output monitoring: score distribution, alert volume per unit per shift, threshold crossing rates.
  3. Interaction monitoring: acceptance, override and dismissal rates; time-to-action after an alert; free-text override reasons.
  4. Outcome monitoring: the clinical endpoint the deployment was justified on, plus at least one balancing measure (for example, unnecessary imaging, or unassisted performance in the assistive-detection case).

Two human-factors risks deserve standing attention. Automation bias is the tendency to accept a system's output without independent verification; the FDA's 2026 CDS guidance places notable weight on transparency of inputs and logic precisely so clinicians can independently evaluate a recommendation. Deskilling is the slower erosion of unassisted capability, now supported by at least one real-world dataset (Budzyń et al., 2025) and flagged in CHAI's education and training playbook.


Frequently misunderstood concepts

  • "FDA-cleared" ≠ "clinically validated in your population." Clearance addresses intended use and substantial equivalence, not local generalisability.
  • A high AUC does not mean a usable alert. With low disease prevalence, a strong AUC can still yield a positive predictive value that generates unmanageable false-alert volume.
  • Non-device CDS is not unregulated risk. Falling outside FDA device oversight does not remove malpractice exposure, accreditation expectations, or state-level requirements.
  • A PCCP is not a licence for silent change. It permits pre-specified modifications; organisations should still be notified and should re-run monitoring after an update.
  • "Human in the loop" is a design claim, not a safety guarantee. It works only if the clinician has the information, time and incentive to disagree.

Latest developments

  • 6 January 2026 (reissued 29 January 2026): FDA published revised final CDS guidance, superseding the 2022 version, and introduced enforcement discretion for certain software producing a single clinically appropriate recommendation where other non-device criteria are met (FDA, 2026b). Industry response has been mixed; the EHR Association publicly welcomed the enforcement-discretion policy while arguing that other parts of the guidance overreach (EHR Association, 2026). Status: final guidance, non-binding.
  • 27 May 2026: CHAI released governance playbooks developed with more than 100 healthcare organisations, structured around eight elements including lifecycle management, third-party management, and education and training (CHAI, 2026). Status: published, voluntary.
  • Late May 2026: The Joint Commission launched its Responsible Use of AI in Healthcare certification, organised around governance, data management, risk and bias reduction, monitoring and validation, and transparency and training. Organisations need not be Joint Commission-accredited to apply (Joint Commission, 2026). Status: live, voluntary; certifies organisations, not products.
  • May–June 2026: EU institutions agreed and adopted the Digital Omnibus on AI, deferring high-risk obligations for AI embedded in regulated products, including medical devices, from 2 August 2027 to 2 August 2028, and Annex III standalone systems to 2 December 2027 (Gibson Dunn, 2026). Status: adopted; readers should confirm Official Journal publication and any subsequent MDR amendments before relying on the dates.
  • 27 February 2026: The first multicentre prospective validation of the second-generation Epic Sepsis Model, across four US health systems and 227,091 encounters, reported improved discrimination alongside high institutional variability, low positive predictive value and high alert burden, and concluded that implementing institutions should still validate locally (Wong et al., 2026). Status: peer-reviewed prognostic study.
  • March 2026: The AMA reported 81% physician AI use in its 2026 survey wave (AMA, 2026). Status: self-reported survey data, voluntary participation — indicative of direction rather than a precise adoption measure.

Practical takeaways

For clinical and quality leaders

  1. Require a silent evaluation on local data before any AI output reaches a clinician, and define in advance the performance floor at which you will not proceed.
  2. Set thresholds yourself, and document the false-positive/false-negative trade-off you accepted.
  3. Add one balancing measure to every deployment, including unassisted clinician performance for detection tools.

For informatics and technical teams 4. Instrument input, output, interaction and outcome monitoring before go-live, not after the first incident. 5. Treat vendor model updates as change-control events: ask what the PCCP allows, require notification, and re-run monitoring afterwards.

For procurement, legal and compliance 6. Ask for subgroup performance, the training population description, and any recall history in writing; absence of an answer is itself information. 7. Do not treat FDA clearance status as a sufficient quality signal, particularly for tools now falling under enforcement discretion.

For executives 8. Fund the unglamorous half of the programme — validation, monitoring, training and governance — at the same time as the licence, or the deployment will silently degrade. 9. Use the CHAI playbooks and Joint Commission certification standards as a maturity model even if you do not pursue certification.


Key insights

  1. Regulatory authorisation establishes market entry, not local clinical benefit.
  2. Most clinical AI is cleared without prospective human-in-the-loop testing.
  3. Public documentation rarely reports subgroup performance, so buyers must ask.
  4. External validation has repeatedly shown large drops from developer-reported performance.
  5. Threshold selection is a clinical and ethical decision, not a technical default.
  6. Alert burden, not AUC, determines whether frontline staff can use a tool.
  7. Recalls cluster early after authorisation, making the first year post-deployment high-risk.
  8. Well-evidenced tools can still cause harm through second-order effects such as deskilling.
  9. Ambient documentation has promising but limited randomised evidence and uneven uptake.
  10. Governance is shifting from regulators to organisations, with accreditation bodies formalising the expectation.

Frequently asked questions

What is local validation of clinical AI? Testing a model's discrimination, calibration, subgroup performance and alert burden on the deploying organisation's own patient data and workflow before it influences care, then repeating that assessment periodically.

Does FDA clearance mean an AI tool improves patient outcomes? No. Clearance provides reasonable assurance of safety and effectiveness for a stated intended use. Outcome improvement requires separate clinical evidence, which is often absent at the time of authorisation.

How many AI-enabled medical devices has the FDA authorised? More than 1,400 as of the agency's 2026 data updates, with roughly three-quarters in radiology. Counts differ between snapshots and analyses, and the FDA states the list is not comprehensive.

What is a silent or shadow trial? A period during which the model runs on live data and its outputs are recorded but hidden from clinicians, allowing performance and alert volume to be measured without affecting care.

What is a Predetermined Change Control Plan? An FDA-authorised plan specifying in advance how a manufacturer may modify an AI model post-market without a new submission, including the modifications, methods and acceptance criteria.

Did the FDA reduce oversight of clinical decision support in 2026? It clarified the boundary and extended enforcement discretion to certain single-recommendation CDS meeting other non-device criteria. The practical effect is that fewer tools carry an FDA authorisation for buyers to reference.

Is the Joint Commission AI certification mandatory? No. It is voluntary, open to non-accredited organisations, and certifies organisational practices rather than individual AI products.

When do EU AI Act high-risk rules apply to medical devices? Following the 2026 Digital Omnibus, obligations for AI embedded in regulated medical devices are set to apply from 2 August 2028. Confirm the current position in the Official Journal before relying on this date.

What is model drift? Degradation in performance caused by changes in the input data or the clinical environment — new equipment, altered documentation practices, or a shift in patient population — without any change to the model itself.

Do AI scribes reduce documentation time? Evidence is mixed. In one randomised trial, one of two products significantly reduced time-in-note while both were associated with improved burnout-related measures; effects depend heavily on how often clinicians actually use the tool.

Can AI make experienced clinicians worse? One observational study found lower adenoma detection in unassisted colonoscopies after routine AI exposure. This is a single real-world signal and cannot establish causation, but it justifies monitoring unassisted performance.

Who is liable when an AI recommendation contributes to harm? Liability frameworks remain unsettled and jurisdiction-specific. In most current arrangements the treating clinician retains professional responsibility, with contractual allocation between organisation and vendor determined case by case. This is a legal question requiring qualified advice.

What should we ask a vendor before purchase? Training population characteristics, subgroup performance, calibration data, recommended and defensible thresholds, PCCP scope, update notification practice, recall history and monitoring support.

Is more data always better for clinical AI? No. Representativeness, label quality and relevance to the deployment setting matter more than volume; large but unrepresentative datasets can entrench performance gaps.


Glossary

510(k) — US premarket submission demonstrating substantial equivalence to a legally marketed predicate device. AUC (area under the receiver operating characteristic curve) — A measure of a model's ability to rank cases above non-cases; 0.5 is chance, 1.0 is perfect. Automation bias — The tendency to over-accept automated outputs and under-weight contradictory evidence. Calibration — The agreement between predicted probabilities and observed event rates. CDS (clinical decision support) — Software providing information or recommendations to support clinical decisions. CPT — Current Procedural Terminology, the AMA-maintained US procedure coding system used for billing. De Novo — US pathway for classifying novel low-to-moderate-risk devices without a predicate. Drift — Performance degradation caused by changes in input data or context over time. GMLP (Good Machine Learning Practice) — Guiding principles for developing safe, effective medical-device machine learning. MDR/IVDR — EU Medical Device Regulation and In Vitro Diagnostic Regulation. PCCP (Predetermined Change Control Plan) — Pre-authorised plan for post-market model modifications. PPV (positive predictive value) — Proportion of positive alerts that are true positives; depends on prevalence. SaMD (Software as a Medical Device) — Software intended for a medical purpose that is not part of a hardware device. Silent/shadow deployment — Running a model on live data with outputs hidden from clinicians for evaluation. TPLC (Total Product Life Cycle) — Regulatory approach spanning design, deployment and post-market monitoring


References

Academic papers

Abràmoff, M. D., Lavin, P. T., Birch, M., Shah, N., & Folk, J. C. (2018). Pivotal trial of an autonomous AI-based diagnostic system for detection of diabetic retinopathy in primary care offices. npj Digital Medicine, 1, 39. https://doi.org/10.1038/s41746-018-0040-6

Budzyń, K., Romańczyk, M., Kitala, D., Kołodziej, P., Bugajski, M., Adami, H.-O., … Mori, Y. (2025). Endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy: A multicentre, observational study. The Lancet Gastroenterology & Hepatology, 10(10), 896–903. https://doi.org/10.1016/S2468-1253(25)00133-5

Lee, B., Kramer, P., Sandri, S., Chanda, R., Favorito, C., Nasef, O., Ross, J. S., Sharfstein, J., & Dai, T. (2025). Early recalls and clinical validation gaps in artificial intelligence–enabled medical devices. JAMA Health Forum, 6(8), e253172. https://doi.org/10.1001/jamahealthforum.2025.3172

Lukac, P. J., Turner, W., Vangala, S., Chin, A. T., Khalili, J., Shih, Y.-C. T., Sarkisian, C., Cheng, E. M., & Mafi, J. N. (2025). Ambient AI scribes in clinical practice: A randomized trial. NEJM AI, 2(12). https://doi.org/10.1056/AIoa2501000

Ostermayer, D. G., Braunheim, B., Mehta, A. M., Ward, J., Andrabi, S., & Sirajuddin, A. M. (2024). External validation of the Epic sepsis predictive model in 2 county emergency departments. JAMIA Open, 7(4), ooae133. https://doi.org/10.1093/jamiaopen/ooae133

Sivakumar, R., Lue, B., & Kundu, S. (2025). FDA approval of artificial intelligence and machine learning devices in radiology: A systematic review. JAMA Network Open, 8(11), e2542338. https://doi.org/10.1001/jamanetworkopen.2025.42338

Soleymanjahi, S., Huebner, J., Elmansy, L., et al. (2024). Artificial intelligence-assisted colonoscopy for polyp detection: A systematic review and meta-analysis. Annals of Internal Medicine, 177(12), 1652–1663.

Teng, C. W., Patel, S. D., Barkmeier, A. J., Liu, T. Y. A., Myung, D., Henderer, J., Liu, J., Hansen, E., & Al-Aswad, L. A. (2026). Autonomous artificial intelligence in diabetic retinopathy testing — Lessons learned on successful health system adoption. Ophthalmology Science, 6(1), 100935. (Published online 3 September 2025.)

Windecker, D., Baj, G., Shiri, I., et al. (2025). Generalizability of FDA-approved AI-enabled medical devices for clinical use. JAMA Network Open, 8(4), e258052. https://doi.org/10.1001/jamanetworkopen.2025.8052

Wong, A., Otles, E., Donnelly, J. P., Krumm, A., McCullough, J., DeTroyer-Cooley, O., … Singh, K. (2021). External validation of a widely implemented proprietary sepsis prediction model in hospitalized patients. JAMA Internal Medicine, 181(8), 1065–1070. https://doi.org/10.1001/jamainternmed.2021.2626

Wong, A., Currey, D., Schwinne, M., et al. (2026). Multicenter prospective validation of an updated proprietary sepsis prediction model. JAMA Network Open, 9(2), e260181. https://doi.org/10.1001/jamanetworkopen.2026.0181

Official documentation and government sources

U.S. Food and Drug Administration. (2024). Marketing submission recommendations for a predetermined change control plan for artificial intelligence-enabled device software functions: Guidance for industry and FDA staff.

U.S. Food and Drug Administration. (2025). Artificial intelligence-enabled device software functions: Lifecycle management and marketing submission recommendations — Draft guidance for industry and FDA staff. Federal Register, 7 January 2025.

U.S. Food and Drug Administration. (2026a). Artificial intelligence-enabled medical devices [Device list]. https://www.fda.gov/medical-devices/software-medical-device-samd/artificial-intelligence-enabled-medical-devices

U.S. Food and Drug Administration. (2026b). Clinical decision support software: Guidance for industry and Food and Drug Administration staff (issued 29 January 2026, superseding the 6 January 2026 version). https://www.fda.gov/media/109618/download

Standards, accreditation and industry guidance

Coalition for Health AI. (2026, May 27). CHAI releases comprehensive governance playbooks to streamline AI implementation for health systems. https://www.chai.org

Joint Commission & Coalition for Health AI. (2025, September 17). Guidance on the responsible use of AI in healthcare (RUAIH).

Joint Commission. (2026). Responsible Use of AI in Healthcare (RUAIH) certification. https://www.jointcommission.org

Industry reports and professional bodies

American Medical Association. (2026). 2026 physician survey on augmented intelligence. AMA Center for Digital Health and AI. https://www.ama-assn.org/practice-management/digital-health/physician-survey-augmented-intelligence

EHR Association. (2026, July 1). FDA's revised CDS guidance: Opportunities and the path forward for EHR developers.

Gibson Dunn. (2026). EU AI Act omnibus agreement — Postponed high-risk deadlines and other key changes.

MedTech Dive. (2026). AI in medtech is booming: Track new devices here [Database analysis of the FDA AI device list, downloaded 11 May 2026].

Commentary

Ahmad, O. (2025). Endoscopist deskilling: An unintended consequence of AI-assisted colonoscopy? The Lancet Gastroenterology & Hepatology, 10(10).


Editorial note: This article is based on published research, official regulatory documents and publicly reported implementations. It does not draw on unpublished institutional experience. Regulatory positions described here were current at the time of writing (August 2026) and should be verified against primary sources before use in compliance decisions. Nothing here constitutes legal, regulatory or clinical advice.

One Tech & AI · Sunday, August 2, 2026 · 24 min read

Advanced Diagnostic Solutions – Clinical technologies use modern medical devices, imaging systems, and laboratory equipment to improve diagnostic accuracy and support faster clinical decision-making.

Enhanced Patient Care – Smart monitoring systems, digital health records, and connected medical technologies enable personalized treatment, continuous monitoring, and improved patient outcomes.

Efficient Healthcare Operations – Automation, AI-powered tools, and integrated clinical systems streamline hospital workflows, reduce errors, and increase the efficiency of healthcare services.

The centre of gravity in clinical AI has moved. For most of the last decade the interesting question was whether a model could match expert performance on a benchmark. That question is now largely settled for a wide class of narrow tasks, and the FDA's authorisation list is the visible evidence of it. The harder question — the one that determines whether patients benefit — is what happens in the six months after a contract is signed.

The published record supports a measured position. Autonomous retinal screening shows that a well-specified task, a prospective trial and a reimbursement pathway can produce durable clinical value. Sepsis prediction shows how far real-world performance can fall short of developer figures. Colonoscopy shows that a genuinely effective tool can still produce an unwelcome second-order effect. Ambient documentation shows a promising technology whose randomised evidence base is still one trial deep. None of these is a verdict on clinical AI as a whole, and treating them as one is the error to avoid.

Important uncertainties remain. There is no consensus method for continuous post-deployment performance monitoring, no standard liability allocation between developers and deploying organisations, and limited evidence on how AI-assisted workflows affect long-term clinical skill. Independent assurance infrastructure is emerging but unproven.

What the evidence does support is unglamorous and actionable: the organisations most likely to benefit are those treating clinical AI the way they already treat a new drug protocol or surgical device — with a defined intended use, a local evidence check, a named owner, measured outcomes, and a plan to stop. Validation is not a hurdle before the value arrives. In clinical AI, it is where the value is actually produced.

TOPIC

Health Tech