A randomized noninferiority trial published in Nature Communications compared human-only prescreening against human-plus-AI augmentation using a pretrained language model. Across 355 patients with non-small-cell lung or colorectal cancer, chart-level accuracy came out at 76.1% with AI assistance versus 71.5% without — a superiority p of 0.002.

That is the strongest randomized evidence in the field. A 4.6-point accuracy gain. And it comes with two qualifications that rarely survive the retelling: the trial is randomized but its charts were collected retrospectively, and the frequently quoted claim that this translates to “10–20 additional patients screened per week” appears nowhere in the paper.

Now the other end of the distribution. A 2025 scoping review in Digital Health screened the clinical-trials literature and included 108 records. Of those, two studies had implemented machine learning for site selection. Ninety-three trials — about 86% — used ML for analysis. Nineteen per cent touched trial design, 6% patient selection, 2% site selection.

Set those two findings beside the numbers that circulate in vendor decks — “3× faster enrollment,” “50% reduction in site selection timelines,” “30–50% timeline acceleration,” “1.7× more patients” — and the shape of this field becomes clear. The peer-reviewed evidence is small, retrospective, and concentrated in one workflow. The narrative is large, self-reported, and everywhere.

This is an evidence audit.

 EVIDENCE MATURITY: WHAT IS ACTUALLY ESTABLISHED
 (bars denote maturity of evidence, not effect size)

 Retrospective chart accuracy gain  ████████████  randomized, N=355, 4.6 pts
 Screening-time reduction           █████████░░░  pilot user study, 2 physicians
 Pooled screening sensitivity       ████████░░░░  meta-analysis; 0.70-0.92 by stage
 ML used for site selection         ██░░░░░░░░░░  2 of 108 records
 Prospective outcome improvement    ░░░░░░░░░░░░  none identified

What genuinely holds up

Patient–trial matching is the one mature application. The reference system is TrialGPT (Jin et al., Nature Communications 15, 9074), and its reported figures verify individually:

Figure Status
Criterion-level accuracy 0.873 Verified — against expert performance of 0.887–0.900
>90% recall of relevant trials using <6% of the collection Verified
Screening time reduced 42.6% Verified — but a pilot user study with 2 physicians
Ranking/exclusion 43.8% better than baselines Verified
Sentence-level explanations Verified — sentence-location F1 88.6%
Evaluation on 1,015 patient–criterion pairs, 3 domain experts Verified
“ChatGPT-4.0 alone achieved 88% agreement with physicians” Not in the paper — a secondary-source distortion

That last row is the pattern in miniature. The real TrialGPT number is a model at 0.873 against experts at 0.887–0.900 — honest, close, and defensible. The version that circulates is a bare 88% attributed to the wrong system.

One more distinction worth carrying: the same paper’s 183 synthetic patients and >75,000 trial-eligibility annotations are not a contradiction of the 1,015 figure. They are the overall evaluation across three public cohorts; the 1,015 pairs are the manually annotated subset used to score the matching module.

The pooled evidence, read correctly. A 2026 scoping review and meta-analysis (J Clin Transl Sci 10(1):e98) screened 21,573 records and included 121 studies, reporting pooled screening sensitivity of 0.91 (95% CI 0.84–0.95). The number that gets dropped is the heterogeneity, which is the useful part: by workflow stage, eligibility identification is 0.80, eligibility classification 0.92, and cohort identification only 0.70. “Sensitivity 0.91 for AI in recruitment” averages tasks that differ by more than twenty points.

And one real-world result worth noting. Paradigm Health, presented at the 2025 ASCO Quality Care Symposium, ran LLM reasoning over full-text clinical notes against rule-based matching across four cancer trials in a large community oncology practice. In three trials it reduced the number of patients surfaced for screening by 31–87% — fewer charts reviewed, same eligible patients. In the fourth, the patient count went up 32% and produced a doubling of eligible patients after human screening. Fewer false positives in three cases, fewer false negatives in the fourth, which is exactly the trade-off you would expect and rarely see reported honestly.

The provenance problem

This is where the field diverges most sharply from every other AI domain I’ve audited. Four statistics repeated across consultancy reports, vendor pages and trade press have no traceable source at all:

Claim Status
ML accelerates trial timelines 30–50%, cuts costs up to 40%, improves enrollment 65% Traces to a 2026 narrative review that states them with no primary citation, echoed verbatim by vendors
Site selection improves forecast accuracy 45%→75%; feasibility planning 8–12 weeks→4–6 weeks No source exists; vendor pages give different numbers
68% of clinical trial data is unstructured No trace; likely a conflation with “68% of sites fail enrollment goals”
EHR integration costs $250K–$500K per health system over 6–8 months No trace of this pairing anywhere

Two more are false as stated:

“Novartis and IBM Watson improved screening accuracy from 45% to 78%, cutting screening time 94% (8 hours to 30 minutes).” The real study (Haddad et al., JCO Clinical Cancer Informatics 2020) reports the “78%” as a time reduction — 110 minutes to 24 minutes across 90 patients in three trials. Accuracy was sensitivity 91–95% / 46.7% and specificity 76–99%. The 94%, the 8-hour figure and the 30-minute figure do not exist.

“Lo et al., PNAS 2020, achieved ~93% AUC predicting trial success.” The paper is Lo, Siah & Wong in Harvard Data Science Review 1.1 — October 2019, not PNAS, not 2020 — and the reported AUC is 0.78 (Phase 2→approval) and 0.81 (Phase 3→approval). The benchmark caveat matters as much as the correction: the evaluation was k-fold plus a held-out set from the same 2003–2015 dataset, retrospectively, at drug–indication level, and the “features” include trial outcome, status, accrual and duration — which are not design features.

And one that inverts completely. The canonical cautionary tale — “IBM Watson for Clinical Trial Matching failed at Mayo, proving AI-in-trials failures are data-plumbing failures, not algorithmic ones” — is wrong. Mayo’s clinical-trial-matching project was reported a success, with an 80% increase in breast-cancer trial enrollment. The famous failure was Watson for Oncology, a different product, in 2018. The two have been conflated for years, and the conflation has been doing rhetorical work for a narrative the primary record does not support.

Site selection: vendor percentages vs two retrospective studies

The market-facing numbers here are all self-reported:

Claim Nature
Amgen ATOMIC: 3× faster enrollment at top-ranked sites Amgen’s own story (Apr 2025); internal analysis of 13 studies
AbbVie: >50% more sites meeting enrollment goals Conference talk (2023); no published method
McKinsey: top-enrolling sites +30–50%, enrollment +10–15% Consultancy analysis of pilots
Parexel: 50% faster site selection Vendor press release
Novartis: 4–6 weeks compressed to a 2-hour meeting Company claim via Reuters
IQVIA: 33% faster startup, 1.7× patients Vendor press release

The peer-reviewed work is thinner and more instructive.

FRAMM (Patterns 2024) is the fairness-aware exception: up to +20% minority enrollment for some groups, +14% Hispanic, +27% Black, +60% Asian, with diversity improved 9% at similar enrollment. But it was evaluated on 4,392 historical trials, and the paper describes itself as “a case study and proof of concept.”

The site-risk model (Int J Med Inform 2026) is the best illustration of why validation splits matter more than headline accuracy. It reports 98.91% accuracy, correctly classifying 91 of 92 sites. That 92-site set is an internal held-out split. The external validation — 761 sites across 89 Duchenne muscular dystrophy studies — comes in at 81.21%. The paper itself acknowledges there was no ground-truth outcome validation.

Read that as the general rule for this entire field: a number measured on held-out data from the same programmes is a fit statistic. The external number is the performance estimate.

Design side: one genuinely strong result, one fabricated breakdown

Trial Pathfinder (Nature 592:629–633, 2021 — not 2022) is real and important: 10 Phase III NSCLC trials compared against 61,094 Flatiron patients, showing the eligible pool “more than doubled” when lab-value exclusions were relaxed, with safety supported across 22 Roche trials. But the popular gloss — “without changing hazard ratios” — is wrong. The data-driven broadening decreased OS hazard ratios by about 0.05 on average. Small, directionally favourable, and not zero.

TrialMap (ISPOR 2026 abstract, not peer-reviewed) reports that original criteria retained only 18–43% of real-world patients across 15 first-line oncology trials. Lee et al. (JMIR AI 2024) built an eligibility ontology from 3,281 phase 2/3 trials and a Bi-LSTM-CRF model at precision 0.91 / recall 0.79 / F1 0.83. AutoTrial (arXiv:2305.11366) was evaluated on — not trained on — over 70,000 trials, with roughly a 60% win rate against GPT-3.5.

The strongest design-side result is AMARANTH. The lanabecestat trial was declared futile. Re-stratifying the same patients with an interpretable prognostic model using only baseline data revealed a 46% slowing of cognitive decline in slow-progressing patients. That is a real signal rescued from a failed trial by better stratification — and the most compelling argument in the corpus that AI belongs in trial design rather than only in operations.

And a fabrication worth naming. The widely cited finding that “50% of front-line leukemia-trial eligibility criteria did not align with known drug safety profiles, 30% were overly permissive and 10% overly restrictive” does not appear in the Haematologica paper it is attributed to. What that paper actually reports is criterion-specific concordance — bilirubin 68.4%, renal 68.4%, AST/ALT 58.8%, HIV 54.8%, hepatic B/C 42.0%/41.2%, ejection fraction 34.8%, QTc 22.4% — and that criteria limits were more restrictive than safety data in 51–76% of trials. The real finding is arguably more damning than the invented one. Someone tidied it into a memorable split, and it has been propagating ever since.

The regulatory position, corrected

Five dates here are commonly wrong.

Instrument Correct position
FDORA diversity action plans The requirement is real (Dec 2022). ⚠ The draft guidance is June 2024, and FDA withdrew it in January 2025 — it was never finalized. Do not treat DAP guidance as live
FDA eligibility criteria guidance ⚠ “Enhancing Participation in Clinical Trials,” December 2025 — FINAL (15 Dec 2025), superseding the 2020 diversity guidance. No eligibility guidance was issued in September 2024. The likely confusion is the April 2024 cancer-eligibility drafts, finalized July 2026
ICH E6(R3) Step 4 adoption 6 January 2025 (Principles + Annex 1); Annex 2 (decentralized elements, real-world data) adopted 3 June 2026
ICH M11 (electronic structured protocol) ⚠ Adopted 19 November 2025; FDA’s Federal Register notice of availability was 22 May 2026. Not “finalized May 2026”
FDA computerized systems guidance ⚠ A trap. The 1999 “Computerized Systems Used in Clinical Trials” guidance was superseded by the May 2007 guidance on computerized systems used in clinical investigations. Citing 1999 as current is a live error
PROCOVA EMA CHMP qualification opinion, 20 September 2022 — correct. ⚠ “FDA concurrence” is overstated: FDA declined the ISTAND letter of intent, saying only that it concurs PROCOVA is a special case of ANCOVA and “does not appear to deviate” from guidance. And the “40% control-arm reduction” is Unlearn’s own Alzheimer’s figure, not in the EMA opinion, which cites up to 15% sample-size reduction and is disease-generic
FDA AI real-time review pilot Announced 28 April 2026, RFI 29 April 2026, explicitly covering “improve patient recruitment” and “enhance safety monitoring.” Two proof-of-concept trials running with AstraZeneca and Amgen

The direction of travel is genuinely encouraging — E6(R3), M11, the decentralized-trials guidance and the AI pilot all point toward structured protocols and risk-proportionate oversight. But the specific documents are being cited with dates and statuses that don’t survive a Federal Register check, which is a problem when the same claims get recycled into validation plans.

And the one structural implication for anyone building here: identification is not the binding constraint. Consent willingness, logistics, trust and eligibility documentation are — missing labs, unavailable biopsies, patients who never reach the site. Matching tools that only shorten the chart-review step run into a different wall. Which is why the credible programmes pair matching with decentralized elements rather than treating matching as the whole solution.

Two final cautions. The fairness impossibility results are real but narrower than advertised — calibration and balance conditions cannot coexist except trivially (Kleinberg et al.), and calibration with equal false-positive and false-negative rates is incompatible under unequal prevalence (Chouldechova, 2017). It is not true that “all four fairness definitions are mathematically incompatible,” and SPIRIT-AI contains no concept of “two-level selection criteria” for AI-assisted patient selection. And retention models should be used to provide support — transport, remote visits, coordinator contact — not to exclude patients, which would reproduce the very disparities these models are meant to correct.

If you are buying or building

  1. Ask for the external validation split before the accuracy figure. 98.91% internal becomes 81.21% external. That is the field’s defining ratio.
  2. Separate randomized from prospective. The strongest trial here is randomized on retrospectively collected charts. That is meaningful design rigour attached to weak temporal validity, and it should be described as such.
  3. Treat the vendor percentages as hypotheses. Amgen, AbbVie, Parexel, IQVIA, McKinsey and Novartis figures are self-reports; none is independently validated. Label them when you use them.
  4. Trace the training-data rights before the model. Site EHR access is the actual moat and the actual blocker. Most public AI-in-trials failures in this space were data-access failures wearing an algorithm costume.
  5. Validate against context of use, not “the model.” Document triage for a data manager and eligibility ranking for a coordinator are different risk classes with different evidence burdens.
  6. Measure the metrics that survive process change: screen-failure rate, time-to-first-patient-in, queries per patient, deviation rate, enrolment diversity. Anything else is activity, not outcome.
  7. Audit the claims you inherit. Two of the numbers most likely to appear in your own business case — the 30–50% timeline figure and the 45%→75% accuracy figure — have no source behind them.

The bottom line

ML in clinical trials is real, valuable, and smaller than its own marketing. The mature application is patient–trial matching, where the evidence is randomized, retrospective, and worth about four and a half accuracy points. The design applications — stratified rescue of failed trials, criteria impact modelling, site risk scoring — are genuinely promising and almost entirely retrospective. Everything else in the deck is a self-report.

The pattern connecting this to the rest of this series is uncomfortable and consistent: in molecular design, virtual cells, and now clinical trials, the generative and predictive capacity has run ahead of the verification capacity by an order of magnitude — and the gap is filled by numbers that nobody has checked.

The unglamorous work is the same everywhere. An external cohort. A calibration report. A provenance trail. A named human accountable for the decision the model informed. For a field where the pooled sensitivity is 0.91 and the honest effect is 4.6 points, that discipline is the entire difference between a capability and a claim.


Related: Two of Eighty-Six: The Validation Deficit in AI Biomarker Discovery · Representation Is Solved, Simulation Isn’t · 117 AI-Built Drug Programs, Zero Approvals

Sources: TrialGPT, Nature Communications 15, 9074 (2024) · randomized prescreening trial, Nature Communications 17, 2306 (2026) · Bian et al., JCO Clinical Cancer Informatics 2025;9:e2500071 · Yin et al., J Clin Transl Sci 2026;10(1):e98 · Paradigm Health, JCO Oncology Practice 21(suppl 10):603 (ASCO Quality Care Symposium 2025) · Kanapari et al., Digital Health 2025;11:20552076251393272 · Teodoro et al., npj Digital Medicine 2025;8:486 · Wójcik et al., Healthcare (Basel) 2026;14(16):2519 · FRAMM, Patterns 5(3), 2024 · Yang, International Journal of Medical Informatics 2026;211:106314 · Bieganek et al., PLOS ONE 17(2):e0263193 (2022) · DocTr, npj Health Systems 2026 · Trial Pathfinder, Nature 592:629–633 (2021) · TrialMap, ISPOR 2026 · Lee et al., JMIR AI 2024;3:e50800 · AutoTrial, arXiv:2305.11366 · Journal of Thoracic Oncology 2017;12(10):1489–1495 · Haematologica (2024) · Scientific Reports 13:121 (2023) · AMARANTH, Nature Communications 16:6244 · PROCOVA EMA CHMP qualification opinion, 20 Sep 2022 · Getz et al., Therapeutic Innovation & Regulatory Science 2016;50(4):436–441 · Huang et al., Contemporary Clinical Trials 2018 (CTTI) · Tufts CSDD 2013 release and 2024 update (Ther Innov Regul Sci 58:696) · Idnay et al., JAMIA 29(1):197–206 · Davoudi et al. 2024 · Olawade et al., International Journal of Medical Informatics 2026;206:106141 · Cummings 2022, Alzheimer’s & Dementia · Haddad et al., JCO Clinical Cancer Informatics 2020 · Lo, Siah & Wong, Harvard Data Science Review 1.1 (2019) · Kleinberg et al., arXiv:1609.05807 · Chouldechova, Big Data 2017 · FDA: 89 FR 54012; Dec 2025 eligibility guidance; 91 FR 30310 (M11); 91 FR 23100 (AI RFI); 72 FR 26638 (2007 computerized systems) · EMA/CHMP/CVMP/83833/2023 · ICH E6(R3) (6 Jan 2025; Annex 2, 3 Jun 2026) · NMPA GCP No. 50/2026 · Amgen, McKinsey, Parexel, IQVIA, Reuters and AbbVie announcements as cited inline.

Research notes: [[ML-Clinical-Trials-Evidence-Audit-2026]]