Five language models refine a clinical protocol by arguing with each other. A Supervisor agent dispatches work; a Trialist agent drafts the design; an Informatician maps it onto real-world data; a Clinician agent pressure-tests the eligibility criteria; a Statistician agent interrogates the analysis plan. They iterate, critique and rewrite until the protocol holds together.
That system is called EmulatRx, it was published in Nature Communications in July 2026, and it is the most technically serious thing in the applied layer of AI in clinical trials. It was also, as the authors state plainly, evaluated as retrospective case studies on existing real-world data — acute conditions against MIMIC-IV, and Alzheimer’s and Parkinson’s across five New York City health systems. No prospective trial has ever run on a protocol it produced.
Both halves of that paragraph are the story. When you check the applied layer of AI in clinical trials — what is actually deployed, in production, at named organisations — the pattern is consistent and almost mechanical:
Capability claims verify. Magnitude claims do not.
The systems are real. The announcements are real. The percentages attached to them, with a handful of honourable exceptions, are company self-reports that either have no methodology behind them or no source at all. And a large block of the “AI in clinical trials” statistics circulating in 2026 traces back to a single narrative review that states them without citing a primary study, after which every vendor deck in the field repeats them verbatim.
This is what the deployed layer looks like when you take it apart.
EVIDENCE TIER BY LAYER OF THE APPLIED STACK
Protocol as structured data ── peer-reviewed EmulatRx, Nat Commun 2026
│ (retrospective evaluation)
▼
Computable eligibility ── peer-reviewed Lee 2024 (P 0.91 / R 0.79)
│
▼
Cohort simulation / TTE ── peer-reviewed Nat Commun 2023, 2025
│ (causal framework, RWD)
▼
Patient–trial matching ── peer-reviewed + one randomized trial on
│ retrospectively held charts
▼
Site ranking ── 2 of 108 studies; everything else is vendor
│
▼
Monitoring / RBQM ── standard of care; the attached percentages
are vendor self-reports
The protocol is now a target the models can iterate against
The reason agentic protocol refinement became possible in 2026 is not that the models got smarter. It is that the protocol started becoming machine-readable.
ICH M11 — the Clinical Electronic Structured Harmonised Protocol — was adopted on 19 November 2025, with FDA’s notice of availability following on 22 May 2026. M11 is not just a document format; it is a guideline plus a standardised template plus a technical specification, and its stated purpose is to enable electronic exchange of protocol information. ICH E6(R3) sits alongside it, adopted 6 January 2025 with Annex 2 — covering decentralised elements and real-world data — following on 3 June 2026.
Put those two instruments next to an agentic system and the logic is clear. A protocol that exists as a structured object with typed fields can be queried, diffed, simulated against cohorts, and revised agent-by-agent. A protocol that exists as a 140-page PDF cannot.
EmulatRx’s architecture is worth understanding because it is the template the rest of the field will copy: five role-specialised agents on a graph runtime with reinforcement learning from human feedback, iterating over real-world data sources rather than over a knowledge base. The domain separation matters more than the model choice. Assigning eligibility-criteria critique to one agent and statistical-power critique to another is what prevents the single most common failure of LLM output in regulated work — a fluent document that is internally consistent and operationally impossible.
What it does not yet have is a prospective arm. And that gap is the honest headline for this entire layer.
Target trial emulation is the methodology the applied layer usually skips
If you want to use real-world data to make a causal claim about a treatment, there is a framework for doing it correctly, and it is not new.
Target trial emulation (TTE) was formalised by Miguel Hernán and James Robins in American Journal of Epidemiology (2016), and refined for practitioners in JAMA in 2022. The idea is simple and demanding: before you analyse observational data, write down the protocol of the randomised trial you would have run — eligibility, treatment strategies, assignment, follow-up window, outcome, causal contrast — and then emulate that protocol explicitly. Any deviation from the specification is a documented, auditable decision.
That framework is what the credible RWD work in trials actually rests on, and it produces results that look like science rather than like dashboards:
- Alzheimer’s drug repurposing at population scale. A high-throughput TTE across two real-world warehouses — OneFlorida EHR plus MarketScan claims, more than 170 million patients over more than a decade — ranked repurposing candidates and surfaced pantoprazole, gabapentin, atorvastatin, fluticasone and omeprazole (Nat Commun 14:8180, 2023). Retrospective, hypothesis-generating, and explicitly framed as such.
- Corticosteroids in sepsis, stratified by predicted trajectory. A two-stage design — first learn organ-dysfunction subphenotypes, then emulate a target trial within them — found the corticosteroid–mortality association varied by predicted trajectory and differed between cohorts (Nat Commun 16:4450, 2025). The authors call it a retrospective multicentre proof of concept.
Here is the practical implication, and it is the single most useful question a buyer can ask. If a vendor claims treatment-effect estimates derived from real-world data, ask for the target trial specification. A system that cannot produce one is not doing causal inference; it is doing retrieval and ranking while using causal language. The distinction is not academic — it determines whether the output can support a regulatory decision or only an internal prioritisation.
A TARGET TRIAL SPECIFICATION — WHAT TO ASK FOR
eligibility ──┬── who is in
├── who is out, and on what evidence
treatment ── what the strategies are, incl. the comparator
assignment ── how it is emulated (and what confounding it assumes)
follow-up ── start, end, censoring rules
outcome ── the causal contrast, named
analysis ── the estimator, pre-specified
Two stratification results worth citing — and how to read them
The rare study with a genuine external cohort
ARDS, npj Digital Medicine, 2021. Schwager and colleagues took 51,555 ARDS patients from the Philips eICU Research Institute database (roughly 3.18 million ICU stays), tested twelve architectures across 220 variables, and stratified patients into three outcome subpopulations: rapid death, spontaneous recovery, and long-stay. The best model was a multiclass gradient-boosting classifier.
Its reported AUC is 0.77. The exact figures are 0.768 internal and 0.751 external.
That second number should not be skimmed. In this literature, a study that reports a separate external validation cohort is close to an endangered species — and the ARDS paper is one of the very few here that does. Internal 0.768 becomes external 0.751: a modest, believable degradation, which is what a real model looks like.
The percentage that is measuring something else
Phase III prostate cancer, Algorithms, 2021. Beacher and colleagues report 76% accuracy predicting outcomes of Phase III trials. That number circulates as evidence that machine learning can forecast trial success.
Read the method and it becomes three different things at once:
| What the 76% actually is | Detail |
|---|---|
| Not external | It is internal stratified 10-fold cross-validation. Chance is 54%. |
| External validation drops it | Train on two studies, test on the third: 70% |
| Not trial-level | It predicts individual-patient 2-year good/bad responder status from baseline data — not whether the trial succeeds |
| Small | n = 3,653, three AstraZeneca Phase III bicalutamide trials, from Project Data Sphere |
The last row is the one that matters most. A tool that identifies which patients will respond is a stratification tool. A tool that predicts whether your Phase III will read out is a portfolio oracle. They are different products with different failure modes, and in vendor material they are routinely the same sentence.
And the sample-size claim is a round number
“ML-driven stratification reduces required sample sizes by 20–30%” appears everywhere. No primary source reports it.
The real published figures are inconsistent with each other and mostly larger: the ARDS paper states up to a 50% reduction; AMARANTH reports up to 90% (82 patients per arm versus 762 at 90% power); the EMA’s PROCOVA qualification opinion cites up to 15% and is disease-generic. When a claim’s number is smaller than every real measurement and tidier than all of them, it is a summary someone wrote rather than a result someone obtained.
The commercial layer: real announcements, unattached numbers
Every one of these deployments is real. The dates check out against company newsrooms. What frequently does not check out is the number in the same sentence.
| Company | What shipped | Date | Number attached | What the number is |
|---|---|---|---|---|
| ICON + Anthropic | Multi-year collaboration; Claude across the trial lifecycle in Orbis, including site intelligence and study planning (OneSearch, OnePlan) and predictive intelligence for real-time enrolment-risk detection | 28 Jul 2026 | “up to 26% increase in subject recruitment, 24% better first-patient-in” | ⚠ Self-report. “Our customers have experienced.” No methodology, no sample, no named customer, no independent validation |
| TCS ADD RBQM | Risk-based quality management platform, four AI/ML modules (RACT, QTL, Clinical Trial Analytics, Subject Data Analytics) | 24 Nov 2025 | 30% monitoring efficiency gain, 20% lower site monitoring costs | ⚠ Vendor-cited from one top-3 pharma customer |
| Triomics → PRISM | Systemwide oncology trial matching at Mount Sinai Tisch — first NCI-designated Comprehensive Cancer Center in NYC to deploy it | 8 Jan 2026 | None published | ✅ Institution says outcomes will follow in peer-reviewed publications. Honest by omission |
| Tempus acquires Deep 6 AI | Consolidation of the EHR-matching layer | 11 Mar 2025 | 750+ provider sites, 30M+ patients | Scale of the dataset — not an effect size |
| Inovalon Patient Finder | EHR-integrated AI pre-screening | 25 Feb 2025 | “40–60% faster Phase 2/3 recruitment” | ⚠ Unverified. Appears in no release, product page, or traceable source |
| AiCure H.Code | Smartphone computer vision for dosing confirmation | Sep 2024 | None | Capability only; no accuracy evidence located |
| Microsoft Azure Trial Matcher | Patient↔trial matching, patient-centric and trial-centric | GA | Algorithm is “recall-optimized”; output “should be reviewed by a human decision maker” | ✅ The disclosure pattern to demand |
| CluePoints IMC | Agentic deep learning for MedDRA/WHODrug coding, human-in-the-loop, trained on >100 trials | Sep 2026 | “up to 99% MedDRA, 95% WHODrug”; ~50% less manual effort | ⚠ Vendor — but the caveat is in their own PDF: “depending on input quality, new information, and coding standards” |
| Medidata Plus / Dot | Unified platform (Dot orchestration engine, Mar 2026; Plus launched 23 Jul 2026) | 2026 | “>38,000 trials, 12 million patients” | ⚠ Boilerplate describing Medidata’s cumulative historical dataset, not a platform metric. Prior milestone: 30,000 studies / 9M participants (2023) |
Two things are worth extracting from that table.
Microsoft and CluePoints show what disclosure looks like. One publishes a transparency note stating the algorithm optimises for recall and requires a human decision maker. The other prints its own accuracy caveat in its own product PDF. Neither claims a business outcome it cannot source. That is the standard to hold vendors to, and it is achievable — because two of them already meet it.
Dataset scale is being substituted for effect size. “750+ provider sites spanning more than 30 million patients” is a moat and a data-rights position. It says nothing whatsoever about whether the model using that data improves a trial. The two get merged in the same sentence, and the reader absorbs the credibility of the first as if it belonged to the second.
The percentages nobody measured
This is the part worth keeping. Each claim below is circulating as fact in 2026 vendor and consultancy material. Each has no traceable source.
| Claim | Status |
|---|---|
| Medidata “35% cut in monitoring costs” | No Medidata primary source exists |
| Medidata “50% faster data cleaning” | No source; site-scoped and phrase searches empty |
| Inovalon “40–60% faster Phase 2/3 recruitment” | Zero hits |
| “Cuts AE coding backlogs by up to 70%” | No source. The nearest real figure is CluePoints’ ~50% manual-effort reduction |
| “ML stratification cuts sample size 20–30%” | No source. Real figures: 15%, 50%, 90% |
| “Digital biomarkers cut serious adverse events by 20%” | No source |
| “AI recruitment cuts delays by up to 50%” | Untraceable. Nearest figure is a 50% screening-failure reduction |
| “65% better enrollment,” “85% accuracy forecasting trial outcomes,” “resolves 30–50% of data-entry errors” | Trace only to a single 2026 narrative review that states them in its abstract citing no primary study — then repeated verbatim by vendors |
| “Flags integrity issues in 24–48h vs 4–6 weeks”; monitoring costs “30–40% lower”; biomarkers at “90% AE-detection sensitivity” with “15–20% false positives” | Same narrative review, with citations to blogs and registry trend pages |
| “Up to 80% of EHR content is unstructured” | Review-level; not traceable to a rigorous primary measurement |
| “Automated study builds cut DB construction from 10–12 weeks to days” | Vendor claims — and the vendors’ own stated baseline is 12–16 weeks, not 10–12 |
The concentration is the finding. It is not that these numbers are wrong in eleven independent ways. It is that eleven of them come from the same place — one narrative review that asserts them without primary citations — and the field has been quoting each other’s quotations of it. This is how a number becomes a benchmark without ever becoming a measurement.
Market forecasts that disagree with themselves
Two forecasts are in wide circulation for adjacent parts of this space.
- Fortune Business Insights: AI in clinical trials at $3.8B (2025) → $5.5B (2026) → $77.3B (2034), a 39.14% CAGR.
- Grand View Research: clinical-trials matching software at $210.7M (2025) → $595.3M (2033), a 14.0% CAGR.
Those are different segments, so the absolute figures are not comparable. The growth rates should be, and they differ by twenty-five points. Then there is the detail that settles the question of what these documents are for: Grand View’s own earlier edition, published in April 2025 on a 2025–2030 horizon, gave $187.09M → $396.13M at 13.45%.
A publisher can revise a forecast. What it cannot do is have two internally consistent answers to the same question. The rates are also the kind that only make sense in a market where every participant is assumed to buy: a 39% CAGR sustained for eight years implies the buyer population grows faster than the trials it serves. Use these numbers for slide headings, not for capacity plans.
Three claims that change shape when you check them
The n2c2 result is real, and the dataset year is wrong. A prompt-based LLM doing cohort selection from unstructured notes reports micro/macro F-measures of 0.9061 and 0.8060 — exact, and among the best reported on that benchmark. It is the n2c2 2018 cohort-selection set (311 patients, 13 eligibility criteria). There was no clinical-trial-matching task in n2c2 2022; that year’s tracks were contextualised medication event extraction, social determinants of health, and progress-note assessment/plan. A citation naming the wrong year of the wrong shared task is how a genuine result acquires an unearned halo of recency.
Methodological rigour is being relabelled as machine learning. Hierarchical composite endpoints, generalised pairwise comparisons and the win-ratio framework behind Net Treatment Benefit are real and important — Finkelstein and Schoenfeld in 1999, Buyse in 2010, the standard JACC treatment in 2023. They are also frequentist non-parametric statistics, not ML. Folding them into the AI column makes the AI evidence base look fuller than it is, and it confuses two things with entirely different validation requirements: a pre-specified statistical method and a trained model.
One correction runs the other way, and it is the more common failure. The oft-quoted oncology figure — “more than 10,000 actively recruiting trials, 60% enrolling fewer than five patients per site, more than 20% enrolling none” — turns out to be accurately attributed. The paper it comes from does contain it verbatim. That paper then cites a 2008 editorial in The Oncologist. So the claim is not fabricated; it is eighteen years old and presented as current, in a field where the number of recruiting trials and the operational economics around them have changed materially. Fabrication is loud and rare. Staleness is quiet and everywhere — and it is harder to catch, because the citation chain is intact.
What actually is verified — and it happens to be the business case
The claims that survive citation in this domain are unglamorous, and they are the ones that justify investment.
Protocol amendments. Across 836 protocols, 57% carried at least one substantial amendment, with a mean of 2.2 global amendments in Phase II and 2.3 in Phase III. Median direct cost: $141,000 for a Phase II substantial amendment, $535,000 for Phase III. A single amendment in a typical Phase III adds roughly three months of unplanned time.
Data-collection waste. Across 105 Phase II/III protocols from 14 companies, non-core and non-essential procedures together account for up to 32.5% of the Phase III data collected per patient, representing 25–30% of participant and site burden. Phase III now averages 5.96 million data points per study.
Read those two paragraphs together and the operational case for AI in trials writes itself, without needing a single unverified percentage: more than half of protocols get substantially amended at six-figure cost and multi-month delay; roughly a third of the data collected per patient serves neither an endpoint nor a safety purpose. EmulatRx-style protocol critique and criteria-impact simulation are aimed exactly at the first problem. Automated query generation and data-quality triage are aimed at the second.
That is a defensible investment thesis. It is also a smaller claim than “30–50% faster timelines,” which is why the bigger claim keeps getting made instead.
If you are buying or building
- Label the self-reports in your own business case. Most of the numbers you will inherit are vendor claims. That is not disqualifying — using them internally as unlabelled facts is.
- Separate dataset scale from effect size. “30 million patients” is a moat. It is not an outcome.
- Ask whether the model output is patient-level or trial-level. The 76% prostate result is patient-level responder prediction; it is quoted as trial-success prediction.
- Require the external cohort. 0.768 internal becomes 0.751 external in the one study here that reports both. Ask what the external number is; if there isn’t one, you are reading a fit statistic.
- Check the dataset vintage. n2c2 2018 is cited as 2022. A 2008 oncology estimate is cited as current. A 2012 site-enrolment benchmark is cited as the baseline. Staleness is the default failure mode, not fabrication.
- Demand the disclosure pattern Microsoft and CluePoints already use — stated limitations, recall-optimised framing, a named human decision gate. If a vendor cannot articulate what its model does not do, you are not buying a validated tool.
- Distinguish a causal framework from a matching engine. If the claim is about treatment effects from real-world data, ask for the target trial specification. If there isn’t one, the output supports prioritisation, not evidence.
The bottom line
The applied layer of AI in clinical trials is real, and its centre of gravity has moved. Protocol authoring is becoming a machine-readable object under M11; agentic systems refine protocols against real-world data; target trial emulation gives that work a causal spine; and the matching layer is consolidating into a handful of platforms with genuine data positions.
What has not moved is the verification layer. Every deployment figure in this space is a self-report, the methodology behind the headline numbers is frequently a different methodology, and eleven of the circulating statistics trace to one review that cited nothing. The strongest externally validated result in the place you would least expect it — ICU stratification — reports 0.768 internal and 0.751 external, and nobody is quoting it.
That is the same pattern as the rest of this series, one level down. In molecular design, virtual cells, biomarker discovery and now trial operations, the capability is genuine and the magnitude is marketing — and the work that closes the gap is not more modelling. It is an external cohort, a documented target trial specification, a stated context of use, and a named human accountable for the decision the model informed.
Five agents can now write your protocol. None of them can validate it for you.
Related: Four Points of Accuracy, Fifty Points of Marketing: The Evidence Base for AI in Clinical Trials · 117 AI-Built Drug Programs, Zero Approvals · Two of Eighty-Six: The Validation Deficit in AI Biomarker Discovery · Representation Is Solved, Simulation Isn’t
Sources: EmulatRx, Nat Commun 17:5501 (2026), DOI 10.1038/s41467-026-74501-2 · Hernán & Robins, Am J Epidemiol 2016;183(8):758–764 · Hernán, Wang & Leaf, JAMA 2022;328(24):2446–2447 · Zang et al., Nat Commun 14:8180 (2023) · Rajendran et al., Nat Commun 16:4450 (2025) · Schwager et al., npj Digit Med 4:133 (2021) · Beacher et al., Algorithms 14(5):147 (2021) · Tayebi Arasteh et al., Nat Commun 15:1603 (2024) · Rahmanian et al., arXiv:2404.16198 / Healthc Inform Res 2025;31(4):367–377 · Clin Transl Sci, DOI 10.1111/cts.70183 · Finkelstein & Schoenfeld, Stat Med 1999;18(11):1341–1354 · Buyse, Stat Med 2010;29(30):3245–3257 · JACC 2023, DOI 10.1016/j.jacc.2023.06.047 · Asher et al., BMC Med Res Methodol 2022;22:5 · Curt & Chabner, The Oncologist 2008;13(9):923–4 · Bennette et al., JNCI 2015;108(2):djv324 · Getz et al., Ther Innov Regul Sci 2016;50(4):436–441 · TransCelerate/Tufts CSDD, 105 protocols across 14 companies · Olawade et al., Int J Med Inform 2026;206:106141 · AMARANTH, Nat Commun 16:6244 · PROCOVA EMA CHMP qualification opinion, 20 Sep 2022 · ICON press release, 28 Jul 2026 · TCS newsroom, 24 Nov 2025 · Mount Sinai newsroom, 8 Jan 2026 · Tempus investor release, 11 Mar 2025 · Inovalon launch release, 25 Feb 2025 · Business Wire (AiCure H.Code), 10 Sep 2024 · Microsoft Learn, Azure AI Health Insights Trial Matcher transparency note · CluePoints Intelligent Medical Coding product documentation, Sep 2026 · Medidata/3DS releases, 11 Feb 2026 and 23 Jul 2026 · Fortune Business Insights, AI in Clinical Trials Market (SKU FOB21037582, 1 Mar 2026) · Grand View Research, clinical-trials matching software report · ICH M11 (adopted 19 Nov 2025; 91 FR 30310) · ICH E6(R3) (6 Jan 2025; Annex 2, 3 Jun 2026) · Veeva Study Builder Agent release, 24 Sep 2026.
Research notes: [[ML-Clinical-Trials-Applied-Layer-2026]]
Saram Consulting