Here is the number that should govern how you read everything else in AI precision medicine.
A 2025 bibliometric and scoping review in npj Digital Medicine surveyed digital biomarkers for Alzheimer’s disease: 431 studies across five databases, 86 AI models scoped in detail. Among the models that reported performance, 21 Alzheimer’s-focused models averaged AUC 0.887, and 45 models for mild cognitive impairment averaged AUC 0.821. Respectable discrimination.
Now the validation columns. Only 2 of the 86 studies incorporated external validation. Only 3 assessed model calibration — and the paper’s own Results section says 4, so call it three-to-four. Only 1 of 86 acknowledged TRIPOD reporting guidelines. Only 2 of 86 shared code.
This is stated three separate times in the paper, in the abstract, the Results and the Discussion, which is unusual care for a negative finding. It is not a soft result.
Put plainly: an AUC of 0.887 computed on a model whose transportability to a new site, a new device, or a new population has never been tested is not evidence of clinical utility. It is evidence of internal fit.
That is the honest centre of this field, and it connects directly to what we found in the molecular-design and virtual-cell pieces: candidate generation has outrun verification.
The framework, precisely stated
Before the critique, the definitions — because several of them are routinely garbled.
| Element | Status | Detail |
|---|---|---|
| BEST definition | Glossary — non-binding | “a defined characteristic that is measured as an indicator of normal biological processes, pathogenic processes, or biological responses to an exposure or intervention, including therapeutic interventions” — the qualifiers are usually dropped |
| Seven categories | Glossary — non-binding | susceptibility/risk, diagnostic, monitoring, prognostic, predictive, pharmacodynamic/response, safety |
| “Surrogate endpoint” | Not an eighth category | BEST defines it separately as an endpoint type — candidate / reasonably likely / validated |
| Biomarker Qualification Program stages | Statute | Letter of Intent → Qualification Plan → Full Qualification Package, established by FD&C Act §507 (21st Century Cures Act, 2016) — not merely program process |
| Qualification ≠ assay qualification | Program FAQ | “Biomarkers … are independent of the specific test.” Assay performance is considered during LOI/QP; any analytically validated assay may be used afterwards |
Two distinctions in that table carry most of the practical weight.
A biomarker is not a test. A qualified biomarker is a context-of-use determination — the biomarker may reliably support a specified interpretation for a specified purpose. Your assay still needs its own analytical validation. Any AI-biomarker commercial plan that assumes qualification blesses the measurement method is misreading the instrument, and this is the single most consequential regulatory misunderstanding in the space.
A prognostic biomarker is not a predictive one. A model that predicts survival has not shown that a treatment works better in the patients it selects. Predictive claims require evidence of treatment-effect heterogeneity, which is a strictly harder evidentiary problem — and the reason the field is moving toward causal methods.
The validation deficit
Two reviews, taken together, map the gap.
The 2025 Alzheimer’s digital-biomarker review: 2 of 86 externally validated; 3–4 calibrated; 1 of 86 acknowledging TRIPOD; 2 of 86 sharing code. Two figures from that paper’s abstract should be avoided, incidentally, because they contradict its own Results section: “1,403 institutions” (the Results say 912, and the percentages confirm 912) and “224 grants” (the Results report 1,345 funding instances across 539 sources; 224 counts government projects only).
The 2021 scoping review of omics-based stratification — Glaab et al., BMJ Open 11:e053674 — included 352 articles. From that literature, 13 clinically approved or cleared tests emerged: 9 cancer and 4 non-cancer, comprising MammaPrint, Prosigna/PAM50, Oncotype DX, Decipher, Afirma, FoundationOne Heme, PGDx Elio, AlloMap, Corus CAD and Vectra DA among others. Four were FDA-cleared, eight laboratory-developed tests, one a device. And the validation discipline: 78% validated internally only, only 12% used internal cross-validation plus an external cohort, roughly 10% external-only.
WHERE THE EVIDENCE ACTUALLY STOPS
Candidate signatures published ████████████ every study
Internal validation only █████████░░░ ~78% (2021 review)
Any external validation ██░░░░░░░░░░ ~22% (2021 review)
Externally validated + calibrated ░░░░░░░░░░░░ 2 of 86 (AD digital)
Code shared for reproduction ░░░░░░░░░░░░ 2 of 86
TRIPOD reporting acknowledged ░░░░░░░░░░░░ 1 of 86
What the credible studies have in common
They all have a held-out cohort. Without exception.
| Study | Verification discipline |
|---|---|
| Necrotizing fasciitis vs osteomyelitis (npj Digital Medicine 2026) | 3,415 patients, 10 routine blood biomarkers, LightGBM — AUC 0.926 on an independent external cohort from a second centre, deployed as a public web tool |
| EGFR mutation from H&E in lung adenocarcinoma (Nature Medicine 2025) | 8,461 international cases; internal AUC 0.847, external 0.870, prospective 0.890 |
| Kawasaki disease single-cell panel (Front. Immunol. 2025) | Training AUCs 0.914/0.958/0.985; genuine external validation cohort GSE73461 at 0.872/0.861/0.893 |
| Radiologic foundation model (Nat. Mach. Intell. 2024) | 11,467 lesions; immune-pathway enrichment independently recovered interferon, MHC class II and PD-1 signalling |
| Longitudinal EHR stratification (ML4H 2021) | 29,229 diabetes patients; clusters differing in trajectory and outcome |
| PBMF (Cancer Cell 2025) | Retrospective application to immuno-oncology trials: 15% survival-risk improvement in one Phase 3, ≥10% in two others with synthetic control arms |
| Autism severity from motor features (Front. Psychiatry 2026) | 84.6±10.9% cross-validated, 86.4% held-out — and the paper itself notes its 100% severe-class accuracy came from synthetic data, calling it a proof-of-concept rather than clinical validation |
That last row is worth pausing on. A study that volunteers the limits of its own strongest-looking number is doing something the field needs more of. Contrast it with an “AUC 0.98” headline that turns out to describe a different comparison than the one you assumed.
Which brings us to a correction worth stating separately: plasma p-tau217 is not an AUC 0.96–0.98 marker for discriminating Alzheimer’s from other neurodegenerative diseases. The defensible figure is roughly 0.89–0.96, with a pooled estimate of 0.95 across a 2026 meta-analysis of 33 studies and 6,138 patients. The 0.98 figure describes Alzheimer’s versus cognitively unimpaired or MCI — an easier comparison. And specificity against dementia with Lewy bodies is only about 0.69. The number is still excellent. It is not the number being quoted.
Where AI genuinely expands what can be a biomarker
This is the upside, and it is real. The most commercially advanced area is computational pathology, where a morphological biomarker can substitute for or triage an expensive molecular assay.
| Model | Verified scale and performance |
|---|---|
| CHIEF (Nature 2024) | 60,530 whole-slide images across 19 sites (44 TB); validated on 19,491 WSIs from 32 slide sets at 24 hospitals; IDH in glioma, MSI in colorectal; survival validated in 15 cohorts |
| Virchow (Nature Medicine 2024) | 0.95 specimen-level AUC across nine common and seven rare cancers; 1.5M WSIs, 632M parameters. The “largest to date” superlative was accurate in July 2024 and has since been superseded |
| CONCH (Nature Medicine 2024) | Vision-language model over 1.17M image-caption pairs; highest overall performance in a 19-model benchmark, marginally ahead of Virchow2 |
| RIDGE (BMC Medical Imaging, Sept 2026) | 4,983 WSIs across 12 solid tumour types, overall AUC 0.763 — a useful reminder that breadth across tumour types costs per-task accuracy |
| Stratipath Breast | CE-IVD marked prognostic stratification from H&E, validated specifically in ER-positive/HER2-negative early breast cancer classified intermediate risk |
Two corrections for anyone citing the radiomics literature: PyRadiomics’ default output is 107 features, not 105 (verified by running v3.0.1), and the widely repeated “six clusters” grouping — global variance, local dimness, global intensity, morphologic structure, bright region, local uniformity — has no published source and is not part of the PyRadiomics taxonomy or the IBSI standard.
The regulatory path is narrower than the summaries suggest
This is where circulating commentary diverges most from the instruments, and where the difference between a strategy and a hope lives.
The kidney-injury panel. Six urine biomarkers — clusterin, cystatin C, KIM-1, N-acetyl-β-D-glucosaminidase, NGAL and osteopontin, each normalized to urine creatinine — were qualified in 2018. But the qualification is narrower than the retelling: they are interpreted as a composite measure (geometric mean of fold-changes), qualified for adjunctive safety monitoring in Phase 1 healthy-volunteer trials. The qualification explicitly does not cover replacing standard measures and is not qualified for individual-patient safety monitoring. “Prespecified safety endpoint” overstates it.
The ADPKD qualification. Baseline total kidney volume as a prognostic enrichment biomarker, to select patients at high risk of a confirmed 30% eGFR decline, usable with age and baseline eGFR in patients ≥12 years. Published 16 September 2016 (Docket FDA-2015-D-2843).
Digital biomarkers. There is no dedicated digital-biomarker qualification track, and the claim that verification evidence is simply “folded into LOI/QP/FQP” is not supported by any FDA source. What exists instead is a separate apparatus: the DHT Framework (March 2023), a DHT Steering Committee, DHT guidance, and ordinary CDRH device regulation. A digital biomarker reaches the clinic through device clearance or through device/DHT qualification for a specific use — not by riding along inside a biomarker submission.
Regulatory research signals. FDA’s CPIM topics have included “Use of Artificial Intelligence (AI) to Identify Biomarkers of Toxic Liver Disease” (March 2026) — but CPIM is a non-regulatory meeting and its advice is non-binding. The agency’s CERSI project on AI for adaptive enrichment in clinical trials dates to October 2021. Both are signals of interest, not pathways.
Companion diagnostics remain the regulated end-state: identifying patients most likely to benefit, at increased risk of serious side effects, or for monitoring response to adjust treatment.
A provenance audit
Here is the part that should change how you read the surrounding narrative. Every claim below was checked against primary sources.
| Claim | Actual status |
|---|---|
| COMPASS — predicts checkpoint-inhibitor response | Peer-reviewed and strong. Nature Medicine 32:3010–3022, 3 July 2026. 10,184 tumours / 33 cancer types pretraining; 16 cohorts, 1,133 ICI-treated patients; +8.5% accuracy, +15.7% AUPRC; held-out Phase 2 urothelial responders with HR 4.7, P = 1.7×10⁻⁷, beating TMB and PD-L1. The concept layer is 44 tumour-immune-microenvironment concepts, not the 43 usually quoted |
| Tempus Immune Profile Score | Peer-reviewed (peer-reviewed J. Immunother. Cancer 2025) |
| Tempus HR 4.536 on osimertinib; xM MRD survival claim | Press-release only — Tempus investor release, 29 May 2026 |
| MutationProjector — resistance prediction, KMT2D / SMARCA4+STK11 | Peer-reviewed — Cancer Discovery, 27 May 2026 |
| Tufts CSDD / DIA 2025 — 36 activities, 302 respondents; 68% time reduction on targeted patient identification; 18% mean cycle-time reduction | Peer-reviewed — Ther. Innov. Regul. Sci. 59:1074–1086 |
| McKinsey: 340 AI deployments, median 210% ROI, 16-month payback | FALSE — misattributed. The cited McKinsey article contains no such figures. It reports 6-month timeline compression, 10–20% enrolment gains and up to 50% cost reduction. The 340/210%/16-month numbers exist only on AI-marketing blogs |
| Immunai AMICA: 25% shorter AstraZeneca trial | Unsourced. No corroboration; the sole origin is a LinkedIn article. The real Immunai–AstraZeneca relationships are an $85M IBD collaboration and an up-to-$37.5M oncology expansion |
| Biomarker market sizes; “AI in clinical trials $3.8B → $8.5B → $14B” | Press-release only, and internally inconsistent. One firm’s own North America page says $45.25B where its release says $42.25B. The AI-in-trials chain stitches three different forecasts from three different firms with different base years into one continuous-looking line |
| “142 AI companies across 12 categories and 9 therapeutic areas” | Do not cite. A LinkedIn Pulse post, self-labelled “AI-assisted by Claude,” with no methodology or dataset. It is internally contradictory on market size and is the origin of both the misattributed McKinsey statistic and the unsourced Immunai figure |
Read the bottom three rows together. A fabricated McKinsey statistic and an unsourced vendor performance claim, both originating in a single low-reliability analysis, are now circulating through AI-pharma content as established facts. That is not a problem with the models. It is a problem with the evidence supply chain — and it is the same failure mode the 2-of-86 finding measures in the academic literature.
Vendor device claims need the same scrutiny. “NeuroRPM received FDA Breakthrough Device status in 2024” is wrong on both counts: the actual record is a 510(k) clearance (K221772, 17 March 2023), and no Breakthrough Device designation exists. “physIQ has six FDA clearances” is unsupported — the openFDA record shows four. “Quibim QP-Brain cleared for early neurodegeneration detection” is a cleared brain-MR image processing and quantification product; the disease-detection framing is marketing.
If you are building or buying here
- Ask for the external cohort before the AUC. If a number comes from a random split on data from the same institution, it is a fit statistic, not a performance estimate.
- Ask about calibration separately. Discrimination and calibration fail independently, and almost nobody reports the second.
- Separate prognostic from predictive claims. Treatment-effect heterogeneity requires a different study design and a different evidentiary standard.
- Read qualification as context of use. The biomarker is qualified for a purpose. Your assay, your cut-off and your platform still need analytical validation — and the qualification may be narrower than the headline (adjunctive, composite, Phase 1 only).
- Check the device record, not the press release. openFDA will tell you whether something is a 510(k), a De Novo, an approval, or a Breakthrough designation — and those are four different things.
- Trace every market and ROI figure to its origin. If it traces to a press release, label it as vendor-sourced. If it traces to a blog post, drop it.
The bottom line
AI has genuinely widened what can serve as a biomarker: multimodal signatures, learned morphology from routine H&E, continuous sensor streams, and protein panels interpreted per-sample rather than as fixed lists. The strongest result in the field — COMPASS at Nature Medicine, with a held-out Phase 2 survival signal beating both TMB and PD-L1 — is exactly the kind of evidence that deserves attention.
But the field’s binding constraint is not discovery. Two of 86 studies externally validated. Three or four checked calibration. Seventy-eight per cent of the prior generation validated internally only. And the regulatory path is a context-of-use determination that leaves the assay, the cut-off, and the platform still to be validated.
The discipline that separates a biomarker from a biomarker-shaped number is unglamorous: a held-out cohort, a calibration curve, a shared model card, and a provenance trail that survives being checked. Everything else is a fit statistic with a clinical-sounding name.
Related: 117 AI-Built Drug Programs, Zero Approvals · The Generalization Wall · Representation Is Solved, Simulation Isn’t
Sources: FDA BEST resource (glossary NBK338448) · FD&C Act §507 (21st Century Cures Act 2016) · FDA Biomarker Qualification Program pages · PSTC drug-induced kidney injury decision letters · ADPKD total kidney volume qualification, FR 81 FR 63764 (Docket FDA-2015-D-2843) · FDA companion diagnostics page · FDA CPIM topics (March 2026) · FDA CERSI AI adaptive-enrichment project (Oct 2021) · FDA DHT Framework, FR 2023-06066 · Qi et al., npj Digital Medicine 8:366 (2025) · Glaab et al., BMJ Open 11:e053674 · Yasin et al., npj Digital Medicine 9:507 (2026) · Feng et al., Front. Immunol. 16:1541939 · Fırat, Front. Psychiatry 17:1751654 · Carr et al., ML4H 2021, PMLR 158:220-238 · Arango-Argoty et al., Cancer Cell 2025 · CHIEF, Nature 2024 · Virchow, Nature Medicine 2024 · CONCH, Nature Medicine 2024 · EGFR-from-H&E, Nature Medicine 2025 · MGB radiologic foundation model, Nature Machine Intelligence 2024 · RIDGE, BMC Medical Imaging 2026 · Stratipath Breast (CE-IVD) · PyRadiomics v3.0.1 documentation · Palmqvist et al., JAMA 2020 · dementia plasma proteomics, Nature Medicine 2026 · adaptive proteomics framework, Nature Communications 2026 · openFDA 510(k) records K221772, K232231 · COMPASS, Nature Medicine 32:3010–3022 · Tempus Immune Profile Score, J. Immunother. Cancer 2025;13(5):e011363 · MutationProjector, Cancer Discovery 2026 · AMARANTH, Nature Communications 16:6244 · Tufts CSDD/DIA, Ther. Innov. Regul. Sci. 59:1074–1086.
Research notes: [[AI-Biomarkers-Stratification-Precision-Medicine-2026]]
Saram Consulting