Two numbers define the state of AI in pharma in 2026, and they point in opposite directions.
117. That is the number of AI-enabled therapeutic assets from 63 companies that have entered interventional clinical trials — 60 of them finishing Phase I, 8 finishing Phase II. AI-discovered molecules clear Phase I at 80–90%, against a historical baseline of 40–65%. The chemistry works. The prediction works. The deal flow is real: Lilly committed up to $1B over five years to an NVIDIA co-innovation lab, BMS pushed agentic AI to more than 30,000 employees with Anthropic, and Roche committed to a 2,176-GPU Blackwell AI factory on the way to what it calls autonomous labs.
Zero. That is the number of AI-originated drugs with FDA approval. Phase II success sits at about 40% — statistically indistinguishable from the historical average. No organization has demonstrated that AI compressed total development timelines, and drug discovery is only 20–30% of a 10–15 year path to market.
Both numbers are true simultaneously. That tension is the whole story, and it is not a story about model capability.
The most advanced candidate is Schrödinger and Nimbus’s zasocitinib (TAK-279), which met all endpoints in Phase 3 LATITUDE-PsO and had its NDA accepted under Priority Review on 14 September 2026 with a PDUFA date in Q1 2027. If it approves, it will likely be the first drug anyone can credibly call AI-designed — and the credibility of that label depends entirely on how much human medicinal chemistry you subtract before it becomes dishonest. FDA has no category for “AI-discovered.” Every such claim is an attribution argument, not a fact.
Here is where the technology genuinely stands, where it stops, and why the bottleneck moved somewhere most AI programmes are not looking.
Where AI is actually deployed
Deployment maturity tracks exactly one variable: how clean, structured, and verifiable the underlying data is. Everything else — enthusiasm, budget, executive mandate — is noise.
DEPLOYMENT MATURITY BY VALUE-CHAIN SEGMENT
Target ID ████████████ production
Molecule design ████████████ production
Protein structure ████████████ production
ADMET prediction ██████████░░ production, benchmark-bound
Preclinical / phenomics ████████░░░░ piloting → production
Regulatory writing ████████░░░░ fastest ROI, human-signed
Pharmacovigilance ████████░░░░ mature but conservative
Clinical development ██████░░░░░░ growing, uneven
Quality / deviations █████░░░░░░░ assist-only, never approves
CMC / manufacturing ████░░░░░░░░ piloting, rarely submitted
A few of those bars deserve specifics.
Structure prediction became infrastructure. AlphaFold DB carries more than 200 million predicted structures, and targets that never had an experimental structure now at least have a druggability hypothesis. RFdiffusion, Chai-1, ESMFold, and ProteinMPNN have moved de novo binder and enzyme design from demo to routine in biologics groups.
ADMET prediction works — on the right test. Commonly reported performance sits around 0.85–0.92 AUC for CYP inhibition and 0.89–0.96 for hERG. Those numbers are almost always measured on random splits. On scaffold splits — that is, on genuinely novel chemistry, which is the only chemistry you care about — degradation is common, though not universal across models. Remember that whenever someone quotes a benchmark at you.
Clinical development is the uneven middle. Around 26% of organizations use AI for trial automation, adverse event detection, or treatment response prediction; roughly 42% are piloting site selection. The reported wins are enrolment compression and protocol drafting. They are real but incremental, and they do not touch the cost of proving a mechanism works in humans.
CMC and manufacturing is further behind than the vendor decks suggest. An industry survey of digital and AI-enabled CMC tools found widespread experience with them — but fewer than 15% had been included in regulatory submissions. The barrier is not capability. It is culture change, skills, and uncertainty about what evidence regulators will accept.
Quality and pharmacovigilance show the pattern that matters most. In PV, machine learning for signal detection has been studied extensively and frequently outperforms traditional disproportionality methods — yet classical statistical methods still dominate the installed base, and only 12.3% of reviewed PV studies used ML or deep learning. The advanced technology exists. The operational system is older. That gap between what is available and what is deployed recurs everywhere in this industry, and it is the thing to budget against.
What it can and cannot do
| Stage | Can do today | Cannot do today |
|---|---|---|
| Biology | Cluster high-dimensional omics to prioritise targets; predict rigid-to-moderate protein conformations | Reconstruct emergent systemic human biology; model complex allostery, multi-protein complexes in cellular flux, epigenetics |
| Chemistry | Generate lead-like hits against a static structure; propose retrosynthetic routes for standard chemotypes | Guarantee synthetic accessibility for non-standard scaffolds; predict in vivo efficacy from in vitro affinity |
| Safety / ADMET | Flag known cardiotoxicity and mutagenicity motifs; kill molecules with extreme PK liabilities | Predict idiosyncratic DILI, immune-mediated adverse events, secondary organ damage |
| Clinical | Accelerate patient matching, protocol drafting, real-world evidence extraction | Compensate for an ineffective target; predict systemic immunological response |
| Metric | AI-originated | Traditional baseline |
|---|---|---|
| Phase I success | 80–90% | 40–65% |
| Phase II success | ~40% (limited sample) | ~40% |
| Total development timelines | Not shown to improve | — |
| FDA approvals | 0 | — |
The Phase I advantage is not mysterious. Phase I mostly tests safety and pharmacokinetics, which is precisely what computational design is good at filtering. Phase II tests whether the target was right in human physiology, which is precisely what no current model can answer. AI accelerates making molecules. The bottleneck was never molecule-making.
The 2026 failures make the point better than the wins. DSP-1181 — the first AI-designed drug dosed in humans, from Exscientia and Sumitomo — was discontinued after Phase I. Recursion’s REC-994 met its primary endpoint, because its primary endpoint was safety, and then failed on efficacy before being discontinued in May 2025. REC-2282, same day. Schrödinger’s SGR-2921, discontinued in August 2025. Isomorphic Labs, holding roughly $3B in partnership value, still had no clinical candidate as of mid-2026.
Four claims to stop repeating
“AI models understand binding.” A 2025 Nature Communications study tested four co-folding models — AlphaFold 3, RoseTTAFold All-Atom, Chai-1, Boltz-1 — against deliberately disruptive changes at the binding site, and found they produced largely unchanged predictions. The important detail is that the disruptions were in silico point mutations, not experimentally destroyed binding sites; protein fold was retained, AlphaFold 3 actually performed best, and several challenges produced only partial ligand displacement. The legible conclusion is not “the models are broken.” It is that these systems recognise patterns they have seen before rather than modelling the physics of binding — which tells you exactly when to trust them (interpolation within known chemistry) and when not to (novel chemotypes, induced fit, cryptic pockets).
“AI is improving drug development outcomes.” Nature Reviews Drug Discovery published an assessment in August 2026 arguing that evidence for clinically meaningful improvement remains limited. Its stated causes are worth reading verbatim because they are not algorithmic: conditional life science data, insufficient problem definitions, integration barriers, siloed data, and the difficulty of translating computational predictions into real decisions. “Conditional” is the load-bearing word. A cell assay can shift a tenfold IC₅₀ on passage number, serum batch, or temperature drift — and most training corpora silently flatten that away.
“Hallucination will be patched.” A 2025 formal treatment argues that hallucination is undecidable for current and foreseeable architectures. That argument is contested — there is a published rebuttal — so treat it as a serious unresolved position rather than settled mathematics. The operational consequence is not in dispute: on regulated output, human accountability must be 100%, because you cannot engineer the failure class away.
“The problem is not enough data.” The problem is the shape of the data. Deep learning assumes abundance; drug discovery produces small, expensive, confidential, highly conditional datasets. High-quality experimental data scarcity is not an anomaly to be fixed with scraping. It is a structural property of the domain.
The disconnect
The AI sits in one place and the work sits in another. Four reference implementations of a molecule are easier to obtain than one comparable assay result. That is the terrain.
SPECIALIST (chemist · clinician · QA · regulatory)
│
▼
┌─ AI interface ──────────────────────┐ ← where vendors compete
└──────────────────┬──────────────────┘
│
┌──────────┴──────────┐
▼ ▼
reasoning agents
└──────────┬──────────┘
▼
┌─ Knowledge layer ───────────────────┐ ← where value compounds
│ ontology · provenance · │
│ entity registry · claims │
└──────────────────┬──────────────────┘
│
┌──────────────┼──────────────┐
▼ ▼ ▼
ELN LIMS QMS ← where the data actually is
Five layers of friction sit in that seam, and only one of them is about technology.
1. The qualification gap
Pharma validates equipment, assays, and manufacturing processes. Nothing goes live without a validation protocol; nobody operates it without documented training; there is a paper trail proving the operator understands failure modes and limitations.
AI implementations routinely skip all of it. They arrive as software rather than as a regulated system requiring operator qualification, challenge conditions, and documented human oversight.
The enforcement proof is not theoretical. In April 2026, FDA issued Warning Letter 320-26-58 to Purolea Cosmetics Lab — a CDER drug-cGMP action despite the company name — with a dedicated section titled “Inappropriate Use of Artificial Intelligence in Pharmaceutical Manufacturing.” The finding was not that AI was used. It was that AI was relied on to generate specifications, procedures, and master production and control records without authorised human Quality Unit review.
Read that as the general case: regulators are not waiting for finalised guidance. They are enforcing existing CGMP authority against AI-mediated processes now.
2. The verification tax
A report that once took thirty minutes to write now takes three minutes to generate and fifteen minutes to check. That is a real twelve-minute gain, and it is also a completely different kind of work.
The metric that survives contact with a quality system is not generation speed. It is time-to-trusted-output. This is the single most important reframing for anyone buying AI in a regulated environment: a less capable model whose output you can audit beats a more capable model whose output you must re-derive from scratch.
3. Metric mismatch
Computational teams optimise RMSD, docking score, cross-validation R², and benchmark ROC-AUC on public datasets. Bench scientists care whether the compound dissolves, crashes out in buffer, kills mouse hepatocytes, or yields a clean enantiomer on HPLC.
The ADMET numbers are the cleanest illustration. A 0.9 AUC on a random split tells you the model interpolates well inside chemistry it has seen. It tells you nothing about a novel scaffold, and novel scaffolds are the entire point.
4. Workflow fragmentation
The AI lives in Jupyter, Databricks, a vendor portal, or a personal ChatGPT account. The work lives in ELN, LIMS, QMS, ERP, MES, and DMS. The scientist exports, cleans, uploads, prompts, interprets, copies, and re-enters.
Friction of that size kills adoption regardless of model quality — and it does something worse. It pushes capability into ungoverned individual accounts instead of auditable organizational systems, which is exactly the Shadow AI problem every quality function is now trying to quantify.
5. Pilot to production
This is where the numbers get blunt.
| Finding | Figure | Source |
|---|---|---|
| Gen AI pilots delivering no measurable financial return | ~95% | MIT Project NANDA, 2025 (preliminary — 52 interviews plus a conference survey) |
| Executives seeing deployment advance | 71% | Deloitte, Apr 2026, 150 life sciences execs |
| Reporting any measurable improvement | 45% | same |
| Reporting improvement at scale | 13% | same |
| Companies reporting significant gen-AI financial value | ~5% | McKinsey, 100+ pharma/medtech leaders |
| Employees with approved gen-AI access | ~1 in 3 | ZS, 2025, n=127 tech executives |
| Of those, using it weekly | ~10% | same |
| Employees who believe AI helps their team | 77% | Cognizant |
| Employees who understand how to use it in their role | 37% | same |
| ML understanding: executives vs employees | 89% vs 56% | same |
| Clinical-development activities fully implementing AI | 10.7% | Tufts CSDD, Ther Innov Regul Sci 2025 — peer-reviewed, 302 orgs |
| Citing lack of qualified people as an adoption barrier | 34%, up from 23% | Pistoia Alliance, Sep 2025 |
Two of those rows deserve to be read twice.
Access is not adoption. One in three employees have approved tools; roughly one in ten of them use the tools weekly. Every “AI seats licensed” metric in circulation is measuring nothing.
The strategy owners may understand the older AI better than the workforce understands the new AI. Executives rate their own ML understanding at 89%; employees rate their ML understanding at 56% — while employees actually rate themselves higher than executives on generative and agentic AI. The people signing the AI strategy and the people operating it have non-overlapping mental models of what the technology is.
And note the direction of the skills gap: the share citing lack of qualified people as a barrier went up, from 23% to 34%, in a single year. The talent problem is widening, not closing, and 67% of organizations’ AI talent is coming from internal upskilling rather than hiring. The scarce species is not the ML researcher. It is the person who can hold a scientific problem and a model’s failure modes in their head at the same time.
The regulatory position, accurately stated
This is the part of the landscape most often reported wrong, including by people who should know better.
| Instrument | Status as of October 2026 | Substance |
|---|---|---|
| FDA Considerations for the Use of AI to Support Regulatory Decision-Making (Jan 2025) | Still draft. Stamped “Not for Implementation” | A 7-step risk-based credibility assessment tied to context of use. The question is not “is this AI validated” but “is this AI credible enough for this intended use” |
| FDA + EMA Guiding Principles of Good AI Practice (Jan 2026) | Non-binding | Ten principles: human-centric design, risk-based approach, adherence to GxP standards, clear context of use, multidisciplinary expertise, data governance, model design, risk-based performance assessment, lifecycle management, clear information |
| FDA qualification of AIM-NASH (Dec 2025) | Qualified | First AI drug development tool qualified, for liver histology scoring in MASH trials — with pathologists retaining final interpretation |
| FDA Warning Letter 320-26-58 (Apr 2026) | Enforcement | AI-generated specifications and procedures relied on without authorised QU review |
| EU GMP Annex 22 | Draft, not finalised. Consultation closed Oct 2025; final text expected Q4 2026 | Draft language states dynamic, probabilistic-output, and generative AI “should not be used in critical GMP applications” — provisional wording, subject to change |
| EU AI Act high-risk obligations | Delayed — not in force | Regulation (EU) 2026/1744 postponed the high-risk obligations to 2 December 2027 and 2 August 2028. Only general transparency provisions applied on 2 August 2026 |
That last row is a live correction to a great deal of published commentary. If you have read that EU AI Act high-risk rules for AI systems took effect on 2 August 2026, you have read something false. The Digital Omnibus on AI pushed them out.
Two things this table says together. First, the regulatory frame has shifted from “is this AI validated” to “is this AI credible for this context of use” — which is a much better question and one that most validation SOPs are not written to answer. Second, the absence of final guidance is not a safe harbour. Purolea is the proof. Enforcement is running ahead of guidance, and it is being run through existing CGMP authority rather than new AI-specific rules.
Worth noting for anyone building validation strategy: AIM-NASH is the model to copy. AI produces a scoring; the pathologist remains accountable for the interpretation; the audit trail carries both. AI → expert review → accountable decision rather than AI → decision.
What actually closes the gap
-
Qualify the operator, not just the software. Extend the equipment qualification instinct to AI: documented training, evidence the operator understands the model’s failure modes and applicability domain, and a named accountable human for every AI-influenced output.
-
Validate to context of use, not to the model. The FDA credibility framework is the right shape — proportionate evidence for a specific decision, rather than a binary verdict on an algorithm. A model that triages documents and a model that informs a batch release decision do not deserve the same evidence burden, and pretending otherwise is why validation projects stall for a year.
-
Embed translators inside scientific teams. Pair one computational scientist into a disease-biology or medicinal-chemistry group and have them co-own the scientific milestone, not the algorithm milestone. The centre-of-excellence service desk is the failure mode; the embedded pod is what works.
-
Close the dry–wet loop and log negative data. Failed reactions, flatline potency, and insoluble compounds are systematically under-recorded, which starves structure-activity models of their most informative signal. Automated active learning with standardised negative capture is the highest-leverage infrastructure investment in discovery.
-
Ship calibrated uncertainty instead of point predictions. Conformal prediction intervals and out-of-distribution alerts tell a chemist when the molecule has left the model’s applicability domain. A number without an interval invites exactly the over-trust that produces bench disappointment.
-
Treat data readiness as the prerequisite, not the follow-up. The organizations reaching enterprise scale at 9–10× the rate of their peers are the ones that had governance and data foundations before the AI programme, not the ones that promised to fix data afterwards.
-
Own the verifiable middle layer. Models are swappable in an afternoon. What a competitor cannot reproduce is the ontology, the canonical entity registry, the provenance chain, the validation rules, and the audit trail that let you answer “why is this here?” with a document, a version, and a verbatim evidence span. That is the layer the AI interface sits on, and it is where value compounds.
The blind spots
The bottleneck moved from generating hypotheses to verifying them. AI now produces more testable hypotheses than any lab — robotic or human — can assay. The binding constraint is experimental verification capacity, not compute. Nobody’s AI roadmap has a line item for that.
“AI adoption” is being measured with the wrong instrument. Seats licensed and prompt volume measure nothing that survives contact with a QMS. The only metric that does is time-to-trusted-output.
Attribution is unresolved. No regulator classifies a drug as AI-discovered. Every claim rests on how much human chemistry you subtract before the label becomes marketing.
The moat is data and provenance, not model novelty. Foundational models are commoditising on a quarterly cadence. Federated fine-tuning on proprietary data, curated biological datasets, and assembled real-world evidence estates are not.
Nobody owns the seam. Quality and IT are asked to validate probabilistic systems with frameworks built for deterministic software. AI teams rarely have the regulatory context to build for GxP from day one. That unowned seam — model → data → workflow → evidence → human decision → institutional memory — is where the value leaks out of otherwise excellent programmes.
The bottom line
The technology is ahead of the organization, and the organization is ahead of the workforce. AI has demonstrably accelerated the front half of the pipeline and a great deal of knowledge work, marginally optimised clinical operations, and remains an unproven but now genuinely testable bet on clinical efficacy.
The organizations that extract disproportionate value from the same models everyone else can rent will not be the ones with the best algorithm. They will be the ones that treated data readiness, operator qualification, context-of-use validation, and workflow integration as first-class engineering disciplines — and that accepted, early, that the bottleneck has moved from the model to the system around it.
Related: The Decision Layer: Type-Safe Models Outside Your Validated Write Path · Knowledge Graphs Without Hallucinated Edges · How Auditors Are Tackling AI
Sources: Jayatunga et al., Drug Discovery Today 2024;29(6):104009 (doi:10.1016/j.drudis.2024.104009) · ASCO/JCO 2026 abstract doi:10.1200/JCO.2026.44.16_suppl.11072 (117 AI-enabled assets across 63 companies) · Bender et al., Nature Reviews Drug Discovery (doi:10.1038/s41573-026-01496-2) · Masters, Mahmoud & Lill, Nature Communications 16, 8854 (2025) · Banerjee et al., IntelliSys 2025 (doi:10.1007/978-3-031-99965-9_39) · Insilico rentosertib Phase IIa, Nature Medicine (nature.com/articles/s41591-025-03743-2; NCT05938920) · FDA draft guidance, docket FDA-2024-D-4689 (fda.gov/media/184830/download) · FDA AIM-NASH qualification, 8 Dec 2025 · FDA Warning Letter 320-26-58, 2 Apr 2026 · FDA CDER AI programme page · Draft EU GMP Annex 22, EudraLex Vol. 4 · Regulation (EU) 2026/1744 (EUR-Lex) · Tufts CSDD, Ther Innov Regul Sci 2025 · Pistoia Alliance, Sep 2025 · ZS 2025 survey · Cognizant life sciences AI maturity study · Deloitte 2026 life sciences AI survey · Benchling 2026 AI report · Lilly–NVIDIA co-innovation lab, Jan 2026 · Roche–NVIDIA AI factory, Mar 2026 · BMS–Anthropic, May 2026 · Reuters, Roche autonomous labs, 28 Sep 2026 · International Journal of Pharmaceutics 2026, digital and AI-enabled CMC tools survey.
Research notes: [[AI-in-Pharma-Landscape-2026]]
Saram Consulting