Four findings from the last eighteen months describe the same wall from four different directions.
PoseBusters showed that deep-learning docking produces poses within 2 Å RMSD that are nevertheless physically invalid — on one Astex benchmark, TankBind placed 59% of ligands within 2 Å but only 5.9% passed basic physical validity checks. GenBench3D found that across the standard 3D generative models, only 0–11% of generated molecules had valid conformations at all. BOOM, a benchmark spanning more than 150 model–task combinations, found that not one model generalizes strongly across tasks — even the best model averaged 3× higher error out of distribution than in distribution. And Runs N’ Poses, testing co-folding models on 2,600 systems released after their training cutoffs, concluded they largely memorize ligand poses they have already seen.
None of this means the models are bad. It means they are interpolating, and drug discovery is a search for things that don’t exist yet.
Here is the molecular design stack as it actually stands, with the parts that are reproducible separated from the parts that are marketing.
The stack both modalities converged on
Small molecules and biologics started as separate disciplines and have converged on one architecture: foundation model → conditional generation → property prediction → multi-objective ranking → experimental selection → retraining. The objects differ; the loop does not.
TARGET
│
▼
Structure / co-folding model
(AlphaFold 3, Boltz, Chai)
│
▼
Binding site or epitope
│
▼
┌─ GENERATOR ───────────────────────────────────┐
│ small molecules: diffusion, transformers │
│ biologics: backbone diffusion, pLMs │
└───────────────────────┬───────────────────────┘
│ 10^4 – 10^6 candidates
▼
┌─ PREDICTORS ──────────────────────────────────┐
│ potency · ADMET · affinity · developability │
└───────────────────────┬───────────────────────┘
│ rank + uncertainty
▼
┌─ GATES ───────────────────────────────────────┐
│ 3D validity · stereochemistry │
│ retrosynthesis (AiZynthFinder) │
└───────────────────────┬───────────────────────┘
▼
10–100 compounds
│
▼
WET LAB
│
▼
proprietary data ──► retrain
The last arrow is the only part that compounds. Everything above it is increasingly commoditized.
| Small molecules | Biologics | |
|---|---|---|
| Representation | Graphs, SMILES/SELFIES, 3D point clouds | Amino-acid sequence, backbone frames in SE(3) |
| Search space | ~10⁶⁰ drug-like structures | 20^L sequence space |
| Generative engine | Diffusion, autoregressive transformers, RL | Backbone diffusion + inverse folding, protein LMs |
| Design step | Atom placement / fragment assembly in a pocket | Backbone synthesis → sequence design |
| Kill gate | Synthesizability, ADMET | Developability: aggregation, viscosity, immunogenicity |
| Field maturity | Mature in method, gated by synthesis | Newer, better lab hit rates, weaker verification |
The three tools that keep getting conflated
This is the most commonly garbled part of the literature, including in otherwise careful writing.
A model can produce a geometrically plausible pose that is chemically impossible. Three separate tools measure three separate failure modes, and they get cited interchangeably:
| Tool | What it actually measures | Representative finding |
|---|---|---|
| PoseBusters | Physical validity of a predicted pose: bond lengths, angles, planarity, clashes, volume overlap | Deep-learning docking reaches sub-2 Å RMSD on poses that are physically invalid |
| GenBench3D | Validity3D of a generated molecule: CSD-based bond/angle likelihood, intramolecular clash | Only 0–11% of generated molecules had valid conformations |
| PoseCheck | Intermolecular protein–ligand clashes and pocket volume overlap | Steric clashes with the target, distinct from internal geometry |
Two corrections follow directly. First, the widely repeated “up to 95% invalid poses” is not a PoseBusters figure — it is GenBench3D’s paraphrase of the LiGAN/3D-SBDD results. Second, the common claim that “3D generators fail PoseBusters checks mainly from pocket clashes” is misattributed: GenBench3D’s driver is unrealistic bond lengths and angles, not protein contact. Intermolecular clashes belong to PoseCheck.
That distinction is not pedantic. If the failure is internal geometry, the fix is a valence-aware decoder. If the failure is pocket sterics, the fix is clash-aware guidance. Solving the wrong one is a wasted quarter.
Gate two: synthesizability
Generation is the easy half. A generated molecule is worthless if nobody can make it within a reasonable number of steps.
The field has moved from filtering synthesizability after generation to encoding it during generation:
- RxnMol operates on synthesis routes rather than molecular structures directly — the molecule is a route.
- SynFormer (PNAS 2025) generates synthetic pathways as the design object, “to ensure that designs are synthetically tractable.”
- SynFlowNet paired with AiZynthFinder succeeds on 62% of length-3 and 40% of length-4 retrosynthetic routes. Comparable reaction-based generators land in the 53–87% band depending on route length and the checker used.
Both numbers are the point: even reaction-aware generation loses a third to a half of its output when an external retrosynthesis engine audits it. In-model validity is not real-world validity.
Then there is the screening side. REvoLd, an evolutionary search over the Rosetta docking space, reported hit-rate improvements of 869× to 1,622× over random selection across five targets — the strongest published case that ML-guided selection captures most of the value even when generation is unremarkable. Schrödinger markets a similar result for its active-learning workflow: recovery of roughly 70% of the top-scoring hits for 0.1% of the cost, citing 1 billion compounds at 17,361 compute-days and $1.25M for brute-force Glide versus 37 days and $2,041 with active learning. That one is vendor marketing, not peer review — but the direction is consistent with the open literature, and it has a strategic implication: if a surrogate model can find 70% of the good molecules for a thousandth of the cost, generation quality stops being the differentiator.
Property prediction: dependable, and the place the leaderboards lie
Property prediction is the least glamorous and most reliable contribution ML makes to discovery. It is also where benchmark claims should be read most carefully.
What holds up:
- In the 2025 ASAP-Polaris-OpenADMET blind challenge, “classical methods remain highly competitive for predicting potency, modern deep learning algorithms significantly outperformed traditional machine learning in ADME prediction.” That is a clean, useful split: deep learning buys you ADME, not potency.
- A study training 62,820 models found random forests on RDKit2D descriptors beat MolBERT and GROVER on BACE, BBBP, ESOL and Lipophilicity, at p<0.05.
- Activity cliffs — one-atom changes that move IC₅₀ by three orders of magnitude — remain unsolved in the general case, and descriptor methods specifically outperform deep learning there.
What does not:
- The BACE dataset has 71% of molecules with at least one undefined stereocenter. MoleculeNet’s most-cited ADMET benchmark cannot distinguish enantiomers for most of its own entries.
- Public ADMET data carries inconsistent SMILES, duplicates and contradictory labels; BACE was compiled from 55 heterogeneous assay papers, with aggregation noise that may exceed the signal.
- The TDC ADMET leaderboard does not survive audit. The top-ranked entry is Chemprop-RDKit — a D-MPNN plus RDKit features, not a pure deep model — and a 2026 audit found only 3 of the top 10 models reproducible: one had a data leak, two were evaluated on splits containing TDC test molecules. Chemprop-RDKit’s verified remaining SOTA is solubility.
The practical translation: property models are ranking tools inside a chemical series, not quantitative predictors across chemotypes. That distinction is exactly what limits how aggressively anyone should let a generative model autonomously optimize a multi-parameter profile.
Co-folding: the widest gap between claim and evidence
Structure and affinity prediction is the fastest-moving and most heavily marketed area in the field. It is also where the reproducibility ledger matters most.
| Claim | Status | Detail |
|---|---|---|
| Boltz-2 is open source (MIT / Recursion) | Verified | Public repo and weights |
| Boltz-2 top of CASP16 affinity, r = 0.65 | Company-reported | Authors’ own retrospective table — Boltz-2 was not an official CASP16 entrant. An independent reproduction got ~0.52 |
| Boltz-2 approaches FEP accuracy at ~1000× speed | Company-reported | The compute ratio is real; the accuracy parity is a self-selected benchmark |
| Boltz-2 + generative model found TYK2 binders | Verified | Confirmed by absolute FEP simulation, not wet lab. A later agent-run wet-lab attempt got 3 actives from 23, IC₅₀ ~240 nM |
| Boltz-2.1 is closed and API-only | False | The Boltz-2 line stays open (v2.1.0 → v2.2.1, Sept 2026). Only the newer BoltzMol-1 and BoltzProt-1 are API-only |
| Co-folding models largely memorize ligand poses | Verified | 2,600 post-cutoff systems; models place ligands they have seen |
| Models keep ATP in place after the site is mutated | Verified | Confirmed independently by funnel metadynamics |
| Physics-based docking beats co-folding on novel pockets | Verified | Attracting Cavities and AutoDock Vina outperform co-folding for novel ligands and pockets |
| Even simple fragments default to the orthosteric site | Verified | Allosteric and cryptic sites remain a structural weak spot |
| IsoDDE more than doubles AlphaFold 3 on the hardest novel systems | Company-reported | Feb 2026 technical report, no code, no weights. It additionally claims to beat physics-based FEP+ and OpenFE — a stronger claim than is usually reported |
The honest summary: zero-shot co-folding is a hypothesis generator, not a decision tool. The published practical pattern is to fine-tune the affinity head on project-specific data rather than trust the pretrained head.
Biologics: better hit rates, weaker verification
Protein language models are the enabling layer. IgLM was trained on 558 million antibody heavy/light variable sequences; BALM on 336 million 40%-nonredundant sequences. AlphaFold 3 extended structure prediction to complexes of protein, nucleic acid, ligand and ion, and explicitly models antibody–target binding.
On top of that sits the de novo design program — and here the field splits cleanly into results anyone can reproduce and results anyone can announce.
Reproducible:
- RFdiffusion (Nature 2023): success rates “roughly two orders of magnitude higher” than prior methods on IL-7Rα, InsR, PD-L1 and TrkA. A designed influenza hemagglutinin binder was confirmed by cryo-EM, matching the computational design model to 0.63 Å RMSD. (That is a model-agreement figure, not a map resolution — a distinction that gets lost in retellings; the cryo-EM map itself was 2.9 Å.)
- BindCraft (Nature 2025): reported success rates of 10–100% across targets. Independent critique: of 65 reported binders from 212 designs, only 20 have measured Kd.
- RFdiffusion3 (open preprint): atom-level design, ~10× faster than RFdiffusion2. Of 190 enzyme designs screened, 35 were multi-turnover active, best kcat/Km = 3557 M⁻¹s⁻¹.
- BoltzGen (open, MIT-licensed preprint): 15 designs per target across nine novel targets → nanomolar binders for 66% (6 of 9).
Announced:
| Result | Cautions |
|---|---|
| Chai-2: 16% hit rate across 52 targets with ≤20 designs each; 68% miniprotein success; follow-up reports >86% of full-length mAbs with therapeutic-grade developability | Closed model, own technical report. Independent review flagged survivor bias: 7 of 88 IgGs and 10 of 27 VHH-Fcs were dropped for expression or purity failure before developability statistics were computed |
| Latent-X2: binders for 9 of 18 targets from 4–24 designs; “first” low immunogenicity confirmed in human donor panels | Only 4 VHHs from a single target, tested as Fc fusions, in a 10-donor panel skewed 60% toward HLA-B44, with no positive controls |
| AlphaProteo: 3–300× affinity improvements; sub-nanomolar binders | Closed technical report; not independently reproducible |
| Schrödinger active learning: ~70% of top hits at 0.1% cost | Vendor page, not peer review |
The pattern is not that anyone is lying. It is that the flashiest numbers come from closed platforms, and the moment an outside reviewer recomputes them, the denominator changes.
And one peer-reviewed result deserves to be read as the strategic headline of the whole biologics section: Biogen evaluated protein language models against internal antibody data from 33 historical therapeutic programs, and found that domain-adaptive fine-tuning on internal sequences consistently improved developability prediction over pretrained representations alone. The foundation model is a commodity. The 33 programs are not.
What the molecules actually did in patients
- Rentosertib (Insilico, TNIK inhibitor, IPF). The Phase 2a reported a placebo-arm mean change of −20.3 mL FVC (95% CI −116.1 to 75.6) against +98.4 mL at 60 mg QD (95% CI 10.9–185.9). A widely circulated −62.3 mL placebo figure comes from Insilico’s November 2024 topline and does not appear in the published paper. Seven patients discontinued for liver injury or dysfunction, four of them concurrently on nintedanib. Phase 3 (320 participants, 47 centres, 52 weeks) dosed its first patient on 9 September 2026.
- Zasocitinib (TAK-279, Takeda). Beat deucravacitinib head-to-head in Phase 3 and had its NDA accepted under priority review in September 2026, PDUFA Q1 2027. Nimbus used ML plus free energy perturbation across more than 13,000 computationally assessed compounds. The “first AI-approved drug” framing is contested by the people who made it — Nimbus’s R&D head describes a compound identified in 2020 and rejects the label.
- GB-0895 (Generate:Biomedicines). An anti-TSLP antibody in two Phase 3 trials, ~1,600 patients combined, dosed every six months, which went straight to Phase 3 without a Phase 2 efficacy trial. Worth a precision note: Generate’s own wording is “engineered with AI,” so calling it strictly de novo is arguable — but it does mean the claim “no AI-designed antibody is in late-stage trials” is no longer true.
- Recursion, after its Exscientia merger, deprioritized or terminated REC-2282, REC-994 and REC-3964, paused REC-4539 and cut a preclinical program in May 2025.
- Pipeline count: 117 AI-enabled assets across 63 companies — peer-reviewed with an explicit methodology, 60 having completed Phase I, 8 Phase II, 82% small molecules. The “more than 75” figure is a valid but stale 2024 floor; the “200+” figure discloses no methodology. No AI-discovered drug has FDA approval.
What to do with this
- Classify every tool before you compare its numbers. Peer-reviewed with open code, open preprint, or closed technical report. The category predicts the replication outcome better than the benchmark table does.
- Audit the split before the metric. Time-based and scaffold-based splits, not random ones. If a paper reports one split, treat the number as an upper bound.
- Separate internal from external validity. In-model synthesizability checks, in-model confidence scores, and in-model affinity heads all flatter themselves. Run AiZynthFinder, run PoseCheck, run the docking baseline.
- Fine-tune the head you care about. Pretrained models are starting points; project data is the product. The Biogen result is the template.
- Instrument experimental selection, not just generation. When a model proposes 10⁶ candidates and the lab assays 10², the value is entirely in the ranking and the uncertainty estimate.
- Record the metadata with the measurement. A compound and an IC₅₀ is a weak training example. Compound, target, construct, assay, cell line, protocol, instrument, batch, replicate, raw value, uncertainty — that is what turns a lab into a learning system.
The bottom line
Molecular AI’s architecture has converged, and the generative layer is genuinely capable: diffusion models design in pockets, protein language models and backbone diffusion produce binders that no natural evolution wrote, and reaction-aware generation is closing the synthesizability gap. What the field has not solved — and what every independent evaluation keeps finding — is generalization beyond the training distribution, which is precisely the territory discovery operates in.
Which means the reliable near-term leverage is not in a bigger generator. It is in the gates, the evaluation discipline, the uncertainty estimates, and the proprietary experimental data that makes the next round of fine-tuning better than the last.
Related: 117 AI-Built Drug Programs, Zero Approvals · Knowledge Graphs Without Hallucinated Edges · The Decision Layer: Type-Safe Models Outside Your Validated Write Path
Sources: PoseBusters, Chem. Sci. 2024 (arXiv:2308.05777) · GenBench3D (arXiv:2407.04424) · PoseCheck (arXiv:2308.07413) · DiffDock (ICLR 2023) · TargetDiff (ICLR 2023) · Pocket2Mol (ICML 2022) · DiffSBDD, Nat. Comput. Sci. 2024 · RxnMol (github.com/snu-lcbc/RxnMol) · SynFormer, PNAS 2025 (arXiv:2410.03494) · SynFlowNet (arXiv:2405.01155) · RGFN (arXiv:2406.08506) · REvoLd, Commun. Chem. 2025 (arXiv:2404.17329) · Fischer et al., JCIM 2025 (doi:10.1021/acs.jcim.5c01982) · Deng et al., Nat. Commun. 14:6395 · BOOM, NeurIPS D&B 2025 (arXiv:2505.01912) · van Tilborg et al., JCIM 2022 · Li et al. (arXiv:2604.16586) · Koleiev et al. (bioRxiv 2026.02.26.708193) · Boltz-2 (github.com/jwohlwend/boltz) · Runs N’ Poses, Nat. Struct. Mol. Biol. 2026 · Masters, Mahmoud & Lill, Nat. Commun. 16:8854 · Attracting Cavities benchmark (bioRxiv 2025.12.09.693161) · CAFE (PubMed 42539081) · Isomorphic Labs IsoDDE technical report, 10 Feb 2026 · RFdiffusion, Nature 2023 · BindCraft, Nature 2025 · RFdiffusion3 (IPD, Dec 2025) · BoltzGen (bioRxiv 2025.11.20.689494) · AbDiffuser, NeurIPS 2023 · IgLM, Cell Systems 2023 · BALM, Brief. Bioinformatics 2024 · Biogen antibody PLM study, mAbs 2026 (doi:10.1080/19420862.2026.2647489) · Chai-2 technical report · Latent-X2 (arXiv:2512.20263) · rentosertib, Nature Medicine (s41591-025-03743-2; NCT07687459) · Takeda zasocitinib releases · Generate:Biomedicines GB-0895 PR (NCT07276724) · Schrödinger active-learning applications page.
Research notes: [[AI-in-Early-Drug-Discovery-Molecular-ML-2026]]
Saram Consulting