Bayes Pharma · Synthesis Intelligence

Retrosynthesis has moved beyond one model, one route.

Explore how expert systems, template models, sequence transformers, graph editors, ensembles, flow models and chemistry LLMs differ—and where each earns a place in modern synthesis planning.

Science-first: reported benchmark numbers are not treated as directly comparable across datasets, splits or evaluation protocols.
20models & planners mapped
7architectural families
3planning layers
1rule: no single leaderboard tells the whole story
The landscape

Every generation solved a different bottleneck.

Retrosynthesis systems differ in what they represent, how they propose disconnections, whether they use explicit reaction precedent, and whether they can plan full routes. Newer does not automatically mean better for every chemistry problem.

01Expert era

Rules & precedent

Human-coded transformations and analog retrieval. Interpretable, but coverage depends on the rule base.

LHASA · RetroSim
02Neural-symbolic

Learn which rule fits

Neural networks rank templates instead of applying every possible transformation.

NeuralSym · GLN · LocalRetro
03Sequence era

Chemistry as translation

SMILES-to-SMILES generation unlocked flexible template-free prediction at scale.

Molecular Transformer · Chemformer · R-SMILES
04Graph era

Reason on molecular structure

Graph models encode bonds, atoms and reaction centers directly instead of inferring them from text alone.

Graph2Edits · NAG2G
05Frontier era

Diversity + convergence

Ensembles and flow models explicitly exploit complementary biases and diverse feasible proposals.

RetroChimera · Retro SynFlow
06Reasoning era

Chemistry-specialist LLMs

Large models add broader chemical priors, Top-K generation and increasingly explicit reasoning—but still need chemistry validation.

RetroDFM-R · Insilico SSRS · C3LM
Interactive model atlas

Compare the model, not just the score.

Search by model, architecture or use case. Add up to four models to the comparison tray to inspect their scientific differences side-by-side.

0 models shown ·
Direct comparison

What actually changes between models?

Representation and search behavior matter as much as headline accuracy. The table below highlights the dimensions that affect real synthesis-planning behavior.

DimensionTemplate / precedentSequence / TransformerGraph / editFlow / ensembleChemistry LLM
Primary representationReaction templates / fingerprintsSMILES tokensMolecular graph, atoms, bondsMultiple model spaces or stochastic pathsTokens + learned chemical priors
InterpretabilityHighMediumHigh–mediumMediumVariable
Uncommon reaction coverageLimited by template libraryCan generalize beyond templatesStrong structural generalization potentialDesigned for breadth / diversityPotentially broad; OOD reliability must be tested
Common failure modeTemplate coverage gapInvalid / implausible sequence generationIncorrect edit sequence / attachmentDiverse but chemically weak proposalsPlausible-sounding hallucination
Best scientific rolePrecedent-aware anchorFast flexible challengerStructural / reaction-center challengerDiversity & consensusReasoning / broad proposal challenger
Needs downstream validation?YesYesYesYesAbsolutely
Model selection

Which family is strongest for which job?

The best practical stack is role-based. Use a model where its inductive bias helps, then challenge it with a fundamentally different method.

Known chemistry & precedent

Favor explicit templates, analog retrieval and systems trained on broad reaction databases.

RetroSim · GLN · LocalRetro · ASKCOS

Complex structural disconnections

Graph-first models make atom/bond changes visible and preserve molecular topology.

Graph2Edits · NAG2G

Diverse hypothesis generation

Use models designed to explore more than one plausible precursor set.

Retro SynFlow · C3LM · Insilico SSRS

Reasoned explanation

LLM-style systems can add explicit rationale and broader chemical context, but require verification.

RetroDFM-R · C3LM

Robust single-step ranking

Use complementary inductive biases rather than trusting one architecture.

RetroChimera

Full route planning

Single-step models need a search layer, stock definition, stopping criteria and route ranking.

AiZynthFinder · ASKCOS · Retro*
BayesARC concept

The Retrosynthesis Convergence Engine.

No single retrosynthesis model becomes “truth.” Independent proposal engines compete, chemistry checks prune weak routes, and route-level evidence stays inspectable through human review.

STARTTarget moleculeSMILES · structure · programme context
LAYER 1Independent proposal enginesDifferent biases · different chemistry
RetroChimeraensembleNAG2GgraphSynFlowflowRetroDFM-Rreasoning LLMC3LM / SSRSchemistry LLMLocalRetrotemplate
GATE 01Canonicalize + deduplicatesame reaction ≠ six votes
GATE 02Forward validationcan precursors plausibly regenerate product?
GATE 03Precedent & noveltyknown transformation vs unsupported leap
GATE 04Chemistry plausibilitystereo · selectivity · protecting groups · conditions
LAYER 2Multistep searchRetro* · MCTS · AiZynthFinder / ASKCOS-style planners
LAYER 3Route-level ranking
stepscoststockhazardstereoprecedentyield riskscale-up
FINAL AUTHORITYQualified chemist reviewPredicted routes remain hypotheses until expert and experimental evidence supports them.
Read the benchmark carefully

A Top-1 number is not a synthesis strategy.

Retrosynthesis is intrinsically one-to-many. Multiple precursor sets can be chemically valid even when only one appears in a benchmark record.

01

Dataset matters

USPTO-50K, USPTO-FULL, Pistachio, CAS-derived and proprietary corpora create very different difficulty and coverage.

02

Split matters

Random, temporal, reaction-class and out-of-distribution splits answer different generalization questions.

03

Ground truth is incomplete

Exact-match evaluation can punish chemically valid alternatives that differ from the recorded literature route.

04

Planning ≠ one step

A strong one-step precursor model can still perform poorly inside a full search tree if it generates weak branches or poor calibration.

Compare modelsSelect up to 4 model cards
Selected models

Scientific comparison