← Back to blog

Make In Silico Selectivity Predictive for R&D With Physics and ML

September 16, 2026
Make In Silico Selectivity Predictive for R&D With Physics and ML

In silico selectivity is the computational prediction of how strongly a compound prefers one biological target over related off-targets, using docking, molecular dynamics, alchemical free-energy calculations, and machine learning. These methods are reliable enough to prioritize synthesis and assay order when validated against a known ligand set for your target family. Treat any unvalidated score as a hypothesis, not a verdict, and confirm the top calls experimentally before committing resources.


TL;DR:

  • Validated ensemble docking and molecular dynamics can reliably produce per-target binding profiles, but calibration against known ligands is essential for accurate selectivity ranking.
  • Shortcomings such as single-structure bias, inconsistent grid alignment, and uncalibrated scoring functions often lead to false selectivity signals and should be carefully managed.
  • Free-energy calculations improve prediction for congeneric series but require targets to share structural similarity for error cancellation and are limited by systematic versus statistical error.
  • Machine learning models are most effective when trained on comprehensive, consistent data, and generate rapid multi-target affinity predictions that complement physics-based methods.
  • A stepwise approach combining docking, rescoring, MD, and selective FEP, with validation against experimental data, optimizes the efficiency and reliability of in silico selectivity campaigns.

Innovabiotech
Build Better Selectivity Predictions
Innovabiotech provides tailored bioinformatics and computational biology solutions for virtual screening, molecular design, and drug discovery projects.
Explore Innovabiotech

Table of Contents

What in Silico Selectivity Means and How You Measure It

Selectivity is a ratio, not a single binding affinity. The most common quantitative expression is the Selectivity Index (SI), typically written as IC50(off-target)/IC50(target), though variants exist depending on the assay context. In antimicrobial and anticancer work, researchers often calculate SI as IC50(normal cells)/IC50(cancer cells) or IC50(normal cells)/MIC, which tells you how much therapeutic window separates efficacy from toxicity rather than how tightly a molecule binds in isolation.

A compound with IC50 = 5 nM against your target and IC50 = 500 nM against a homologous off-target has an SI of 100. That number looks clean on a spreadsheet, but it hides a lot. SI compresses two separate questions, how potent is this molecule, and how much does it discriminate, into one figure. A weak binder against both targets can post a respectable SI while being useless in a cellular assay. A potent pan-inhibitor can post a poor SI while still being druggable if the off-target activity is clinically irrelevant.

That is why target-specific compound selectivity work frames the problem as bi-objective optimization instead of collapsing everything into one ratio. You track potency and relative selectivity as two axes and look for compounds that clear a minimum threshold on both, rather than chasing the single highest SI value in your dataset. This matters most in multi-target drug discovery and repurposing campaigns, where a "top" molecule by raw SI can rank poorly once you also weight absolute potency.

Mapping computational outputs onto these metrics takes some translation work:

  • Docking scores (kcal/mol, or unitless scoring-function values) rank relative binding across a target panel but rarely convert cleanly to IC50 without calibration against known actives.
  • ΔΔG values from alchemical free-energy calculations estimate the difference in binding free energy between two ligands or two targets, which maps directly onto a log-scale selectivity ratio once converted.
  • ML classifier probabilities (a model's confidence that a compound is active against Target A but not Target B) need calibration curves before you can treat them as anything like an SI proxy.
  • Ensemble docking scores across a kinase or protease panel give you a per-target profile you can rank ordinally, even when the absolute numbers don't reproduce experimental IC50 values.

None of these outputs are interchangeable with wet-lab SI. They are ranking tools. Use them to decide what to test first, not to decide what to report as final.

Structure-Based Methods: Docking, Ensemble Docking, and Molecular Dynamics

Single-structure docking against one crystal structure is fast and cheap, and it is also the method most likely to mislead you on selectivity. A single static pocket conformation captures one moment in a protein's conformational life. If your target and its homolog share 70% sequence identity in the binding site but differ in loop flexibility, a rigid-receptor docking run can easily miss the structural feature that actually drives discrimination.

Ensemble docking addresses this by docking the same ligand set against multiple structures or multiple targets simultaneously. An ensemble docking algorithm validated across 14 human kinases achieved performance comparable to single-target docking while generating per-target scores usable for selectivity ranking across the whole panel. That per-target output is the real value here: instead of one binding score, you get a profile showing where a compound is predicted to bind tightly and where it should fall off, which is exactly the shape of data selectivity decisions need.

Molecular dynamics earns its place when docking's static snapshot isn't enough. Running MD on your top-scoring poses lets the receptor breathe. It samples side-chain rotations, loop movements, and occasionally reveals cryptic pockets that don't exist in the deposited crystal structure at all. For kinases and proteases in particular, where selectivity often hinges on a flexible gatekeeper residue or a transiently open subpocket, an MD trajectory can surface discrimination that a single frame never would.

A practical sequence for a target panel:

  1. Prepare all structures with consistent protonation states and identical treatment of conserved waters before any comparison is meaningful.
  2. Dock the full ligand set against the full target ensemble, not one target at a time, so scores are generated under identical conditions.
  3. Cluster docking poses and run short MD (10 to 50 nanoseconds per system, depending on your compute budget) on the top clusters per target.
  4. Re-rank using MD-averaged contacts or MM/GBSA rescoring rather than the raw docking score alone.
  5. Flag any target pair where the ranking flips between docking and MD, since that instability is itself informative.

Grid alignment is where a surprising number of selectivity campaigns go wrong. If you align docking grids to different reference atoms across targets, or use inconsistent box sizes, you introduce a systematic bias that looks like a selectivity signal but is really a setup artifact. Protonation state mismatches are the same problem in a different guise. A histidine that's neutral in one structure and charged in another will shift binding scores by an amount that swamps a genuine selectivity difference.

Pro Tip: Before trusting a docking-derived selectivity ranking, run your entire target panel against a small set of known selective and known promiscuous ligands. If the method can't reproduce that known ordering, recalibrate your scoring function before touching your real compound set. This single sanity check, referenced in guidance on scoring calibration for ensemble docking, catches more bad campaigns than any other single step.

Conserved water treatment deserves its own mention because it's so often handled inconsistently. Some binding sites have a structural water that bridges ligand and protein across an entire target family, in which case removing it in one structure and keeping it in another will bias your comparison. Decide your water-handling rule once, for the whole panel, and document it before you start scoring.

Free-Energy Calculations and Rescoring for Fine Selectivity Tuning

Alchemical free-energy methods, meaning Free Energy Perturbation (FEP) and Thermodynamic Integration (TI), earn their computational cost when you're optimizing selectivity within a congeneric series rather than screening a diverse library. If you have a core scaffold and you're trying to decide whether adding a methyl group here or a fluorine there shifts binding preference toward Target A and away from Target B, FEP is the right tool. It's not the right tool for ranking 50,000 diverse compounds against a single target, that's what docking and ML are for.

The mechanism that makes FEP useful for selectivity specifically, rather than just affinity, is error correlation. Systematic errors in free-energy calculations often correlate between structurally similar targets, meaning if your force field or your simulation setup overestimates binding to Target A by some fixed amount, it frequently overestimates binding to Target B by a similar amount. When you subtract the two ΔG values to get ΔΔG for selectivity, that shared systematic error partially cancels out. The result is that relative selectivity predictions can be more trustworthy than either absolute affinity number alone.

That cancellation is not guaranteed, and it's not free. It depends on your two targets being similar enough in structure and dynamics that the same force field weaknesses manifest the same way in both systems. A kinase and a completely unrelated off-target from a promiscuity panel are not, because there's no shared systematic bias to cancel.

Free-energy calculations for protein-ligand binding now approach accuracy on the order of 1 kcal/mol in favorable cases, and longer simulation times reduce statistical error and improve relative selectivity predictions when systematic error is already correlated between the targets being compared.

That 1 kcal/mol figure matters because it sits right at the edge of what separates a 10 fold selectivity window from a 50 fold one on a log scale. It's precise enough to guide medicinal chemistry decisions on a congeneric series, but not precise enough to treat any single ΔΔG value as gospel without replicate runs.

MM/GBSA and similar endpoint methods sit between docking and full alchemical FEP on the cost curve. They're faster because they skip the alchemical pathway and estimate binding free energy from MD snapshots directly, at the cost of accuracy. Use MM/GBSA as a filtering step to shrink a candidate list from hundreds down to the dozens you'll actually run through full ensemble rescoring or FEP. Treating MM/GBSA output as final selectivity data is a common and avoidable mistake.

The distinction between systematic and statistical error is worth internalizing precisely because it determines how you interpret your error bars:

  • Statistical error comes from finite sampling, insufficient simulation time, or too few independent replicates, and you reduce it by running longer or running more replicates.
  • Systematic error comes from force field inaccuracy, incomplete conformational sampling of a slow degree of freedom, or protonation state assignment, and no amount of additional sampling time fixes it.
  • A ΔΔG value with a small statistical error bar but large uncorrected systematic error will look precise and be wrong.
  • The safest selectivity predictions come from target pairs where systematic error is likely to be correlated, and where you've run enough replicates to shrink the statistical component to below the signal you're trying to detect.

Machine Learning and De Novo Design for Selectivity Prediction

Machine learning models complement physics-based methods rather than replacing them, particularly when you have enough historical assay data to train on. A supervised model trained on high-throughput screening or biophysical binding data across a target panel can output a multi-target affinity profile for a new compound in milliseconds, something no amount of docking or MD throughput can match.

The strongest recent demonstration of this comes from enzyme engineering rather than small-molecule screening. An ML-based method used HTS data to design an N-TIMP2 variant with enhanced selectivity for MMP-9 over related matrix metalloproteinases, and the design was validated experimentally. The catch, and it's an instructive one, is that the selectivity gain came with a tradeoff in absolute affinity. The variant discriminated better between MMP family members, but it wasn't simply "better" in every dimension. That's the bi-objective tension from the metrics section showing up again at the protein engineering level, not just the small-molecule level.

Protein variant balancing affinity and selectivity

Graph Neural Networks (GNNs) and AutoML pipelines have become the default tools for ADMET and off-target risk prediction, largely because they handle the messy, sparse, multi-assay datasets that selectivity work generates better than fixed-descriptor models do. If you want a working sense of where these models are reliable and where they aren't for a real pipeline, the practical breakdown in when AutoML and GNN predictions work for ADMET is a useful reference point, and the same caveats about training data density apply to off-target selectivity models.

A few things to keep in mind when you're evaluating or building ML models for selectivity work:

  • Model performance degrades fast outside the chemical space it was trained on, so a model trained on kinase inhibitors will not reliably score GPCR ligands.
  • Off-target panels with sparse assay coverage (a handful of tested compounds per off-target) produce models that overfit to noise rather than learning real discrimination.
  • Predicted probabilities from classifiers need calibration curves against a held-out test set before you treat them as anything like a selectivity ratio.
  • Ensemble approaches combining multiple ML architectures with physics-based rescoring outperform any single method in most published benchmarks.

De novo design pushes this further by generating candidate structures conditioned directly on a selectivity objective, rather than screening an existing library and hoping something selective turns up. 3D diffusion models and joint 2D and 3D generative architectures now outperform earlier 2D-only generation methods for conditional molecule generation, including generation conditioned on selective binding profiles. The unresolved problem is synthetic accessibility. A generative model can propose a beautifully selective molecule that no synthetic chemist would touch, and scalability to larger, more complex scaffolds remains a genuine limitation of the current generation of these tools rather than a solved problem.

Building a Selectivity Campaign: A Stepwise Protocol

A selectivity campaign that skips validation checkpoints tends to produce confident-looking rankings that fall apart the moment you run the first assay plate. The protocol below balances throughput against accuracy at each stage, front-loading the cheap filters and reserving expensive calculations for the compounds that survive them.

Seven-stage selectivity campaign workflow

Step 1: Assemble your target panel and dataset. Decide which off-targets actually matter clinically or mechanistically before you start, rather than testing against every homolog you can find a structure for. Pull curated bioactivity data, ideally from a single consistent assay format across the panel, since mixing IC50 values measured under different assay conditions across targets introduces noise you'll mistake for selectivity signal later.

Step 2: Prepare every structure to the same standard. This is the step most campaigns underinvest in. Your checklist should cover consistent protonation state assignment, resolved metal coordination geometry where relevant, explicit handling of missing loops (model them or exclude affected residues from scoring, but don't leave them ambiguous), and a documented water-treatment rule applied uniformly across the panel. The practical guidance on protein-ligand docking setup covers most of this ground in more depth than a checklist can.

Step 3: Run ensemble docking across the full panel. Score every compound against every target under identical grid and parameter settings. This is your first-pass filter, cheap enough to run against thousands of compounds and reliable enough to eliminate the clearly non-selective bottom half of your library.

Step 4: Rescore the survivors. Apply MM/GBSA or ML-assisted rescoring to the top-ranked poses from Step 3. This is where you catch cases where the raw docking score got the ranking wrong because it couldn't properly weight a polar contact or a desolvation penalty. A virtual screening workflow built around this kind of staged filtering typically cuts a starting library by one to two orders of magnitude before anything expensive happens.

Step 5: Run short MD on the finalists. Fifty or so compounds surviving to this stage is a reasonable number for MD-based conformational sampling. This step catches the receptor flexibility and cryptic pocket effects that both docking and rescoring miss by construction.

Step 6: Reserve FEP for the closest calls. By this point you should have a shortlist in the single digits to low tens. Run alchemical free-energy calculations on this final set, focused on the target pairs where discrimination matters most and where you have reason to expect correlated systematic error.

Step 7: Validate against experiment, with statistics, not just a table. Calculate correlation metrics (Spearman rank correlation is usually more appropriate than Pearson for selectivity rankings, since you care about ordering more than absolute values) between predicted and measured selectivity for a known reference set before trusting the method on new compounds. Permutation testing helps you establish whether an observed correlation is likely to be real or could plausibly arise from a randomly shuffled ranking. Expect meaningful effect sizes, not perfect agreement. A validated pipeline that reliably identifies the top quartile of truly selective compounds is genuinely useful, even if its rank-order correlation with experimental data is moderate rather than exceptional.

Pro Tip: Run your validation step before your production campaign, using a retrospective dataset where you already know the experimental selectivity outcome. If your pipeline can't recover known selective and known non-selective compounds in roughly the right order on data you already have, it won't magically work better on compounds where you don't know the answer.

The documented ensemble docking workflow with rescoring and MD follow-up reflects roughly this same staged logic, and it's worth treating as a template rather than reinventing the sequence from scratch for every new target panel.

What the Published Case Studies Actually Show

Three published results give a realistic picture of where in silico selectivity prediction currently delivers and where it still needs a human checking its work.

The kinase ensemble docking study is the clearest positive case. Testing across 14 human kinases, the ensemble method matched single-target docking accuracy while producing the per-target score profiles that make selectivity ranking possible in the first place. That's a meaningful result because kinases are notoriously difficult to discriminate between computationally, given how conserved the ATP-binding pocket is across the family. A method that holds up on kinases is also a reasonable bet to hold up on other conserved-fold target families.

The free-energy analysis published in the Journal of Chemical Information and Modeling offers a more nuanced lesson. It found that systematic errors correlate between similar targets often enough that relative selectivity predictions can outperform absolute affinity predictions, but this only holds when the targets share enough structural similarity for the same force field weaknesses to manifest in both systems. Apply FEP to two dissimilar targets expecting the same error cancellation, and you're extrapolating past what the method has actually demonstrated.

The enzyme engineering case, the ML-designed N-TIMP2 variant targeting MMP-9 selectivity over related metalloproteinases, is the most instructive because of its honesty about tradeoffs. The machine learning method successfully increased selectivity, validated by wet-lab assay, but the variant also showed reduced absolute affinity compared to the parent molecule. That's not a failure. It's exactly the bi-objective tradeoff that a single selectivity index would have obscured if the researchers had only reported one number.

Across these three case studies, the pattern that repeats is not "computation replaced the assay." It's "computation narrowed the field enough that the assay became affordable and interpretable." Ensemble docking cut kinase candidates to a rankable shortlist. Free-energy calculations flagged which target pairs were worth the FEP budget. ML flagged a single engineered variant worth testing rather than a library.

The consistent lesson across all three: pair computational evidence with a targeted experiment and an actual statistical test of whether your predicted ranking matches the measured one, rather than treating a strong-looking in silico score as a stopping point. None of these methods are described in their original literature as a replacement for the assay. They're described as a way to make the assay cheaper by asking it fewer, better-chosen questions.

Where In Silico Selectivity Predictions Go Wrong

Most failed selectivity campaigns trace back to a small set of repeatable mistakes, not to some fundamental limit of the computational methods themselves.

Single-snapshot bias tops the list. Scoring one crystal structure or one MD frame per target and treating that score as representative ignores the fact that binding sites move. A loop that's closed in your reference structure might be open in solution, and a rigid-receptor comparison will systematically misjudge any target where that flexibility differs between your target and its off-targets.

Scoring miscalibration is close behind. Docking scores from different software packages, or even the same package with different parameter sets, are not directly comparable to each other or to experimental IC50 values without a calibration step against known actives. Skipping that calibration is how a technically correct-looking ranking turns out to be arbitrary.

Sparse assay data quietly undermines both ML models and your ability to validate any method at all. If you only have five known active compounds for one target in your panel, no model, physics-based or learned, can reliably tell you where the sixth compound will fall.

Neglecting correlated error shows up specifically in free-energy work, where an unvalidated assumption that systematic errors cancel between targets can produce a confidently wrong ΔΔG.

Mitigations map fairly directly onto each pitfall:

  • Use ensemble docking or MD sampling instead of single-structure scoring whenever target flexibility is a plausible factor in discrimination.
  • Calibrate every scoring function against a reference set of known selective and non-selective ligands before trusting it on new compounds.
  • Curate and, where possible, expand assay data before training any ML model, and be explicit about which chemical space the model is actually trained to cover.
  • Run sensitivity analysis on your FEP results by testing whether small changes in force field or sampling protocol shift your selectivity ranking; if they do, your result is fragile.

Pro Tip: If your in silico ranking survives a scrambled-control test (known non-selective compounds don't falsely score as selective, and known selective compounds don't falsely score as promiscuous) and a sensitivity check, trust it enough to prioritize. If it doesn't survive either check, stop trusting the ranking and go straight to the assay for your top candidates instead.

The common causes of false positives in virtual screening overlap heavily with the pitfalls above, and the mitigations there apply just as directly to selectivity ranking as to primary hit identification.

How Innovabiotech Applies These Methods in Practice

A computational selectivity workflow can follow the staged logic laid out above, scaled to the specific target family and compound class a project calls for. Virtual screening and ensemble docking form the front end of most engagements, generating the per-target score profiles needed to rank a starting library before anything more expensive gets touched.

From there, the service scope typically maps like this:

  • Molecular dynamics and structure preparation work supports campaigns where receptor flexibility or a suspected cryptic pocket is likely driving the discrimination researchers are trying to engineer or predict.
  • Free-energy and rescoring methods get applied selectively, on congeneric series and prioritized target pairs, rather than run indiscriminately across an entire panel.
  • Machine learning models support both off-target risk flagging and, in protein and peptide design projects, direct optimization toward a selectivity objective.
  • Protein engineering and chimeric protein design and enzyme optimization work draw on the same ML plus physics-based hybrid approach described in the enzyme case study above, adapted to each project's specific target family.

For researchers who want to go deeper on any single piece of this before scoping a project, Innovabiotech's technical posts on off-target interaction prediction and computational peptide screening walk through method choices in more detail than a single article can cover. The approach throughout stays grounded in the validation-first discipline this article has argued for: calibrate against known data, quantify your error, and let the computation narrow the field rather than replace the assay.

What Actually Deserves Investment Right Now

The near-term direction in this field isn't more exotic algorithms. It's better integration between what physics-based methods and ML each do well. Hybrid pipelines, where an ML model handles rapid triage across a large library and physics-based FEP handles fine discrimination within the shortlist, are already outperforming either approach run in isolation on the published benchmarks I've cited throughout this piece. Expect that hybridization to deepen rather than get replaced by a single dominant method.

Three-dimensional generative models are the piece worth watching most closely for de novo selectivity work, precisely because they're improving fast on the exact problem, conditional generation toward a selectivity objective, where 2D methods historically struggled. Synthetic accessibility remains the bottleneck, and I'd bet against anyone who tells you that's solved.

My honest advice to any lab standing up a selectivity pipeline for the first time: spend your early budget on curating a multi-target dataset with consistent assay conditions, not on licensing the fanciest new software. A mediocre method run on clean, consistent data will outpredict a state-of-the-art method run on a messy dataset assembled from six different assay formats. Rigorous benchmarking against known selective and non-selective reference compounds, run before your real campaign, is not optional overhead. It's the only way you'll know whether to trust the ranking your pipeline hands you.

The field would benefit enormously from more shared, public multi-target benchmark sets, especially for target families beyond kinases, where most of the validated methodology currently concentrates. Until that exists more broadly, the burden falls on individual labs to build and share their own calibration sets, and that's worth the investment.

— Hooman

A Faster Path to Validated Selectivity Data

Every method in this article, ensemble docking, MD, FEP, ML-driven design, takes real infrastructure and real expertise to run correctly, and getting any one of them wrong quietly produces a confident, wrong answer. These pipelines are often run as a staffed service rather than a toolkit you have to assemble and validate yourself, which can save the months most labs spend debugging scoring calibration before generating a single trustworthy result.

Innovabiotech

An initial engagement typically starts with scoping your target panel and reviewing what data you already have, then moves to a pilot deliverable, often an ensemble docking or MD-based selectivity ranking on your priority compounds, before committing to the deeper FEP or ML design phases. Whether your project centers on small-molecule selectivity across a homologous target family, enzyme variant optimization, or peptide design work where discrimination between related receptors matters most, the same validation-first approach applies. If you have a target panel and a question about which candidates deserve synthesis first, reach out through Innovabiotech's protein design services to scope a pilot.

FAQ

What does "in silico" mean in research?

In silico describes any experiment or prediction performed by computer simulation rather than in a living organism (in vivo) or in a test tube (in vitro), covering methods from docking to machine learning-driven simulation.

What does drug selectivity mean?

Drug selectivity is a compound's tendency to bind and act on its intended target while sparing related off-targets, most often quantified with a Selectivity Index comparing potency against the intended target versus an off-target or normal tissue.

What is an in silico model?

An in silico model is a computational representation, structural, statistical, or learned, that predicts a biological or chemical outcome, such as binding affinity or selectivity, without requiring a physical experiment for every prediction.

What is in silico drug screening?

In silico drug screening uses computational methods like virtual screening and ensemble docking to rank large compound libraries against one or more targets, narrowing thousands of candidates down to a shortlist worth testing experimentally.

How do you assess selectivity computationally?

Assessing selectivity computationally means comparing a compound's predicted binding behavior across a target panel, typically through ensemble docking scores, ΔΔG values from free-energy calculations, or calibrated ML predictions, then validating the resulting ranking against known experimental selectivity data.