There is no universally best scoring function. The reliable path is a target-aware pipeline: a fast docking scorer for pose generation, a complementary rescoring step or target-tuned machine-learning model, and validation using CASF-style metrics and EF1% checks. Services exist that apply this layered logic when advising research teams on which scorer fits a given target and screening goal.
TL;DR:
- Layered scoring pipelines combining fast docking, ML rescoring, and physics-based validation improve hit identification accuracy.
- Empirical and knowledge-based scoring functions excel at pose generation and initial ranking, while ML and physics-based methods refine top candidates.
- Retrospective benchmarks should focus on property-unmatched decoys and out-of-distribution data to avoid overestimating prospective performance.
- Small pilot screens validated with EF1% and RMSD metrics help select and refine the optimal scoring strategy for a specific target.
- Consensus and multi-objective scoring methods are recommended when available data is limited or conflicting, to reduce false positives in screening campaigns.
Table of Contents
- 1. Classification of scoring functions and what they do best
- 2. How researchers benchmark scoring functions with CASF and related metrics
- 3. Selection strategies for target-specific models, consensus, and multi-objective scoring
- 4. A practical virtual screening workflow built around scoring choices
- 5. Avoiding benchmark bias and building a defensible validation plan
- 6. How Innova Biotech evaluates and recommends scoring pipelines
- 7. Where scoring function research is heading next
- Get a validated scoring pipeline for your target
- Sources
- FAQ
1. Classification of scoring functions and what they do best
Choosing among scoring functions starts with knowing what each family is actually built to do. The field groups scoring functions into four broad architectures, and each carries its own trade-off between interpretability, accuracy, and speed, as described in a review of scoring function architectures.
Physics-based scoring functions estimate binding free energy from terms grounded in molecular mechanics: van der Waals interactions, electrostatics, solvation, and sometimes entropy. They are the most mechanistically transparent option, which makes them useful when a researcher needs to explain why a compound binds, not just whether it does. Their cost scales with the rigor of the energy calculation, so full physics-based methods like free energy perturbation (FEP) are rarely used for scanning large libraries.
Empirical scoring functions borrow the same physical intuition but replace exact energy terms with weighted, curve-fit components trained on experimental binding data. They run fast enough for high-throughput docking and remain the default choice in most docking software, precisely because they balance speed against reasonable pose discrimination.
Knowledge-based scoring functions derive interaction preferences statistically from observed atom-pair frequencies in solved protein-ligand structures, converting frequently observed contacts into favorable potentials. They skip explicit physics entirely, which makes them fast and relatively robust to missing parameters, though they can struggle with binding modes underrepresented in the training structures.
Machine-learning scoring functions treat binding prediction as a regression or classification problem over structural and chemical descriptors, using models such as random forests, gradient boosting, or neural networks. Given enough labeled data, machine-learning scoring functions can outperform classical scorers for binding-affinity prediction, and with careful validation can extend to targets outside their training set. The catch is data hunger: without diverse, well-curated training sets, ML scorers tend to memorize dataset-specific patterns rather than binding physics.
Matching class to task looks like this in practice:
- Physics-based: choose when mechanistic interpretation matters more than throughput, such as late-stage lead optimization on a handful of compounds.
- Empirical: choose as the default for pose generation and first-pass ranking across large virtual libraries.
- Knowledge-based: choose when speed matters and the target's binding site resembles well-characterized structural motifs.
- Machine-learning: choose when you have tens or more known actives for the target and need sharper affinity or hit-rate prediction than empirical methods provide.
Most working pipelines do not commit to a single class. A fast empirical or knowledge-based scorer handles pose generation across a full library, while a machine-learning or physics-based method is reserved for a smaller rescoring pass on top candidates. That layering, rather than a single "best" function, is what shows up repeatedly in comparative literature on docking workflows.
2. How researchers benchmark scoring functions with CASF and related metrics
Comparing scoring functions requires a shared vocabulary, and the field has largely converged on the Comparative Assessment of Scoring Functions (CASF) benchmark to supply it. CASF evaluates scoring functions across four distinct powers, and conflating them is one of the most common ways researchers misjudge a scorer's fit for their project.
Scoring power measures how well a function's numerical output correlates with experimentally measured binding affinities, typically reported as Pearson's R across a congeneric or diverse test set. Ranking power asks a narrower question: given a fixed set of ligands for one target, can the function correctly order them by potency, usually assessed with Spearman's rho. Docking power tests whether the function can pick the correct binding pose from a set of decoys, scored by Top-n success rate where a pose within 2 angstroms RMSD of the crystal structure counts as correct. Screening power, the one most relevant to virtual screening campaigns, measures whether a function enriches true actives near the top of a ranked list of a large compound set.
CASF's four-power framework shows that a scoring function strong in one dimension is not automatically strong in another; docking power and screening power are measured with different statistics entirely, which is why a scorer praised for pose accuracy can still underperform in hit enrichment.
Screening power is most often reported as enrichment factor at 1% (EF1%), the ratio of active compounds found in the top 1% of a ranked screening set relative to what random selection would produce. A high EF1% means the scorer concentrates true binders near the top of the list, directly reducing the number of compounds a team needs to test experimentally. A low EF1% on a retrospective benchmark, even alongside a respectable Pearson R for scoring power, is a warning sign that the function will not deliver useful hit rates in a real screen.
Practical interpretation matters more than the raw numbers:
- A Pearson R above roughly 0.6 to 0.7 on a CASF-style scoring test suggests the function tracks affinity trends reasonably well across diverse chemo types.
- Spearman rho closer to 1.0 within a single target's ligand series indicates reliable relative ranking, useful for lead optimization decisions.
- Top-1 docking success above 80% on a benchmark set suggests the function reliably finds native-like poses without needing extensive pose reranking.
- An EF1% noticeably above random selection on a retrospective set is the minimum signal worth carrying into a prospective pilot.
None of these numbers transfer perfectly across targets, which is the core argument for target-aware, rather than blanket, scoring function selection, a point echoed in guidance on validating docking protocol choices.
3. Selection strategies for target-specific models, consensus, and multi-objective scoring
Once you know what CASF's powers measure, the real decision is how to translate that into a pipeline choice for your target. Three practical routes cover most situations, and picking among them comes down to data availability, the outcome you need, and your compute budget.
- Use a general-purpose scoring function when you have few or no known actives for the target and need a quick first pass. Empirical or knowledge-based scorers built into standard docking software work here because they require no target-specific training and generalize reasonably across protein families.
- Build or tune a target-specific machine-learning scorer when you have a meaningful set of known actives, generally in the tens or more, since target-specific ML scoring functions perform best with that scale of labeled data; transfer learning and leave-target-out cross-validation help when the active count is on the low end.
- Apply consensus or hybrid rescoring when no single scorer clearly dominates on your retrospective benchmark, or when you want to reduce sensitivity to any one method's blind spots.
Consensus scoring deserves its own decision logic because it is often the most pragmatic option for teams without the data or time to train a custom model. The simplest version averages normalized scores from several independent scoring functions, and this approach has outperformed individual docking methods across a set of 21 DUD-E targets. More advanced variants use mixture models that combine each scorer's mean and variance to down weight unstable predictors, or gradient-boosting consensus that learns how to combine scorers rather than simply averaging them. Both extensions have been reported to further improve early enrichment compared to plain averaging in benchmark studies.
The trade-offs are straightforward. Plain-average consensus is cheap to implement and requires no additional training data, but it treats all input scorers as equally trustworthy, which is not always true for a given target. Mixture-model and gradient-boosting consensus correct for that but need a validation set to fit the combination weights, adding a data requirement back into the picture. For teams screening against a target with no historical data at all, plain consensus is usually the more defensible starting point; for teams with even a modest retrospective set, the more sophisticated consensus methods tend to pay for themselves in reduced false positives, a theme covered in more detail in guidance on reducing false positives in virtual screening.
Multi-objective training is the newest strategy on this list and addresses a structural weakness in older scoring functions: a model tuned purely for scoring power (affinity correlation) often does not transfer well to screening power (enrichment), because the two tasks reward different aspects of the score distribution. Multi-objective architectures train against several CASF-style objectives simultaneously, which tends to produce a function that generalizes better across scoring, docking, and screening tasks rather than excelling narrowly at one.
Pro Tip: Run your candidate scoring functions through a quick retrospective EF1% check on any historical actives you already have before committing to a full campaign; a scorer with mediocre EF1% on your own target rarely improves once you scale up the library.
The decision rule that ties these together: start from your data budget. Zero known actives points you to a general-purpose or consensus approach. Tens of known actives opens the door to a target-specific ML model. Ambiguous or conflicting retrospective results across multiple candidate scorers is the clearest signal to fall back on consensus rather than betting the campaign on a single function.

4. A practical virtual screening workflow built around scoring choices
The classification and selection logic above only matters if it is embedded in an actual screening pipeline, and the sequence in which scoring functions get applied is as important as which ones you pick.
- Prepare the compound library: filter for physicochemical sanity, remove reactive or unstable groups, and generate consistent protonation states so that downstream scores are comparable across the set.
- Generate poses with a fast docking scorer: an empirical or knowledge-based function handles the initial docking pass across the full library, since speed is the priority at this stage and only rough discrimination is needed.
- Rank by the primary docking score: take the top slice of the library, typically the highest few percent, forward for more expensive treatment.
- Apply an ML rescoring layer: methods in the style of ΔVinaRF20 rerank the top-docking subset using a model trained to correct systematic biases in the primary scorer, which is a lower-cost step than full physics-based rescoring.
- Rescore the smallest surviving set with MM/GBSA or FEP: reserve these physics-based methods for a shortlist, since MM/GBSA and FEP rescoring are higher-cost options typically applied late in a cascade once the candidate pool has already been narrowed.
- Triage for experimental testing: combine the rescoring output with chemical diversity and synthetic accessibility to select the final compounds for assay.
This cascade structure exists because effective pipelines commonly pair a fast Vina-style docking pass with an ML rescoring layer, then reserve MM/GBSA or FEP for a small final set, which keeps total compute cost proportional to how far a compound survives in the funnel.
Compute budget should decide where the cascade stops, not enthusiasm for a particular method:
- ML rescoring on the ΔVinaRF20 pattern is affordable enough to apply to several thousand top-ranked poses.
- MM/GBSA ensembles are affordable for a few hundred compounds when run with modest sampling, but costs climb quickly with more elaborate conformational sampling.
- FEP is reserved for tens of compounds at most, typically the final candidates under serious consideration for synthesis, because its accuracy comes from extensive simulation rather than a fast approximation.
Validation should happen in two stages before committing resources to a full campaign. Second, run a small prospective pilot, typically a few hundred compounds carried through to experimental testing, before scaling to the full library. This two-stage check catches a common failure mode where retrospective enrichment looks strong but does not translate to real hit rates, a gap covered further in a broader virtual screening workflow guide.
5. Avoiding benchmark bias and building a defensible validation plan
The single most common way teams overestimate a scoring function's real-world performance is training or validating on property-matched decoy sets. Standard benchmarks like DUD-E generate decoys that match known actives on bulk physicochemical properties (molecular weight, logP, charge) while differing in topology, which sounds rigorous but lets a model learn superficial physicochemical distinctions rather than genuine binding physics. A scoring function can post excellent retrospective numbers on property-matched decoys and still fail on a real screening library, because real inactive compounds are not filtered to resemble actives on those same properties.
Two corrections address this directly. Property-unmatched decoy sets, or decoys drawn from a more chemically diverse pool, give a more honest estimate of screening power because the model cannot rely on shortcut features. HTS-derived test sets, built from actual high-throughput screening results rather than curated decoy generation, are more predictive of prospective success because they reflect the messy, unfiltered composition of a real compound library.
Overfitting to retrospective benchmarks is a documented risk: strict train-test separation and realistic decoy composition are necessary to avoid overoptimistic claims about a machine-learning scoring function's prospective utility, a finding that applies just as much to consensus and hybrid methods built on top of ML components.
Leave-target-out cross-validation is the other essential guardrail, particularly for ML scorers. Training and testing on the same target family inflates apparent performance, since the model can memorize target-specific patterns rather than learning transferable binding rules. Holding out entire target families during validation, rather than just holding out individual compounds, gives a more realistic estimate of how the scorer will behave on the next, unrelated target you screen.
A practical validation checklist before trusting any scoring function on a new campaign:
- Confirm EF1% on a property-unmatched or HTS-derived retrospective set, not just a standard curated benchmark.
- Check ROC-AUC across the same set to confirm the enrichment signal is not an artifact of one or two outlier actives.
- Verify Top-n docking success with RMSD below 2 angstroms on a crystal structure test set relevant to your target family.
- Visually inspect a sample of top-ranked poses for chemically implausible geometries that a numeric score alone can miss.
- Run a small prospective screen, even a few hundred compounds, before scaling the full campaign.
6. How Innova Biotech evaluates and recommends scoring pipelines
Innova Biotech Solutions approaches scoring function selection as a diagnostic problem rather than a default setting. Engagements typically begin with a target data audit: cataloging known actives, existing structural data, and any prior screening results that can seed a retrospective benchmark.
Where the data supports it, this audit determines whether a target-specific model is worth building or whether a consensus approach across several general-purpose scorers is the more defensible starting point given the available actives. The output of this stage feeds directly into a small pilot screen, structured to test the chosen scoring cascade, ML rescoring, MM/GBSA, or FEP as budget allows, on a manageable compound set before committing to a full campaign.
A typical pilot engagement delivers a benchmarking report showing retrospective performance on the client's own target, a recommended scoring cascade with documented rationale, and a prioritized hit list from the pilot screen ready for experimental follow-up. Clients working across virtual screening and hit-to-lead projects, protein engineering, or enzyme optimization campaigns receive scoring recommendations calibrated to the specific target class and data available, rather than a one-size-fits-all pipeline.

7. Where scoring function research is heading next
The clearest trend worth tracking is the shift away from single-objective scoring functions toward architectures trained across multiple CASF-style tasks at once. Frameworks such as DeepRLI optimize scoring, docking, and screening jointly rather than treating them as separate problems, and early results suggest this reduces the trade-off where a function tuned for pose accuracy underperforms on hit enrichment.
The harder, less glamorous trend is the field's slow correction toward honest validation. Leave-target-out splits and out-of-distribution test sets are becoming less optional and more expected in any serious benchmarking claim, because the earlier generation of retrospective results, built on property-matched decoys, overstated how well many methods would generalize.
My honest read is that the next real gains will come less from a smarter architecture and more from researchers being stricter about what they validate against. A modest ML scorer tested honestly on out-of-distribution data is more useful than an impressive one tested only on its own training distribution. Teams that build a habit of small prospective pilots, rather than trusting retrospective numbers alone, will adapt faster as scoring methods keep evolving.
— Hooman
Get a validated scoring pipeline for your target
Choosing the right scoring cascade for a new target takes retrospective benchmarking, careful decoy selection, and a pilot screen before any full campaign makes sense, and that groundwork is exactly what slows most teams down. Innova Biotech Solutions runs this process as a structured engagement rather than a generic docking run.

A pilot engagement with Innova Biotech typically includes:
- A target-specific retrospective benchmark using EF1% and related CASF-style checks on your own data.
- A recommended scoring cascade, whether target-tuned ML, consensus, or a rescoring sequence, with documented rationale.
- A small pilot screen producing a prioritized hit list ready for experimental testing.
Teams working on enzyme targets can pair this with enzyme optimization services, and those designing novel binders can extend the same benchmarking logic into peptide design projects. Visit the virtual screening and hit-to-lead page to request a pilot evaluation or technical consultation.
Sources
- Comparative Assessment of Scoring Functions (CASF) — community reference
- Review of scoring function architectures and trade-offs
- Performance of machine-learning scoring functions in structure-based virtual screening | Scientific Reports
FAQ
What are scoring functions?
Scoring functions are computational models that estimate how favorably a small molecule binds to a protein target, typically by combining structural and physicochemical features into a single numerical score. They range from physics-based energy calculations to machine-learning models trained on known binding data, as outlined in a review of scoring function architectures.
What are the differences between LBDD and SBDD?
Ligand-based drug design (LBDD) infers what makes a molecule active by comparing it to known active compounds, using similarity or pharmacophore models, without requiring a solved protein structure. Structure-based drug design (SBDD) uses the three-dimensional structure of the target protein directly, docking candidate molecules into the binding site and scoring the resulting poses. Scoring function selection matters primarily in the SBDD context, since it governs how docking poses are ranked.
How do I interpret docking scores?
A docking score reflects the predicted binding favorability under a given scoring function's assumptions, so it is only meaningful relative to other scores from the same function and the same target, not as an absolute energy value. Cross-target or cross-method comparisons require normalization or validation against known actives before the numbers can be trusted for ranking decisions.
What is considered a good docking score?
There is no fixed numeric threshold that defines a good docking score across all targets, since scores depend on the scoring function, the target, and the ligand set used. A more reliable signal is relative performance on validation metrics such as Top-1 docking success with RMSD under 2 angstroms and EF1% enrichment on a retrospective benchmark for your specific target.
How do I choose a scoring function for a new target?
If you have tens of known actives, a target-specific machine-learning model is worth building; if not, a consensus of general-purpose scorers is typically the safer starting point.
