Virtual screening typically returns hit rates far above brute-force high-throughput screening, but the number you get depends heavily on how many compounds you test and which methods you pair with the screen. Rescoring, focused libraries, and machine learning-guided active learning are the levers that most reliably raise and stabilize those numbers. Testing only a few dozen compounds produces unreliable estimates; testing several hundred gets you closer to the truth.
TL;DR:
- Testing several hundred compounds provides a more accurate and stable estimate of hit rates than testing only a few dozen, which can be highly misleading.
- Larger screening libraries, especially those with billions of compounds, yield more actives, better potency distributions, and more novel scaffolds than smaller libraries.
- Using rescoring, focused libraries, machine learning prioritization, and active learning can significantly improve hit rates, with machine learning reducing docking costs by over 500 times.
- Reporting clear experimental details such as the number tested, hit cutoff, and study type is essential to interpret hit rates reliably.
- Confidence in hit rate estimates improves substantially with larger sample sizes and iterative validation, rather than relying on small, single-screen data.
Table of Contents
- What typical hit rates look like and why reported numbers vary
- How library size and scale of testing affect hit rates
- Practical methods that raise hit rates
- Designing an experiment and validation workflow
- How Innova Biotech Solutions views stabilizing hit rates
- How Innova Biotech can help you improve enrichment
- Sources
- FAQ
What typical hit rates look like and why reported numbers vary
Numbers in the literature swing wildly, and that swing is often about methodology rather than the underlying chemistry. A widely cited review of virtual screening literature found that experimental high-throughput screening typically returns hit rates around 0.01% to 0.14%, while prospective virtual screening campaigns in the same survey ranged much more broadly, with a median near 13%, though definitions of "hit" varied enormously between studies. That gap between HTS and VS is real, but it is also partly an artifact of how each field defines success.

Retrospective studies, where a known active is dropped into a library of decoys and the model is asked to find it again, tend to report cleaner enrichment than prospective campaigns, where nobody knows the answer in advance. Prospective numbers are messier and more representative of what a real project will see. The choice of hit definition matters just as much: an IC50 cutoff of 10 micromolar produces a very different hit count than one set at 1 micromolar, and percent-inhibition thresholds at a single concentration behave differently again from full dose-response criteria.
Ligand efficiency helps correct for a common bias in this process. Larger molecules tend to show better raw potency simply because they have more atoms available to make contacts, which can inflate apparent hit rates if potency alone decides who counts as a hit. The PMC review recommends pairing potency with ligand efficiency and standardizing cutoffs before comparing hit rates across campaigns.
When you read a hit-rate claim, three details tell you whether it is trustworthy:
- Number tested: a hit rate from 20 compounds means something different than one from 2,000.
- Hit cutoff used: IC50, percent inhibition, and ligand efficiency thresholds are not interchangeable.
- Prospective or retrospective: retrospective enrichment numbers usually look better than real-world prospective results.
How library size and scale of testing affect hit rates
Bigger docking libraries do not just return more candidates, they change the quality of what you find, and recent large-scale experimental work quantifies that effect directly. A comparison of docking against a 1.7 billion-compound library versus a 99 million-compound library, described in research on library size and testing scale, found that the larger library produced more actives overall, better potency distributions, and more novel scaffolds. In that study, the larger screen yielded substantially more actives out of a few thousand compounds tested, compared to far fewer actives from a very small test of the smaller library, and the larger library uncovered many more inhibitors overall.
One finding stands out for anyone planning a screen: hit-rate estimates only stabilize after testing several hundred compounds, not a few dozen. Bootstrapping analysis in the same study showed that the standard deviation in hit-rate estimates decreases noticeably as the number of sampled molecules increases from a few dozen to several hundred, for one target.
That instability at small sample sizes has direct consequences for how you plan an experiment:
- Testing only 20 to 50 compounds can make a genuinely mediocre library look excellent, or a good one look mediocre, purely by chance.
- Convergence toward a stable hit-rate estimate happens somewhere in the range of a few hundred molecules, though the exact number depends on the target's underlying bindability.
- Larger libraries increase raw false positives among top-ranked candidates even as they increase true actives, which means bigger screens demand more rigorous downstream filtering, not less.
Targets with naturally high bindability converge faster because a larger share of tested compounds are genuine actives; low-bindability targets need much bigger tested sets before the numbers settle down.
Practical methods that raise hit rates
Several concrete interventions move hit rates upward, and they work through different mechanisms, so the strongest campaigns tend to layer more than one.
- Rescoring and consensus scoring. Combining multiple scoring functions or docking programs, along with accounting for protein flexibility, reduces the artifacts that come from relying on a single rigid-docking protocol, and has repeatedly improved pose prediction and early enrichment across published benchmarks, according to a review of enrichment strategies.
- Focused and target-biased libraries. Enumerating compounds around a known chemotype, or including covalent warheads where the target supports them, can push hit rates into double digits for specific campaigns, and has produced nanomolar leads directly from virtual screening in multiple published cases.
- Machine learning-guided prioritization. Workflows built on conformal predictors and gradient-boosted classifiers, such as CatBoost, can shrink the set that needs full docking by orders of magnitude. A machine learning-guided docking study in Nature Computational Science reduced docking cost by more than 500 times against a multibillion-compound library while still recovering more than 90% of the very top-scoring molecules.
- Active learning with bioactivity feedback. Feeding real assay results back into the model and retraining between rounds, rather than screening once and stopping, is described in research on active learning for virtual screening, which reported consistent hit-rate enrichment across benchmark and prospective tests, along with wider chemical diversity among the hits that emerged.
As libraries climb into the billions, the constraint shifts from raw compute to accurate prioritization, and the Nature Computational Science authors argue that resources are better spent on rescoring and machine learning triage than on docking every candidate.
Pro Tip: Before trusting a machine learning prioritization step, check its recall on a held-out, already-docked set and set your conformal prediction threshold to control how many true actives you are willing to risk missing.
Designing an experiment and validation workflow
A reliable hit-rate estimate starts with a clear hit definition, reported alongside ligand efficiency and enrichment factor, not potency alone. Sample size matters as much as method: dozens of compounds swing wildly with each new result, while several hundred tested molecules give an estimate that holds up under bootstrapped resampling. Budget confirmatory assays accordingly rather than treating them as an afterthought once a primary screen is done.
A workable pipeline moves through defined stages, each catching a different class of artifact:
- Primary screen, followed by counterscreens to remove assay interference.
- Orthogonal biophysical confirmation such as SPR or ITC.
- Potency confirmation through full dose-response curves.
- ADMET triage, then novelty and PAINS filtering to remove promiscuous or previously flagged scaffolds.
| Report element | What to include |
|---|---|
| Number tested | Total compounds screened and assayed |
| Actives found | Count meeting the hit cutoff |
| Potency bins | Distribution across affinity ranges |
| Statistical range | Bootstrap confidence interval on hit rate |
| Study type | Prospective or retrospective, stated explicitly |
How Innova Biotech Solutions views stabilizing hit rates
The consistent thread across recent literature is that hit rates are not fixed properties of a target, they are outcomes of experimental design. Innova Biotech Solutions, a San Francisco-based biotechnology company founded in 2024, builds its virtual screening, hit-to-lead, and protein engineering work around that idea rather than treating a single docking run as a final answer.
The honest failure mode in this field is not a low hit rate, it is a confident hit rate built on too few data points. A team that screens 30 compounds and calls the result definitive is making a statistical error dressed up as a scientific conclusion, no matter how sophisticated the docking software behind it. The fix is not more computing power, it is better experimental design: bigger validated sample sizes, iterative feedback from real assays, and skepticism toward any single number reported without its confidence interval.
— Hooman
How Innova Biotech can help you improve enrichment
Some companies build tailored virtual screening and hit-to-lead pipelines that combine rescoring, machine learning prioritization, and iterative bioactivity feedback, aimed at enrichment you can reproduce and report with confidence.

If you are planning a screen and want a pipeline built around your target's specific bindability profile, get a project consultation through our virtual screening service and talk through scope and next steps with our team.
Sources
- Hit Identification and Optimization in Virtual Screening: Practical Recommendations Based Upon a Critical Literature Analysis - PMC
- The impact of Library Size and Scale of Testing on Virtual Screening - PubMed
- Rapid traversal of vast chemical space using machine learning-guided docking screens | Nature Computational Science
- Improving the Hit Rates of Virtual Screening by Active Learning from Bioactivity Feedback - PubMed
FAQ
What is considered high throughput screening?
High-throughput screening, or HTS, is the automated experimental testing of large physical compound libraries against a biological target using robotics and plate-based assays. It typically returns much lower hit rates than virtual screening, around 0.01% to 0.14% in one literature review.
What is hit finding in drug discovery?
Hit finding is the early-stage process of identifying compounds that show measurable activity against a target, using methods like HTS, virtual screening, or fragment-based approaches. Hits then go through confirmation and triage before being considered for hit-to-lead optimization.
What is a virtual screening?
Virtual screening is a computational method that ranks or filters large compound libraries by predicted binding affinity or activity against a target, before any physical testing happens. It is used to shortlist the most promising candidates for experimental follow-up, which is explained further in this overview of virtual screening.
What are the different types of virtual screening methods?
The two broad categories are structure-based methods, which dock compounds against a known or modeled protein structure, and ligand-based methods, which compare candidate molecules to known actives using similarity or pharmacophore models. Many modern pipelines combine both with machine learning prioritization to manage library scale, an approach detailed in this virtual screening workflow guide.
How many compounds should I test to get a reliable hit rate?
A large-scale study found that testing only a few dozen compounds produces highly variable hit-rate estimates, while testing several hundred brings the estimate much closer to its stable value. The exact number needed depends on the target's underlying bindability, so low-bindability targets require larger tested sets.
