← Back to blog

Researchers, Start Cyclic Peptide Permeability Prediction With DMPNN

October 10, 2026
Researchers, Start Cyclic Peptide Permeability Prediction With DMPNN

For cyclic and cell-penetrating peptides, the current best practice is to train multimodal, graph-based models on curated, assay-aligned datasets such as CycPeptMPDB using a regression formulation rather than binary classification. This database aggregates thousands of membrane permeability measurements across multiple assays, giving models enough chemical diversity to generalize. Your fastest path to a working baseline: pull the curated splits, build a DMPNN-style regression model, and benchmark against assay-specific subsets before adding 3D or contrastive features.


TL;DR:

  • PAMPA supplies most CycPeptMPDB labels, so pooled models may reflect passive diffusion better than active transport or efflux.
  • Cell based intestinal assays capture active transport and efflux, while artificial membrane assays measure passive diffusion; compare results within each assay.
  • Published CPMP and graph models report R² values from 0.48 to 0.75 across assays, but results vary by assay and peptide class.
  • Use both random and scaffold splits, since random splits can overstate performance when related peptide scaffolds appear in training and test sets.
  • For peptides above 500 daltons, add a small conformer ensemble because three dimensional shape can capture chameleonic behavior that two dimensional graphs miss.

Innovabiotech
Build Better Peptide Permeability Models
Innova Biotech provides tailored bioinformatics and computational biology solutions for peptide design and other project-specific research needs.
Explore Innova Biotech

Table of Contents

Choosing a Computational Approach: MD, QSPR, ML, and Multimodal Models

Picking a method family comes down to what you need to know and how much compute you can spend. Molecular dynamics (MD) simulations capture the physical mechanism: how a peptide's conformation shifts as it crosses a lipid bilayer, which intramolecular hydrogen bonds form, and how solvent-accessible surface area changes along the way. That mechanistic detail costs hours to days of compute per candidate, which rules MD out for screening large libraries but makes it valuable when you need to understand why a specific peptide permeates or fails.

Classical QSPR models, built on physicochemical descriptors, sit at the opposite end: fast, interpretable, and useful as a sanity check, but they tend to miss nonlinear structural effects that matter for cyclic peptides. Deep learning models, particularly graph neural networks, learn representations directly from molecular structure and have become the workhorse for high-throughput prediction because they scale to thousands of compounds without per-candidate simulation cost.

Multimodal frameworks represent the current frontier: they fuse SMILES strings, 2D molecular graphs, and 3D conformer information into a single model, often using contrastive learning to align the different views of the same molecule. Research in this direction, including a multi-modal contrastive learning framework for cyclic peptide permeability, shows that combining complementary representations improves generalization beyond what any single modality achieves alone.

  • MD simulation: best for mechanistic insight on a handful of candidates, especially chameleonic peptides.
  • Classical QSPR: fast and interpretable, useful as a baseline or sanity check.
  • Graph neural networks: the default choice for screening large peptide libraries.
  • Multimodal/contrastive models: the strongest option when you have the data and compute to train them.

Which Assays and Datasets Train Today's Permeability Models

Every permeability model inherits the biases of its training labels, so understanding assay provenance matters as much as model architecture. PAMPA (parallel artificial membrane permeability assay) measures passive diffusion across a lipid-coated artificial membrane, with no transporters or efflux pumps involved, making it fast and cheap but mechanistically limited. Caco-2 uses a human colon-derived cell monolayer that includes active transport and efflux, giving a readout closer to intestinal absorption but at higher assay cost and variability. RRCK and MDCK are both cell-line based assays similar in spirit to Caco-2, often used to cross-validate permeability rankings when Caco-2 data is scarce for a given peptide class.

CycPeptMPDB is the central public resource tying these assays together for cyclic peptides. The database comprises several thousand cyclic peptides collected from dozens of publications and patents, with a majority of PAMPA measurements and smaller numbers from Caco-2, MDCK, and RRCK assays. The PAMPA-heavy composition of CycPeptMPDB means most publicly trained models are implicitly biased toward passive diffusion behavior, which can understate the contribution of active transport or efflux for peptides destined for oral absorption assessment.

  • PAMPA: fast, cheap, passive diffusion only.
  • Caco-2: includes active transport and efflux, closer to intestinal physiology.
  • RRCK and MDCK: cell-line assays used to cross-check permeability rankings.
  • CycPeptMPDB: the largest public aggregator spanning all four assay types.

Because assay choice changes what "permeability" actually measures, the safest practice is to keep assay identity as a model feature or to train and evaluate assay-specific splits separately, rather than pooling everything into one undifferentiated label.

Representing Peptide Structure: SMILES, Graphs, Conformers, and Descriptors

The representation you feed a model determines what it can possibly learn. A 1D SMILES string is compact and easy to generate but discards three-dimensional information that matters for cyclic peptides, whose permeability often depends on which conformation they adopt in a given environment. A 2D molecular graph preserves connectivity and supports message-passing architectures, giving models a sense of local chemical neighborhoods without the cost of conformer generation. A 3D conformer ensemble goes further, capturing the shape-shifting behavior known as chameleonicity, where a peptide hides its polar groups inside intramolecular hydrogen bonds when crossing a hydrophobic membrane and exposes them again in water. Mechanistic studies of chameleonic behavior show this conformational switching is a key determinant of passive permeation for many cyclic peptides, which is why static 2D graphs alone can miss an important signal.

  1. Generate a canonical SMILES string and verify it round-trips to the same structure.
  2. Build a 2D molecular graph for message-passing or DMPNN-style featurization.
  3. Generate a small 3D conformer ensemble, especially for peptides larger than 500 daltons.
  4. Compute physicochemical descriptors: logD, topological polar surface area (TPSA), molecular weight, rotatable bond count, and hydrogen bond donor and acceptor counts.
  5. Store all representations alongside the assay label, so you can test which combination performs best for your dataset.

Tools like RDKit for graph and descriptor generation, combined with conformer samplers, cover most of this pipeline without custom code. Caco-2 transport data shows transepithelial transport correlates strongly and non-linearly with logD (positively) and molecular weight (negatively), confirming these two descriptors alone carry real predictive signal even before structural features are added.

Pro Tip: Compute logD, TPSA, and molecular weight first as a quick baseline. If a simple descriptor model already separates permeable from non-permeable peptides reasonably well, you know your harder graph or 3D model has a meaningful bar to beat.

Model Architectures That Lead Current Permeability Benchmarks

Graph-based models, particularly Directed Message Passing Neural Networks (DMPNNs), consistently rank at the top of permeability benchmarks because they capture both local chemical environments and the longer-range structural dependencies that matter for macrocycles. Sequence-based models borrowed from natural language processing can work on SMILES strings but tend to underperform graph encodings because they have to relearn connectivity patterns that a graph already encodes explicitly. Convolutional architectures applied to 2D or 3D grids see some use but have mostly been superseded by graph approaches in recent cyclic peptide work.

Directed messages passing through a peptide graph

Recent models that fuse SMILES, 2D graphs, and 3D geometric data, including the CPMP framework and related GNN-based approaches, report R2 values typically between 0.48 and 0.75 across PAMPA, Caco-2, RRCK, and MDCK datasets, according to benchmark results. Reported values for CPMP and GNN models typically fall between moderate and good ranges of R2 across assays, indicating variability depending on assay and peptide class. That spread shows performance depends heavily on which assay and which peptide class you're evaluating on, not just which model you choose.

Multimodal fusion tends to outperform any single-representation model because it lets the network draw on complementary signal: a 2D graph for connectivity, a 3D conformer for shape, and physicochemical descriptors for bulk properties. Multi-task training, where a model predicts permeability alongside auxiliary properties like logP or TPSA, often improves generalization slightly by forcing the shared layers to learn representations useful across related targets rather than overfitting to one narrow label.

  • DMPNN and other graph architectures lead most published benchmarks.
  • Multimodal fusion of SMILES, graphs, and 3D geometry improves generalization over single-representation models.
  • Multi-task prediction of auxiliary properties (logP, TPSA) is a low-cost addition that often helps stability.
  • Regression formulations tend to outperform classification framings for permeability, since they preserve the magnitude information that a binary cutoff throws away.

A Step-by-Step Workflow for Building and Validating Permeability Models

A reproducible permeability model starts with data curation, not architecture choice. Here's the order that avoids the most common failure modes:

  1. Unify molecular formats. Convert all peptides to a consistent representation, whether HELM notation for monomer-level detail or canonical SMILES, and verify round-trip consistency.
  2. Remove duplicates and near-duplicates. Cyclic peptides with different stereochemistry annotations can look like duplicates to naive string matching, so deduplicate on canonical structure, not raw text.
  3. Annotate assay provenance. Tag every label with its source assay (PAMPA, Caco-2, RRCK, MDCK) so you can train assay-aware models or evaluate assay-specific subsets later.
  4. Normalize units. Permeability values across papers appear in different unit conventions; standardize before pooling.
  5. Generate representations. Build SMILES, graphs, and conformer ensembles as described earlier, caching them so retraining doesn't repeat expensive generation steps.
  6. Split the data twice. Run both a random split and a scaffold split, since scaffold splitting simulates how well a model generalizes to genuinely novel chemotypes, while random splitting can overstate performance when similar scaffolds appear in both train and test sets.
  7. Train with a fixed featurization pipeline. Lock down preprocessing before hyperparameter search, so performance differences reflect architecture choices, not pipeline drift.
  8. Report uncertainty. Run multiple training seeds and report mean and standard deviation, not a single best run.
  • Keep a changelog of every dataset version and preprocessing decision.
  • Store model checkpoints and exact data splits alongside any published result.
  • Report R2 and RMSE for regression, and ROC-AUC alongside a confusion matrix if you frame the task as classification.

Pro Tip: Always report both random-split and scaffold-split performance side by side. A model that looks strong only under random splitting is telling you more about scaffold overlap in your dataset than about real predictive power.

For teams building out this pipeline for the first time, a practical guide to peptide binding affinity prediction covers similar validation choices that transfer directly to permeability modeling.

Combining Molecular Dynamics With Machine Learning Efficiently

Running full production MD on every candidate in a screening library isn't realistic, but MD-derived information still adds real value when folded into an ML pipeline as a feature source rather than a per-candidate requirement. The goal is to extract mechanistic signal once, on a representative subset, and let the model generalize that signal to new peptides without resimulating each one.

  • Intramolecular hydrogen bond statistics across a conformer ensemble, which capture chameleonic behavior directly.
  • Solvent-accessible surface area (SASA) distributions, showing how much polar surface a peptide exposes in different environments.
  • Conformer relative energies, flagging which structures are energetically accessible at physiological temperature.
  • Precomputed conformer libraries for common peptide scaffolds, reused across many candidates that share a macrocyclic core.

Replica exchange molecular dynamics (REMD) and other enhanced sampling methods can generate a representative conformer ensemble faster than brute-force simulation, and running them once per scaffold family rather than once per candidate keeps compute costs manageable. The practical shortcut most teams land on: run enhanced sampling on a diverse representative set of peptides, extract conformer-level features, and train an ML model to predict those same features from 2D structure alone for the remaining candidates. Discussion of peptide cyclization strategies is a useful companion reference here, since the cyclization chemistry you choose directly shapes which conformers are even accessible.

Reporting Standards That Make Benchmark Comparisons Fair

Permeability benchmarks are only useful if they're comparable, and right now a lot of published results aren't, because teams use different splits, different assay mixes, and different metrics. A few reporting habits help fix most of that.

Train and evaluate on large curated datasets like CycPeptMPDB rather than small private sets, and report performance separately for each assay rather than pooling PAMPA, Caco-2, RRCK, and MDCK into one number. Report R2 and RMSE for regression tasks, and add ROC-AUC if you also frame the problem as classification, since a model can look strong on one metric and mediocre on another. Run each experiment across multiple random seeds and report the spread, not just a single best run, since a single number hides how sensitive a model is to initialization or data ordering.

  • Report assay-specific performance, not just an aggregate score.
  • Include both random-split and scaffold-split results in every benchmark table.
  • Share code, data splits, and random seeds so other teams can reproduce your exact setup.
  • Publish model checkpoints where licensing allows, so comparisons don't require full retraining.

Pro Tip: Before trusting a benchmark number, check whether it's assay-pooled or assay-specific. A model that reports one R2 across all four assays is usually hiding a much wider performance range underneath.

A 2026 guide to bioinformatics validation for peptides walks through these reporting standards in more depth, including how to present uncertainty in a way reviewers expect.

How We Support Permeability Prediction Projects at Innova Biotech

We build permeability prediction pipelines as part of broader peptide design and virtual screening engagements, which means your permeability model doesn't exist in isolation from the rest of your discovery workflow. Our team curates assay-aligned training data, builds multimodal graph-based models, and runs the MD sampling needed to extract chameleonic conformer features when a project calls for mechanistic grounding rather than just a score.

  • Data curation and assay-provenance tagging for permeability datasets.
  • Multimodal model training combining SMILES, graph, and 3D conformer features.
  • MD-based feature extraction for chameleonic and cell-penetrating peptide candidates.
  • Validation reporting with uncertainty estimates and reproducible splits.

Typical engagements deliver a curated dataset, a trained and validated model, and an interpretability report explaining which structural features influenced each prediction. Clients can reach out to discuss specific projects, with team contacts available to address permeability-focused work.

Where Peptide Permeability Research Needs to Go Next

The biggest limitation in this field right now isn't model architecture, it's data. CycPeptMPDB is a genuine advance, but it's still small relative to the chemical diversity of cyclic and cell-penetrating peptides worth testing, and PAMPA measurements dominate it in a way that understates active transport effects. The next real gains will come from larger, more assay-balanced datasets and from imaging-based mechanistic validation that confirms what MD simulations predict about chameleonic conformational switching, rather than from yet another architecture tweak on the same training set.

My priority list for the community: publish raw data splits alongside every benchmark paper, standardize assay-provenance tagging so pooled datasets stop hiding their biases, and invest in enhanced-sampling infrastructure that's cheap enough to run on hundreds of candidates rather than a handful. Reproducibility is the bottleneck now, not creativity.

— Hooman

Partner With Us on Your Next Permeability Prediction Project

We take the workflow above and run it end to end: curated assay-aligned data, multimodal graph-based models, and MD-informed features built around your actual candidate library rather than a generic benchmark set.

Innovabiotech

  • Peptide design for cyclic and cell-penetrating peptide candidates.
  • Virtual screening and hit-to-lead pipelines that fold permeability prediction into candidate prioritization.
  • Protein engineering support for chimeric constructs that need permeability alongside stability data.

If your team needs a permeability model built around your own data rather than a public benchmark, reach out through our peptide design page to scope the project.

FAQ

How do I estimate the pI of a peptide?

The isoelectric point (pI) is typically estimated from the pKa values of ionizable side chains and terminal groups using the Henderson-Hasselbalch relationship, summed across all ionizable residues. Most researchers use established calculators such as those built into ExPASy or peptide design software rather than computing it by hand, since cyclic and modified peptides often have non-standard termini that need manual adjustment.

Can AlphaFold predict peptide structure?

AlphaFold was trained primarily on globular proteins and performs less reliably on short, cyclic, or heavily modified peptides, where conformational flexibility and non-standard linkages fall outside its training distribution. For cyclic peptide permeability work, conformer ensembles from molecular dynamics or specialized cyclic-peptide modeling tools generally give more reliable structural input than AlphaFold alone.

How long does it take for your body to absorb peptides?

Absorption timing depends heavily on administration route, peptide stability, and permeability, so there's no single figure that applies across peptides. Orally administered peptides face enzymatic degradation and low passive permeability, which is why benchmark permeability work often uses Cyclosporin A as a reference compound given its established oral bioavailability profile.

Is 98% purity good for peptides?

Some structural or clinical-adjacent work calls for even higher purity, so the right threshold depends on the sensitivity of the downstream assay.

Sources