AI-Guided Peptide Design for Pharmaceutical and Biotech Research
Drug Discovery: From known sequences to new candidates for experimental testing
Drug Discovery: From known sequences to new candidates for experimental testing
Peptides are short chains of amino acids that can be designed to bind specific biological targets, making them valuable across therapeutics, vaccines, and diagnostics. However, designing new peptides remains a largely experimental process.
Researchers typically start from a known binder, change one amino acid at a time, synthesize the resulting variants, and test whether binding is retained. The number of possible sequences quickly becomes far larger than any laboratory can test.
The challenge is that amino acids do not act independently. Whether a substitution is tolerated can depend on its position and on the residues surrounding it. Testing substitutions one at a time can therefore miss promising combinations while consuming experimental effort on low-value candidates.
Fraunhofer USA CMA, in collaboration with Fraunhofer IZI, developed an AI system trained on a large corpus of peptide sequences, including thousands of antibody-binding peptides and their variants. Across these sequences, the model learns recurring patterns that characterize families of related binders.
The model learns which positions are conserved, which tolerate variation, and which combinations of amino acids tend to occur together. Given a small set of known binders from a new peptide family, it applies these learned patterns to generate new variants and ranks them by how well each fits that specific family. The result is a prioritized shortlist for experimental testing.
The system works from sequence data alone. It requires no structural model and no receptor information, and accepts either binder/non-binder labels or measured binding affinities, making it applicable to datasets ranging from simple activity measurements to detailed experimental characterization.
We evaluated our approach on six peptide families from published studies, covering different receptor systems and family sizes ranging from 7 to approximately 100 known sequences. For validation, known sequences were withheld from the model and used as blind targets: the model had to recover them using only the remaining family members.
Across all six families, the model recovered more withheld known binders than the strongest conventional baseline. For a typical family, the model's top 10 candidates contained the same number of known binders as approximately 40 candidates from the baseline strategy, a fourfold reduction in candidates to consider.
The generated sequences also reproduce known biological design rules without being given those rules. The model preserved conserved motifs important for binding and concentrated substitutions at positions independently identified by experimental studies as tolerant to change. The resulting ranked lists provide a focused set of variants for synthesis and testing.
Our results show how AI can turn a relatively small set of existing peptide measurements into a prioritized set of testable designs. Rather than relying exclusively on sequential trial-and-error substitution, researchers can use these learned patterns to focus experimental effort on a smaller, more targeted set of candidates.