structural-biology/ai produced the result/Bioinformatics 2022 · v2
Neural networks pick amino acid sequences for short peptides in protein binding sites
Researchers trained two neural networks, PepSeP1 and PepSeP6, to choose the amino acids of six-residue peptide fragments lodged in a protein's binding site. The networks produced the sequences; Rosetta software then refined and scored them.
spectrum · one line per step, placed by what the step does · bright lines used AI
Deep learning of protein sequence design of protein–protein interactions
Bioinformatics, 2022
doi:10.1093/bioinformatics/btac733 · record aix-00153 v2 · checked 2026-10-09
- AI was for
- Candidate generation
- Model family
- Convolutional neural network, Recurrent neural network
- Checked by
- Held-out1245 tested
- Code
- available
The finding the paper is about came from the AI.
What this research was about
Proteins do much of their work by sticking to one another. A short stretch of one protein settles into a groove on the surface of another, and the fit depends on which amino acids — the twenty chemical building blocks of proteins — sit at each position along that stretch. Designing such a binder means answering a hard question: given the shape of the groove and the path the peptide's backbone takes through it, which amino acids should fill the positions? The number of possible combinations is enormous, and the usual approach is to search through them with physics-based scoring software, which is slow and does not always land on sequences resembling those nature uses.
The authors set out to have a neural network make that choice directly. They assembled peptide–binding site pairs from 9002 co-crystal structures in the Protein Data Bank, a public archive of experimentally determined protein shapes. In each pair, the peptide was a six-residue fragment and the binding site a patch on the partner protein. They stripped the peptides back to plain backbones, nudged them away from their natural shapes, and asked whether a model could recover the original amino acids from geometry alone.
Where AI came in
The AI is where the sequence comes from. Each complex was turned into maps of the distances between backbone atoms, within the peptide and between peptide and binding site, plus a description of the binding site's amino acids and shape. A convolutional network — the kind used for reading images — compressed that into a summary, and a second network with an attention mechanism read the summary off one residue at a time, much as image-captioning systems write a sentence describing a picture. PepSeP1 gives one sequence per complex; PepSeP6 reuses PepSeP1's compressed summary and emits six, each conditioned on the one before.
The networks stood in for the physics-based search that normally picks the amino acids. The authors ran Rosetta's FastDesign on the same bare backbones as a comparison, and also used PepSeP1's output to guide Rosetta redesigns. On a held-out test set of 1245 cases, PepSeP1 matched the natural amino acid at 40.83% of positions and at 45.95% of positions flagged as binding hot-spots; PepSeP6 averaged 38.67% across its outputs. Matching was lower on interfaces between different proteins, at 27.71%, and lower again on the antibody–antigen sets, at 26.00% and 16.78%. Code and example data are published.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record
The authors trained two attention-based encoder–decoder networks, PepSeP1 and PepSeP6, to assign amino acid sequences to 6-residue peptide fragments sitting in a protein binding site, using peptide–binding site complexes extracted from 9002 co-crystal structures in the Protein Data Bank. On the independent test set, PepSeP1 recovered 40.83% of native peptide residues overall and 45.95% of residues identified as binding hot-spots, while PepSeP6, which emits six sequences per complex, averaged 38.67% across its outputs. Recovery was lower on hetero-oligomeric interfaces (27.71%) and on the antibody–antigen benchmark subsets (26.00% and 16.78%). Compared with Rosetta's FastDesign applied to the same all-glycine peptides, PepSeP1 recovered more native residues overall, though FastDesign gave higher hot-spot recovery on the T-he and B-ag/ab subsets.
How AI was used
Peptide–binding site complexes were extracted from multichain PDB co-crystal structures, with the peptide defined as a 6-residue fragment of one partner and the binding site as a 24–48 residue patch of the other; peptides were mutated to all-glycine and perturbed from their native conformation, and the complexes were split into training, validation, test and a held-out antibody–antigen benchmark set. Each complex was encoded as intramolecular and intermolecular backbone distance maps plus binding-site secondary structure, binding-site amino acid types and a homo- or hetero-oligomeric flag. These features were passed to a convolutional encoder of two blocks (8 and 4 layers) whose concatenated feature vectors fed a bidirectional Bahdanau-attention LSTM decoder, trained with categorical cross-entropy in three learning-rate stages and repeated 20 times, with the run selected on recovery over hetero-oligomeric subsets. PepSeP6 reused the PepSeP1 encoder with frozen weights and ran its decoder five times, each conditioned on the previous prediction's final hidden state, with the PepSeP1 output as the sixth sequence. Designed sequences were threaded back onto the perturbed backbones, refined with Rosetta FastRelax under harmonic constraints, and scored with InterfaceAnalyzerMover under ref15; Rosetta FastDesign was run by the authors as a baseline both on all-glycine backbones directly and as a redesign constrained by PepSeP1 position-specific scoring matrices in three schemes (RD3, RD5, RD20). Recovery was measured against native sequences overall and at alanine-scanning hot-spot positions, and the models were additionally applied to highly perturbed backbones generated by the iNNterfaceDesign method in antibody case studies.
The shape of the work
Structural · the record, drawn
no AI
Extract peptide–binding site complexes from crystal structures
Obtaining raw data, whether by measurement, download or retrieval.
The complexes were extracted from 9002 co-crystal structures.where the paper describes this · verbatim
no AI
Perturb backbones, glycine-mutate and split datasets
Cleaning, filtering, normalising or labelling data already obtained.
Peptides were perturbed up to 1.07 Å root-mean-square deviation (RMSD) of their native conformationwhere the paper describes this · verbatim
no AI
Encode complexes as distance maps and sequence features
Encoding data into features, descriptors, embeddings or graphs.
We use two types of distance maps as the main geometrical descriptors of the structureswhere the paper describes this · verbatim
AI
Train PepSeP1 and PepSeP6 networks
Fitting model parameters, including fine-tuning an existing model. The AI stood in for conventional algorithm.
Training of the PepSeP1 model was conducted 20 times in three stages: 5 epochs with a learning rate of 0.001where the paper describes this · verbatim
AI
Design peptide sequences for held-out and perturbed backbones
Producing candidate objects that did not previously exist. The AI stood in for conventional algorithm.
Five outputs are generated by passing feature vectors produced by the encoder into the decoder of PepSeP6 five timeswhere the paper describes this · verbatim
no AI
Rosetta FastDesign redesign and baseline design
Iterative search over a space.
Peptide ligands designed by the PepSeP1 method underwent an additional design step after refinement using the FastDesign protocolwhere the paper describes this · verbatim
no AI
Refine complexes and compute binding energies
Numerical or physics simulation, including where a learned surrogate replaces it.
Binding free energies of the complexes were estimated using InterfaceAnalyzerMover with repacking chains after separation.where the paper describes this · verbatim
no AI
Score sequence and hot-spot recovery against native sequences
Testing outputs against ground truth.
To evaluate the performance of PepSeP1, we utilized sequence recovery.where the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
The paper's result is the designed peptide sequences themselves, which are produced by the trained networks PepSeP1 and PepSeP6; the reported recovery rates and binding energies are measurements of those model outputs.
We developed an attention-based deep learning model inspired by algorithms used for image-caption assignments to design peptideswhere the paper describes this · verbatim
The native sequence recovery rate of PepSeP1 is 40.83% on the independent test setwhere the paper describes this · verbatim
All the code and example data are available at https://github.com/strauchlab/iNNterfaceDesignwhere the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- Trained model weightsWhether the trained model is available is not stated.
- ComputeThe hardware or time used is not stated.
- Version of PepSeP1Which version of the model was used is not stated.
- Version of PepSeP6Which version of the model was used is not stated.
About this article
Record aix-00153, version 2, checked by a person on 2026-10-09. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error