structural-biology/ai produced the result/Bioinformatics Advances 2026 · v2
Neural networks trained to find the cutting site on serpin proteins
Serpins are proteins that disable enzymes, and the short loop that does the work is hard to locate from sequence alone. Researchers trained neural networks on expert-labelled serpins to mark that loop residue by residue.
spectrum · one line per step, placed by what the step does · bright lines used AI
U-Net-based reactive center loop-identifier for serpins
Bioinformatics Advances, 2026
doi:10.1093/bioadv/vbag132 · record aix-00152 v2 · checked 2026-10-09
- AI was for
- Segmentation
- Model family
- Convolutional neural network, Recurrent neural network, Protein language model
- Checked by
- Held-out78 tested
- Code
- available
The finding the paper is about came from the AI.
What this research was about
Proteins are chains of amino acids, and serpins are a family of such chains that block enzymes which cut other proteins. The part that does the blocking is a short, exposed stretch called the reactive centre loop, or RCL. The enzyme bites the loop, and the serpin snaps shut around it. Knowing where that loop sits in a given serpin sequence matters for understanding how it works. The difficulty is that the loop's sequence varies greatly from one serpin to another. There is no shared pattern of letters to search for, so the usual trick of scoring a sequence against a matrix of expected amino acids at each position does not find it.
Labelling by hand is slow. Of the more than 48,000 serpins recorded in UniProt, a public catalogue of protein sequences, only 78 carry an RCL annotation. The researchers set out to see whether a model could learn to do the labelling instead. They gathered serpins from over 20 animal genomes, and trained experts marked the loop in each one by looking at three-dimensional structures, including high-confidence models from AlphaFold. That produced 1384 annotated loops, alongside 2048 non-serpin proteins of matching length as counter-examples.
Where AI came in
The task was framed as labelling every amino acid in a sequence as loop or not loop, a job much like marking out a region in an image. Each sequence was first turned into numbers three different ways: a plain code for each amino acid, a code based on the BLOSUM62 table of which amino acids substitute for which, and embeddings from ESM2, a language model trained on protein sequences and used here as released. On each of these the team trained three network types from scratch: a convolutional network, a U-Net, and a bidirectional LSTM, which reads the chain in both directions.
The trained models were then run on the 78 UniProt-annotated serpins, held back and never used in training, and their output compared with those existing annotations. The U-Net using ESM2 embeddings reached a residue-level F1 score of 0.9995 and accuracy of 0.9999, and both the U-Net and the LSTM got the whole loop exactly right in around 97% to 98% of sequences. The AI stood in for the structural inspection and hand-curation that had produced the training labels. Trimming the training set to 70% or 40% sequence similarity did not change accuracy much, and the BLOSUM-based U-Net sometimes flagged loop-like stretches in proteins that are not serpins.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record
The reactive center loop (RCL) of serpin proteins is highly variable and is annotated for only 78 of the more than 48,000 serpins in UniProt, so it cannot be located with position-matrix methods. The authors built an expert-annotated dataset of 1384 RCL regions plus 2048 length-matched non-serpins and trained CNN, U-Net and bidirectional LSTM models on one-hot, BLOSUM62 and ESM2 encodings to label each residue as RCL or non-RCL. On the 78 UniProt-annotated serpins used as an independent test set, the U-Net with ESM2 embeddings reached residue-level F1 = 0.9995, MCC = 0.9990 and accuracy = 0.9999, and per-sequence exact match was around 97%-98% for both U-Net and LSTM. Thresholding the training set to 70% or 40% sequence identity did not significantly change prediction accuracy, and applying the Blosum-Unet model to whole-genome protein sets sometimes flagged RCL-like regions in non-serpins.
How AI was used
Serpin sequences from over 20 animal genomes were downloaded from UniProt, and RCL regions were annotated by trained experts from publicly available 3D structures, including high-confidence AlphaFold models; serpins that already carried a UniProt RCL annotation were held back as an independent test set and length-matched non-serpin proteins were added as negatives. Sequences were truncated or zero-padded to 1024 residues and encoded three ways: one-hot, a BLOSUM62 substitution matrix encoding, and embeddings from the ESM2_650M protein language model used as released. For each encoding, three architectures were trained from scratch for per-residue two-class output: a four-block 1D CNN, a 1D U-Net with four DoubleConv encoder blocks, a bottleneck, a four-stage decoder with skip connections and attention gates, and a bidirectional LSTM. Training used Adam at learning rate 0.001 for up to 50 epochs with early stopping on validation F1, a masked binary cross-entropy loss that excluded padded positions, and batch sizes of 32 for one-hot or BLOSUM and 4 for ESM2, on one or two Nvidia B200 GPUs. The trained models were then run over the held-out sequences and their predicted RCL positions compared with the existing UniProt annotations at residue and sequence level; the training set was additionally thresholded at 70% and 40% sequence identity and the models retested.
The shape of the work
Structural · the record, drawn
no AI
Collect serpin and non-serpin sequences
Obtaining raw data, whether by measurement, download or retrieval.
All serpin proteins encoded by these genomes were batch-downloaded from UniProt.where the paper describes this · verbatim
no AI
Expert annotation of RCL regions
Cleaning, filtering, normalising or labelling data already obtained.
we meticulously annotated the RCL regions using publicly available 3D structural informationwhere the paper describes this · verbatim
AI
Encode sequences
Encoding data into features, descriptors, embeddings or graphs. The AI stood in for conventional algorithm.
Three embedding methods were implemented: one-hot encoding, BLOSUM62 substitution matrix-based encoding, or the ESM2 encoding with the ESM2_650M model.where the paper describes this · verbatim
AI
Train per-residue classifiers
Fitting model parameters, including fine-tuning an existing model. The AI stood in for manual curation.
All models were trained for up to 50 epochs using the Adam optimizerwhere the paper describes this · verbatim
AI
Predict RCL positions on held-out sequences
Running a trained model over new data to predict, classify or score. The AI stood in for manual curation.
Training and inference were conducted on the University of Florida HiPerGator supercomputerwhere the paper describes this · verbatim
no AI
Evaluate against UniProt annotations
Testing outputs against ground truth.
At the per-sequence level, the exact match rate is around 97%–98% for both U-Net and LSTMwhere the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
The paper's result is the trained RCL identifier itself and its per-residue/per-sequence annotation performance; there is no non-AI route to the reported annotations
We evaluated three neural network architectures—CNN, U-Net, and LSTM—to compare performance.where the paper describes this · verbatim
On the independent test dataset, U-Net-based models achieved ∼98% accuracy in identifying the RCL at the per-sequence level.where the paper describes this · verbatim
Source code, training and testing datasets, and models are freely available at https://github.com/leizhou69/RCL-identifierwhere the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- Version of 1D U-Net (onehot-Unet, Blosum-Unet, ESM2-Unet)Which version of the model was used is not stated.
- Version of 1D CNNWhich version of the model was used is not stated.
- Version of Bidirectional LSTMWhich version of the model was used is not stated.
About this article
Record aix-00152, version 2, checked by a person on 2026-10-09. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error