structural-biology/ai produced the result/Briefings in Bioinformatics 2024 · v2
Machine learning scores short protein stretches that mark proteins for destruction
Researchers built MetaDegron, two trained models that read short protein sequences and score whether they are degrons, the tags that mark a protein for disposal. One model learns from computed structural features, the other from sequence alone.
spectrum · one line per step, placed by what the step does · bright lines used AI
MetaDegron: multimodal feature-integrated protein language model for predicting E3 ligase targeted degrons
Briefings in Bioinformatics, 2024
doi:10.1093/bib/bbae519 · record aix-00184 v2 · checked 2026-10-09
- AI was for
- Classification
- Model family
- Gradient-boosted trees, Protein language model, Recurrent neural network, Convolutional neural network, Multilayer perceptron
- Checked by
- Held-out
- Code
- available
The finding the paper is about came from the AI.
What this research was about
Cells do not keep every protein they make. A protein that is damaged, or simply no longer needed, is labelled for disposal by enzymes called E3 ligases, which attach a small marker to it so the cell's recycling machinery takes it apart. The E3 ligase does not grab anywhere on its target. It recognises a short stretch of the protein chain, a few amino acids long, known as a degron. Finding these stretches is hard. They are short, they vary a great deal from one protein to another, and whether a given stretch works as a degron depends partly on how it sits in the folded protein, not just on which letters it spells.
The researchers set out to build a predictor: a program that takes a protein sequence and scores which segments look like degrons for particular E3 ligases. They gathered known degron examples from a motif database and from earlier studies, and paired them with randomly chosen peptides of the same length to serve as a comparison set.
Where AI came in
The AI is the predictor itself, and every result reported is its output. One version, MetaDegron-X, is a set of ten gradient-boosted decision tree classifiers, a kind of model that learns by combining many simple yes-or-no rules. These were trained on ten quantities computed for each segment with existing bioinformatics tools, covering things like how flexible the stretch is, how exposed it is to water, and how well conserved it is across related species. Their scores are averaged into a single degron probability.
The second version, MetaDegron-D, works from sequence alone. It uses SeqVec, an already trained protein language model, which has learned patterns of amino acid usage across a large database of sequences in much the way a text language model learns patterns of words. SeqVec turns each amino acid into a list of numbers that reflects its context, and a network of recurrent and convolutional layers learns from those numbers, alongside a separately trained branch that reads the amino acids without context. Both models were tested by cross-validation, on a held-out split, and against an earlier predictor on substrates of one E3 ligase. They are served from a web page that scores submitted sequences for 21 E3 ligases, standing in place of laboratory testing of each candidate.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record
The authors built MetaDegron, a pair of models that score short protein sequence segments as E3 ligase targeted degrons or as random peptides, trained on more than 300 curated degron instances and a length-matched random peptide background. MetaDegron-X is an ensemble of ten XGBoost classifiers over ten computed structural, evolutionary and physicochemical features, and MetaDegron-D is a deep network that works from sequence alone using embeddings from the SeqVec protein language model together with a jointly trained amino acid embedding branch. In five-fold cross-validation the average AUC was 0.87 for MetaDegron-X and 0.90 for MetaDegron-D, with independent-test AUCs of 0.86 and 0.90; on an additional dataset of experimentally validated β-TrCP2 substrates the reported AUCs were 0.9705 for MetaDegron-D, 0.9670 for MetaDegron-X and 0.9540 for Degpred. The models are deployed as a web server that predicts degrons for 21 E3 ligases in batch and annotates degron-related features.
How AI was used
Degron motifs and instances were collected from the ELM database and prior studies, and a background set of length-matched random peptides was assembled. Ten features per instance — flexibility, solvent accessibility, secondary structure, disorder, anchoring score, conservation, domain and PTM context — were computed with existing structural bioinformatics tools, and the dataset was split 90/10 into training and independent sets. On the feature representation, ten XGBoost classifiers were trained under a bootstrapping strategy with parameters selected by cross-validation, and their scores averaged to give a degron probability (MetaDegron-X). A second, sequence-only model (MetaDegron-D) combined 1024-dimensional per-residue embeddings from the pre-trained SeqVec protein language model, passed through BLSTM and convolution-pooling layers, with a separately trained context-insensitive amino acid embedding feeding its own BLSTM branch; the two branches were joined by dense layers ending in a two-node output for the degron and random peptide classes. Evaluation used five-fold cross-validation, the independent split, a feature-elimination study, and a separate dataset of β-TrCP2 substrates with repeated random background sampling against the published Degpred predictor. Both trained models are served from a web application that accepts FASTA input and selected E3 ligases and returns scored degron instances with feature and structure annotations.
The shape of the work
Structural · the record, drawn
no AI
Curate degron instances and background peptides
Obtaining raw data, whether by measurement, download or retrieval.
We collected and processed a set of human degron motifs, which are E3 binding consensus patterns, from the ELM databasewhere the paper describes this · verbatim
no AI
Compute structural, evolutionary and physicochemical features
Encoding data into features, descriptors, embeddings or graphs.
We calculated 10 features for all motif instances or random peptides.where the paper describes this · verbatim
no AI
Partition into training and independent datasets
Cleaning, filtering, normalising or labelling data already obtained.
the constructed dataset was partitioned into a training dataset, which represented 90% of the total datawhere the paper describes this · verbatim
AI
Train XGBoost ensemble on degron features (MetaDegron-X)
Fitting model parameters, including fine-tuning an existing model.
the XGBoost classifier (called MetaDegron-X) was constructed using these discerning features for E3 targeted degronwhere the paper describes this · verbatim
AI
Train hybrid sequence-only deep network (MetaDegron-D)
Fitting model parameters, including fine-tuning an existing model.
Here, we used SeqVec, a protein-specific adaptation of ELMowhere the paper describes this · verbatim
AI
Evaluate by cross-validation and independent testing
Testing outputs against ground truth.
we employed five-fold CV, a widely used technique in machine learning, to assess the model’s performance on the training datasetwhere the paper describes this · verbatim
AI
Compare with Degpred on β-TrCP2 substrate dataset
Testing outputs against ground truth.
we compared MetaDegron to Degpred and found that MetaDegron could achieve better AUROC valuewhere the paper describes this · verbatim
AI
Serve batch degron prediction and annotation via web server
Running a trained model over new data to predict, classify or score. The AI stood in for physical experiment.
MetaDegron possesses the capability to predict targeted degrons of 21 E3 ligases in a batch mannerwhere the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
The paper's object is the predictor itself; every reported result is the output of the trained XGBoost ensemble and the hybrid deep network.
we employed a bootstrapping strategy to train 10 eXtreme Gradient Boosting (XGBoost) classifierswhere the paper describes this · verbatim
compared the performance of the MetaDegron with previous method using an additional dataset comprising experimentally validated substrates of β-TrCP2where the paper describes this · verbatim
MetaDegron can be accessed freely at http://modinfor.com/MetaDegron/ and https://github.com/BioDataStudy/MetaDegronwhere the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- Trained model weightsWhether the trained model is available is not stated.
- ComputeThe hardware or time used is not stated.
- How many were testedThe paper gives no count of what was tested.
- Version of MetaDegron-X (XGBoost ensemble of 10 classifiers)Which version of the model was used is not stated.
- Version of MetaDegron-D (hybrid BLSTM/CNN deep neural network)Which version of the model was used is not stated.
- Version of SeqVec (ELMo-based protein language model, pre-trained on UniRef50)Which version of the model was used is not stated.
- Version of DegpredWhich version of the model was used is not stated.
- What step 4 replacedThe paper gives no basis for what the AI stood in for.
- What step 5 replacedThe paper gives no basis for what the AI stood in for.
- What step 6 replacedThe paper gives no basis for what the AI stood in for.
- What step 7 replacedThe paper gives no basis for what the AI stood in for.
About this article
Record aix-00184, version 2, checked by a person on 2026-10-09. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error