~/aixsci
200 records · all checked

structural-biology/ai produced the result/Scientific Reports 2022 · v2

Protein language model embeddings used to predict succinylation sites in proteins

Researchers built LMSuccSite, a tool that predicts which lysines in a protein carry a succinyl tag, using only the protein's sequence. Numerical descriptions from a pre-trained protein language model replaced hand-designed sequence features.

1. Reuse curated succinylation site dataset2. Build window peptides and balance negatives3. Extract ProtT5 embeddings for lysine residues4. Train base modules and alternative ML/DL models5. Select architectures and hyperparameters by cross-validated grid search6. Train stacked meta-classifier (LMSuccSite)7. Score independent test set and compare with existing predictors8. Inspect learned features and training-size sensitivity

spectrum · one line per step, placed by what the step does · bright lines used AI

Improving protein succinylation sites prediction using embeddings from protein language model
Scientific Reports, 2022

doi:10.1038/s41598-022-21366-2 · record aix-00181 v2 · checked 2026-10-09

ai-resultrole of AI
AI was for
Classification
Model family
Protein language model, Transformer, Convolutional neural network, Multilayer perceptron, Random forest, Support vector machine, Gradient-boosted trees, Recurrent neural network
Checked by
Benchmark
Code
available

The finding the paper is about came from the AI.

read as

The science is explained before the AI appears. Switch to field specialist to go straight to the method.

Assumes the discipline and goes straight to the method.

The diagram, the record and what the paper did not report are identical in both modes. Only the framing changes — never the evidence.

Introduction by AIxSci · plain language

What this research was about

Proteins are chains of amino acids, and after a cell makes one it often decorates it with small chemical tags. Succinylation is one such tag, attached to the amino acid lysine, and it can change how a protein behaves. Finding which lysines carry the tag matters for understanding how proteins are regulated, but the usual route is laboratory work with mass spectrometry, which is slow and only ever catches part of the picture. That leaves many proteins with lysines whose status is simply unknown.

The alternative is to predict the answer from the sequence itself. The difficulty is that a lysine looks much like any other lysine on paper, and the clues lie in the surrounding stretch of amino acids. Traditionally researchers decided by hand which properties of that neighbourhood to measure. Here the researchers instead set out to let a model work out its own numerical description of each site, and to combine two such descriptions into a single predictor they call LMSuccSite.

Where AI came in

AI supplied the description of each candidate site and the judgement about it. A protein language model, ProtT5-XL-UniRef50, had already been trained on large numbers of protein sequences in the way a text model learns language, by predicting hidden pieces of its input. Feeding a whole protein through it returns a 1024-number vector for each amino acid, and the researchers kept the vector for the lysine in question. Separately, a short 33-residue window around each lysine was fed through a convolutional network that learned its own encoding. Between them these stood in for hand-crafted sequence features.

The two parts were then trained to classify sites as succinylated or not, and their internal outputs were joined and passed to a further neural network that made the final call. Random forests, support vector machines, gradient-boosted trees and other networks were trained on the same features for comparison, with architectures and settings chosen by cross-validation and grid search. The reported scores on a held-aside test set are the model's own predictions. The researchers also projected the learned features into two dimensions with t-SNE and retrained on smaller slices of the data.

Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.

The work

Technical · from the record

The study builds LMSuccSite, a sequence-only predictor of lysine succinylation sites in proteins. It combines a supervised word embedding of a 33-residue window around each lysine, processed by a 2D convolutional network, with per-residue embeddings from the pre-trained protein language model ProtT5-XL-UniRef50, processed by a two-layer feed-forward network; features from the second-to-last layer of each module are concatenated and passed to a neural-network meta-classifier. On an independent test set the authors report MCC, sensitivity and specificity of 0.36, 0.79 and 0.79, and the highest MCC, sensitivity and g-mean among the predictors compared, while other methods scored higher on accuracy and specificity. t-SNE plots of the learned features and a training-size sensitivity analysis are also reported.

How AI was used

A previously published succinylation dataset was reused: sequences were redundancy-filtered, 33-residue windows were taken around each lysine, all non-annotated lysines in the same proteins were treated as negatives, and the negative training set was randomly under-sampled to match the positives. Two encodings were computed from sequence alone: a supervised word embedding learned in a Keras embedding layer over the window peptides (vocabulary size 21, output shape 33 x 21), and 1024-dimensional contextualised embeddings taken from the encoder of the pre-trained ProtT5-XL-UniRef50 language model applied to full-length sequences, with the vector for the central lysine retained. A 2D CNN was trained on the supervised embedding and an ANN with hidden layers of 256 and 128 units on the ProtT5 features; random forest, SVM, XGBoost, CNN1D and LSTM alternatives were also trained for comparison. Architectures and hyperparameters were chosen by tenfold cross-validation with grid search on the training set. The base modules were then frozen and their second-to-last-layer outputs concatenated into a 144-dimensional meta-feature (16 plus 128) used to train a feed-forward meta-classifier, optimised with Adam on binary cross-entropy with dropout and early stopping. The final model was applied to the held-aside independent test set, and learned features were projected with t-SNE (perplexity 30, learning rate 100), with models also retrained on 20-80% subsets of the training data.

The shape of the work

Structural · the record, drawn

ACQUISITIONPREPARATIONREPRESENTATIONTRAININGOPTIMISATIONTRAININGINFERENCEINTERPRETATION12345678AIAIAIAIAIAIReuse curatedsuccinylationsite datasetBuild windowpeptides andbalance negativesExtract ProtT5embeddings forlysine residuesTrain basemodules andalternative ML/D…Selectarchitectures andhyperparameters …Train stackedmeta-classifier(LMSuccSite)Score independenttest set andcompare with exi…Inspect learnedfeatures andtraining-size se…↤ conventional algorithm↤ physical experimentloops back
AI stepNo AI↤ what the AI stood in for
1Acquisition
no AI

Reuse curated succinylation site dataset

Obtaining raw data, whether by measurement, download or retrieval.

We used the dataset used during the development of DeepSuccinylSite to train and test our approach.where the paper describes this · verbatim
in the paper
2Preparation
no AI

Build window peptides and balance negatives

Cleaning, filtering, normalising or labelling data already obtained.

we performed random under sampling on the negative training set to obtain the same number of negative sites (4750)where the paper describes this · verbatim
in the paper
3Representation
AI

Extract ProtT5 embeddings for lysine residues

Encoding data into features, descriptors, embeddings or graphs. The AI stood in for conventional algorithm.

This model takes the overall protein sequence as an input and returns an embedding vector of dimension 1024 for each amino acid.where the paper describes this · verbatim
in the paper
4Training
AI

Train base modules and alternative ML/DL models

Fitting model parameters, including fine-tuning an existing model.

Initially, a 2D-CNN-based architecture was used for supervised embedding features while an artificial neural network (ANN)-based module was used for ProtT5 features.where the paper describes this · verbatim
in the paper
5Optimisation
AI

Select architectures and hyperparameters by cross-validated grid search

Iterative search over a space. Its result feeds back into an earlier step.

This architecture is chosen based on tenfold cross-validation on the training set using different architectures with different combinations of hyperparameters using grid search.where the paper describes this · verbatim
in the paper
6Training
AI

Train stacked meta-classifier (LMSuccSite)

Fitting model parameters, including fine-tuning an existing model.

the embedding module and Prot-T5 module were combined using an ANN as a meta-classifierwhere the paper describes this · verbatim
in the paper
7Inference
AI

Score independent test set and compare with existing predictors

Running a trained model over new data to predict, classify or score. The AI stood in for physical experiment.

we used the independent test set described in Table 1 and computed parameters such as accuracy, MCC, sensitivity, specificity and g-meanwhere the paper describes this · verbatim
in the paper
8Interpretation
AI

Inspect learned features and training-size sensitivity

Extracting understanding from model behaviour.

we created four different training datasets by randomly selecting 20%, 40%, 60%, and 80% of the samples from our training setwhere the paper describes this · verbatim
in the paper

What the record says

Technical · every part carries its own basis

+ in the paper~ our reading− not reported

How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.

~Role of AI
AI produced the resultour reading

The paper's result is a trained predictor of succinylation sites; the reported scores are the model's own outputs, so the finding exists only through the model

+What the AI was for
Classificationin the paper
we used the pretrained ProtT5 model to encode the featureswhere the paper describes this · verbatim
+How it was taught
SupervisedSelf-supervisedin the paper
+Models named
ProtT5-XL-UniRef50 · Off the shelfLMSuccSite (ANN meta-classifier) · Trained from scratchEmbedding module (CNN2D on supervised word embedding) · Trained from scratchProtT5 module (two-hidden-layer ANN) · Trained from scratchRandom forest (ProtT5 features) · Trained from scratchSupport vector machine (ProtT5 features) · Trained from scratchXGBoost (ProtT5 features) · Trained from scratchCNN1D (ProtT5 features) · Trained from scratchLSTM (supervised word embedding) · Trained from scratchin the paper
+How results were checked
Benchmarkin the paper
Importantly, the same training and test sets were used for all of the models tested.where the paper describes this · verbatim
+Code · weights · data
code availableweights availabledata availablein the paper
The source code and trained models are publicly available in the GitHub repositorywhere the paper describes this · verbatim
+Compute
not reportedin the paper

What this paper did not report

Technical · absence is published deliberately

Reported as not stated — 15 items
  • ComputeThe hardware or time used is not stated.
  • How many were testedThe paper gives no count of what was tested.
  • Version of ProtT5-XL-UniRef50Which version of the model was used is not stated.
  • Version of LMSuccSite (ANN meta-classifier)Which version of the model was used is not stated.
  • Version of Embedding module (CNN2D on supervised word embedding)Which version of the model was used is not stated.
  • Version of ProtT5 module (two-hidden-layer ANN)Which version of the model was used is not stated.
  • Version of Random forest (ProtT5 features)Which version of the model was used is not stated.
  • Version of Support vector machine (ProtT5 features)Which version of the model was used is not stated.
  • Version of XGBoost (ProtT5 features)Which version of the model was used is not stated.
  • Version of CNN1D (ProtT5 features)Which version of the model was used is not stated.
  • Version of LSTM (supervised word embedding)Which version of the model was used is not stated.
  • What step 4 replacedThe paper gives no basis for what the AI stood in for.
  • What step 5 replacedThe paper gives no basis for what the AI stood in for.
  • What step 6 replacedThe paper gives no basis for what the AI stood in for.
  • What step 8 replacedThe paper gives no basis for what the AI stood in for.

About this article

Record aix-00181, version 2, checked by a person on 2026-10-09. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error