~/aixsci
200 records · all checked

structural-biology/ai produced the result/Bioinformatics Advances 2024 · v2

Language models trained on protein sequences used to predict sugar-attachment sites

Researchers built StackGlyEmbed, a tool that predicts which points in a human protein carry an attached sugar chain. Three pre-trained protein language models turned each candidate site into numbers, and a stack of classifiers made the call.

1. Compile N-X-[S/T] sequon datasets and splits2. Extract protein language model embeddings and physicochemical features3. Select optimal feature groups by incremental feature selection4. Select base and meta learners5. Tune learner hyperparameters by grid search6. Train stacking ensemble on imbalance-corrected data7. Predict glycosylation sites on independent sequons8. Attribute feature-group importance with SHAP

spectrum · one line per step, placed by what the step does · bright lines used AI

StackGlyEmbed: prediction of N-linked glycosylation sites using protein language models
Bioinformatics Advances, 2024

doi:10.1093/bioadv/vbaf146 · record aix-00081 v2 · checked 2026-10-08

ai-resultrole of AI
AI was for
Classification
Model family
Protein language model, Transformer, Support vector machine, Gradient-boosted trees, Random forest, Multilayer perceptron, Linear model
Checked by
Benchmark
Code
available

The finding the paper is about came from the AI.

read as

The science is explained before the AI appears. Switch to field specialist to go straight to the method.

Assumes the discipline and goes straight to the method.

The diagram, the record and what the paper did not report are identical in both modes. Only the framing changes — never the evidence.

Introduction by AIxSci · plain language

What this research was about

Many proteins in the human body do not work alone: sugar chains are attached to them after they are built. One common form is N-linked glycosylation, where a sugar chain is fixed to an asparagine, one of the twenty amino acid building blocks of proteins. These sugars affect how a protein folds, where it travels and how it is recognised by other molecules. The attachment happens only at short patterns in the sequence, written N-X-[S/T]: an asparagine, then almost any amino acid, then a serine or a threonine. But not every such pattern, called a sequon, actually gets a sugar. Which ones do is hard to tell from the sequence alone.

Finding out experimentally means laboratory work such as mass spectrometry or radioactive labelling, which the authors note is expensive. So the researchers set out to predict the answer computationally. They gathered two existing collections of sequons, drawn from the N-GlyDE and N-GlycositeAtlas resources, re-checked them against the UniProt sequence database, and used them to train and test a predictor they call StackGlyEmbed.

Where AI came in

The prediction itself is the result, and it is made by machine learning throughout. Each candidate site was first described using three protein language models: ProtT5-XL-U50, ESM-2 and ProteinBERT. These are models trained on large numbers of protein sequences without being told anything about glycosylation, in the way a text model learns from reading. They turn a stretch of amino acids into lists of numbers, called embeddings, that capture something about its context. Here they were used unchanged, as fixed feature extractors. The authors describe earlier predictors as relying on hand-crafted input features instead.

A series of fitting-and-scoring rounds then chose which of those number sets to keep, which classifiers to use and what settings to give them. The chosen arrangement was a stacked ensemble: support vector machine, XGBoost and nearest-neighbour models in a first layer, whose output probabilities were passed to a support vector machine in a second layer. This was trained on deliberately rebalanced data, since glycosylated sites are the minority, then run over held-out test sequons and one case-study protein. A method called SHAP was used afterwards to ask which groups of features mattered most.

Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.

The work

Technical · from the record

StackGlyEmbed predicts which N-X-[S/T] sequons in human proteins carry N-linked glycosylation. Sites were represented with embeddings from three pre-trained protein language models (ProteinBERT, ESM-2 window-averaged, ProtT5-XL-U50 per-residue), selected by incremental feature selection, and classified by a stacking ensemble of SVM, XGBoost and KNN base learners feeding an SVM meta learner. On the N-GlyDE independent test set the authors report 98.2% sensitivity, 92.5% balanced accuracy, 89.1% F1-score, 82.6% MCC and 93.1% AUPR, and on the N-GlycositeAtlas independent test set 83.1% sensitivity, 69.3% F1-score, 51.5% MCC and 65.9% AUPR; the paper reports higher sensitivity, F1, MCC, AUROC and AUPR than the compared published methods on both test sets, while cross-dataset training and testing performed poorly. A SHAP analysis on a PCA-reduced feature space found the base learner probability group had the highest mean absolute SHAP value.

How AI was used

Two published datasets of N-X-[S/T] sequons (N-GlyDE and N-GlycositeAtlas) were re-checked against UniProt, sequons whose protein sequences had changed were dropped, and each training set was split into a 60% part for base learners and a 40% part for the meta learner. For every sequon, features were computed as per-residue, window-averaged (31 residues) and global embeddings from three frozen pre-trained protein language models (ProtT5-XL-U50, ESM-2 esm2_t33_650M_UR50D, ProteinBERT) together with averaged physicochemical parameters. Incremental feature selection, scored by the F1 of SVM models, chose ProteinBERT global, ESM-2 window and ProtT5-XL-U50 per-residue embeddings. Ten candidate classifiers were fitted and compared; incremental mutual information between learner outputs and labels selected SVM, XGBoost and KNN as base learners, and cross-validation on the BLP-augmented validation set selected SVM as meta learner, with grid search used for hyperparameters and window size. Features were Yeo–Johnson transformed; base learners were trained on multiple random undersamples of the majority class, and the meta learner was trained on randomly undersampled validation data whose feature vectors were augmented with base learner probabilities. The trained ensemble was then run over the held-out independent test sequons and a case-study protein, and SHAP values were computed on a PCA-reduced version of the feature space to attribute importance to each feature group.

The shape of the work

Structural · the record, drawn

PREPARATIONREPRESENTATIONSCREENINGSCREENINGOPTIMISATIONTRAININGINFERENCEINTERPRETATION12345678AIAIAIAIAIAIAICompile N-X-[S/T]sequon datasetsand splitsExtract proteinlanguage modelembeddings and p…Select optimalfeature groups byincremental feat…Select base andmeta learnersTune learnerhyperparametersby grid searchTrain stackingensemble onimbalance-correc…Predictglycosylationsites on indepen…Attributefeature-groupimportance with …↤ expert judgement↤ expert judgement↤ expert judgement↤ physical experiment
AI stepNo AI↤ what the AI stood in for
1Preparation
no AI

Compile N-X-[S/T] sequon datasets and splits

Cleaning, filtering, normalising or labelling data already obtained.

We further divided N-GlyDE-TS into two parts: 60% (N-GlyDE-TS-60) for training the base layer classifierswhere the paper describes this · verbatim
in the paper
2Representation
AI

Extract protein language model embeddings and physicochemical features

Encoding data into features, descriptors, embeddings or graphs. The AI stood in for expert judgement.

we used embedding features, obtained from various protein language models (PLMs), such as ProteinBERT, ESM-2, and ProtT5-XL-U50 embeddingswhere the paper describes this · verbatim
in the paper
3Screening
AI

Select optimal feature groups by incremental feature selection

Reducing a candidate set by filtering or ranking, in a single pass. The AI stood in for expert judgement.

For the feature selection process, we used IFS, with SVM as the classifier.where the paper describes this · verbatim
in the paper
4Screening
AI

Select base and meta learners

Reducing a candidate set by filtering or ranking, in a single pass. The AI stood in for expert judgement.

IMI resulted in the selection of SVM, XGB, and KNN as the base learners.where the paper describes this · verbatim
in the paper
5Optimisation
AI

Tune learner hyperparameters by grid search

Iterative search over a space.

We used Grid Search to tune the hyperparameters of the different classifiers.where the paper describes this · verbatim
in the paper
6Training
AI

Train stacking ensemble on imbalance-corrected data

Fitting model parameters, including fine-tuning an existing model.

We used the N-GlyDE-TS-60 set to train the base learners, while N-GlyDE-TS-40 was used to train the meta learner.where the paper describes this · verbatim
in the paper
7Inference
AI

Predict glycosylation sites on independent sequons

Running a trained model over new data to predict, classify or score. The AI stood in for physical experiment.

The overall framework for training and inference is shown in Fig. 2.where the paper describes this · verbatim
in the paper
8Interpretation
AI

Attribute feature-group importance with SHAP

Extracting understanding from model behaviour.

SHAP analysis was then performed and SHAP values were obtained for each feature of each sample.where the paper describes this · verbatim
in the paper

What the record says

Technical · every part carries its own basis

+ in the paper~ our reading− not reported

How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.

~Role of AI
AI produced the resultour reading

The reported result is the predictor itself: learned embeddings and fitted classifiers produce the N-linked glycosylation site calls the paper is about.

+What the AI was for
Classificationin the paper
built the stacking ensemble using Support Vector Machine (SVM), Extreme Gradient Boosting (XGB) and K-nearest Neighbor (KNN) learners in the base layerwhere the paper describes this · verbatim
+How it was taught
SupervisedSelf-supervisedin the paper
+Models named
ProtT5-XL-U50 · Off the shelfESM-2 esm2_t33_650M_UR50D · Off the shelfProteinBERT · Off the shelfSupport Vector Machine (base and meta learner) · Trained from scratchExtreme Gradient Boosting (XGB) · Trained from scratchK-nearest Neighbor (KNN) · Trained from scratchExtra Tree · Trained from scratchMulti-layer Perceptron · Trained from scratchPartial Least Square Regression · Trained from scratchLogistic Regression · Trained from scratchNaive Bayes · Trained from scratchRandom Forest · Trained from scratchDecision Tree · Trained from scratchFeed-forward neural network (FNN) baseline · Trained from scratchLMNglyPred · Off the shelfDeepNGlyPred · Off the shelfin the paper
+How results were checked
Benchmarkin the paper
Performance of different state-of-the-art (SOTA) methods has been compared with StackGlyEmbed on the N-GlycositeAtlas-IT and the N-GlyDE-IT datasetswhere the paper describes this · verbatim
+Code · weights · data
code availableweights availabledata availablein the paper
Datasets, StackGlyEmbed model and scripts to reproduce the results are available at https://github.com/nafcoder/StackGlyEmbedwhere the paper describes this · verbatim
+Compute
Hardware is reported only for an illustrative AlphaFold2 structure-prediction estimate (Intel Xeon CPU @ 2.20 GHz, Tesla T4 GPU, 16 GB memory, 15 GB VRAM, about 4 hours for a 1280-residue sequence); no compute for embedding extraction, training or inference is given.in the paper

What this paper did not report

Technical · absence is published deliberately

Reported as not stated — 19 items
  • How many were testedThe paper gives no count of what was tested.
  • Version of ProtT5-XL-U50Which version of the model was used is not stated.
  • Version of ProteinBERTWhich version of the model was used is not stated.
  • Version of Support Vector Machine (base and meta learner)Which version of the model was used is not stated.
  • Version of Extreme Gradient Boosting (XGB)Which version of the model was used is not stated.
  • Version of K-nearest Neighbor (KNN)Which version of the model was used is not stated.
  • Version of Extra TreeWhich version of the model was used is not stated.
  • Version of Multi-layer PerceptronWhich version of the model was used is not stated.
  • Version of Partial Least Square RegressionWhich version of the model was used is not stated.
  • Version of Logistic RegressionWhich version of the model was used is not stated.
  • Version of Naive BayesWhich version of the model was used is not stated.
  • Version of Random ForestWhich version of the model was used is not stated.
  • Version of Decision TreeWhich version of the model was used is not stated.
  • Version of Feed-forward neural network (FNN) baselineWhich version of the model was used is not stated.
  • Version of LMNglyPredWhich version of the model was used is not stated.
  • Version of DeepNGlyPredWhich version of the model was used is not stated.
  • What step 5 replacedThe paper gives no basis for what the AI stood in for.
  • What step 6 replacedThe paper gives no basis for what the AI stood in for.
  • What step 8 replacedThe paper gives no basis for what the AI stood in for.

About this article

Record aix-00081, version 2, checked by a person on 2026-10-08. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error