structural-biology/ai produced the result/Bioinformatics Advances 2024 · v2
Language models trained on protein sequences used to predict sugar-attachment sites
Researchers built StackGlyEmbed, a tool that predicts which points in a human protein carry an attached sugar chain. Three pre-trained protein language models turned each candidate site into numbers, and a stack of classifiers made the call.
spectrum · one line per step, placed by what the step does · bright lines used AI
StackGlyEmbed: prediction of N-linked glycosylation sites using protein language models
Bioinformatics Advances, 2024
doi:10.1093/bioadv/vbaf146 · record aix-00081 v2 · checked 2026-10-08
- AI was for
- Classification
- Model family
- Protein language model, Transformer, Support vector machine, Gradient-boosted trees, Random forest, Multilayer perceptron, Linear model
- Checked by
- Benchmark
- Code
- available
The finding the paper is about came from the AI.
What this research was about
Many proteins in the human body do not work alone: sugar chains are attached to them after they are built. One common form is N-linked glycosylation, where a sugar chain is fixed to an asparagine, one of the twenty amino acid building blocks of proteins. These sugars affect how a protein folds, where it travels and how it is recognised by other molecules. The attachment happens only at short patterns in the sequence, written N-X-[S/T]: an asparagine, then almost any amino acid, then a serine or a threonine. But not every such pattern, called a sequon, actually gets a sugar. Which ones do is hard to tell from the sequence alone.
Finding out experimentally means laboratory work such as mass spectrometry or radioactive labelling, which the authors note is expensive. So the researchers set out to predict the answer computationally. They gathered two existing collections of sequons, drawn from the N-GlyDE and N-GlycositeAtlas resources, re-checked them against the UniProt sequence database, and used them to train and test a predictor they call StackGlyEmbed.
Where AI came in
The prediction itself is the result, and it is made by machine learning throughout. Each candidate site was first described using three protein language models: ProtT5-XL-U50, ESM-2 and ProteinBERT. These are models trained on large numbers of protein sequences without being told anything about glycosylation, in the way a text model learns from reading. They turn a stretch of amino acids into lists of numbers, called embeddings, that capture something about its context. Here they were used unchanged, as fixed feature extractors. The authors describe earlier predictors as relying on hand-crafted input features instead.
A series of fitting-and-scoring rounds then chose which of those number sets to keep, which classifiers to use and what settings to give them. The chosen arrangement was a stacked ensemble: support vector machine, XGBoost and nearest-neighbour models in a first layer, whose output probabilities were passed to a support vector machine in a second layer. This was trained on deliberately rebalanced data, since glycosylated sites are the minority, then run over held-out test sequons and one case-study protein. A method called SHAP was used afterwards to ask which groups of features mattered most.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record
StackGlyEmbed predicts which N-X-[S/T] sequons in human proteins carry N-linked glycosylation. Sites were represented with embeddings from three pre-trained protein language models (ProteinBERT, ESM-2 window-averaged, ProtT5-XL-U50 per-residue), selected by incremental feature selection, and classified by a stacking ensemble of SVM, XGBoost and KNN base learners feeding an SVM meta learner. On the N-GlyDE independent test set the authors report 98.2% sensitivity, 92.5% balanced accuracy, 89.1% F1-score, 82.6% MCC and 93.1% AUPR, and on the N-GlycositeAtlas independent test set 83.1% sensitivity, 69.3% F1-score, 51.5% MCC and 65.9% AUPR; the paper reports higher sensitivity, F1, MCC, AUROC and AUPR than the compared published methods on both test sets, while cross-dataset training and testing performed poorly. A SHAP analysis on a PCA-reduced feature space found the base learner probability group had the highest mean absolute SHAP value.
How AI was used
Two published datasets of N-X-[S/T] sequons (N-GlyDE and N-GlycositeAtlas) were re-checked against UniProt, sequons whose protein sequences had changed were dropped, and each training set was split into a 60% part for base learners and a 40% part for the meta learner. For every sequon, features were computed as per-residue, window-averaged (31 residues) and global embeddings from three frozen pre-trained protein language models (ProtT5-XL-U50, ESM-2 esm2_t33_650M_UR50D, ProteinBERT) together with averaged physicochemical parameters. Incremental feature selection, scored by the F1 of SVM models, chose ProteinBERT global, ESM-2 window and ProtT5-XL-U50 per-residue embeddings. Ten candidate classifiers were fitted and compared; incremental mutual information between learner outputs and labels selected SVM, XGBoost and KNN as base learners, and cross-validation on the BLP-augmented validation set selected SVM as meta learner, with grid search used for hyperparameters and window size. Features were Yeo–Johnson transformed; base learners were trained on multiple random undersamples of the majority class, and the meta learner was trained on randomly undersampled validation data whose feature vectors were augmented with base learner probabilities. The trained ensemble was then run over the held-out independent test sequons and a case-study protein, and SHAP values were computed on a PCA-reduced version of the feature space to attribute importance to each feature group.
The shape of the work
Structural · the record, drawn
no AI
Compile N-X-[S/T] sequon datasets and splits
Cleaning, filtering, normalising or labelling data already obtained.
We further divided N-GlyDE-TS into two parts: 60% (N-GlyDE-TS-60) for training the base layer classifierswhere the paper describes this · verbatim
AI
Extract protein language model embeddings and physicochemical features
Encoding data into features, descriptors, embeddings or graphs. The AI stood in for expert judgement.
we used embedding features, obtained from various protein language models (PLMs), such as ProteinBERT, ESM-2, and ProtT5-XL-U50 embeddingswhere the paper describes this · verbatim
AI
Select optimal feature groups by incremental feature selection
Reducing a candidate set by filtering or ranking, in a single pass. The AI stood in for expert judgement.
For the feature selection process, we used IFS, with SVM as the classifier.where the paper describes this · verbatim
AI
Select base and meta learners
Reducing a candidate set by filtering or ranking, in a single pass. The AI stood in for expert judgement.
IMI resulted in the selection of SVM, XGB, and KNN as the base learners.where the paper describes this · verbatim
AI
Tune learner hyperparameters by grid search
Iterative search over a space.
We used Grid Search to tune the hyperparameters of the different classifiers.where the paper describes this · verbatim
AI
Train stacking ensemble on imbalance-corrected data
Fitting model parameters, including fine-tuning an existing model.
We used the N-GlyDE-TS-60 set to train the base learners, while N-GlyDE-TS-40 was used to train the meta learner.where the paper describes this · verbatim
AI
Predict glycosylation sites on independent sequons
Running a trained model over new data to predict, classify or score. The AI stood in for physical experiment.
The overall framework for training and inference is shown in Fig. 2.where the paper describes this · verbatim
AI
Attribute feature-group importance with SHAP
Extracting understanding from model behaviour.
SHAP analysis was then performed and SHAP values were obtained for each feature of each sample.where the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
The reported result is the predictor itself: learned embeddings and fitted classifiers produce the N-linked glycosylation site calls the paper is about.
built the stacking ensemble using Support Vector Machine (SVM), Extreme Gradient Boosting (XGB) and K-nearest Neighbor (KNN) learners in the base layerwhere the paper describes this · verbatim
Performance of different state-of-the-art (SOTA) methods has been compared with StackGlyEmbed on the N-GlycositeAtlas-IT and the N-GlyDE-IT datasetswhere the paper describes this · verbatim
Datasets, StackGlyEmbed model and scripts to reproduce the results are available at https://github.com/nafcoder/StackGlyEmbedwhere the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- How many were testedThe paper gives no count of what was tested.
- Version of ProtT5-XL-U50Which version of the model was used is not stated.
- Version of ProteinBERTWhich version of the model was used is not stated.
- Version of Support Vector Machine (base and meta learner)Which version of the model was used is not stated.
- Version of Extreme Gradient Boosting (XGB)Which version of the model was used is not stated.
- Version of K-nearest Neighbor (KNN)Which version of the model was used is not stated.
- Version of Extra TreeWhich version of the model was used is not stated.
- Version of Multi-layer PerceptronWhich version of the model was used is not stated.
- Version of Partial Least Square RegressionWhich version of the model was used is not stated.
- Version of Logistic RegressionWhich version of the model was used is not stated.
- Version of Naive BayesWhich version of the model was used is not stated.
- Version of Random ForestWhich version of the model was used is not stated.
- Version of Decision TreeWhich version of the model was used is not stated.
- Version of Feed-forward neural network (FNN) baselineWhich version of the model was used is not stated.
- Version of LMNglyPredWhich version of the model was used is not stated.
- Version of DeepNGlyPredWhich version of the model was used is not stated.
- What step 5 replacedThe paper gives no basis for what the AI stood in for.
- What step 6 replacedThe paper gives no basis for what the AI stood in for.
- What step 8 replacedThe paper gives no basis for what the AI stood in for.
About this article
Record aix-00081, version 2, checked by a person on 2026-10-08. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error