structural-biology/ai produced the result/Frontiers in Plant Science 2022 · v2
Machine learning sorts plant protein sequences into disease-resistance proteins or not
Researchers built StackRPred, a classifier that labels a plant protein sequence as a resistance protein or not. Six machine learning models feed a seventh, which makes the final call from features derived from a residue energy matrix.
spectrum · one line per step, placed by what the step does · bright lines used AI
Prediction of Plant Resistance Proteins Based on Pairwise Energy Content and Stacking Framework
Frontiers in Plant Science, 2022
doi:10.3389/fpls.2022.912599 · record aix-00107 v2 · checked 2026-10-08
- AI was for
- Classification
- Model family
- Support vector machine, Random forest, Gradient-boosted trees
- Checked by
- Held-out92 tested
- Code
- not reported
The finding the paper is about came from the AI.
What this research was about
Plants fight off disease with the help of so-called resistance proteins, or R proteins. These act rather like sentries inside the plant cell, recognising signs of an invading pathogen and setting off a defence response. Knowing which of a plant's many thousands of proteins are R proteins matters to anyone trying to breed hardier crops. But finding them is awkward. A protein is, at bottom, a long string of amino acids, and its job is not written plainly in that string. The usual approach is to compare a new sequence against known R proteins and look for family resemblance, which works less well when the resemblance is faint.
The authors set out to make that judgement from sequence alone, without relying on alignment to known examples. They gathered R proteins from a curated database alongside ordinary plant proteins from a public sequence archive, stripped out near-duplicates so that no two non-R sequences were more than 30 per cent similar, and split what remained into a training set and a separate test set of 92 sequences held back for the final check.
Where AI came in
The machine learning is not a tool used along the way here; it is the result. Each sequence was first turned into numbers by way of a published table of pairwise energies between amino acid types, treating the resulting grid as a kind of signal and summarising it with a wavelet transform, a standard way of describing a pattern at several scales at once. A further procedure, which repeatedly trains support vector machines and discards the least useful measurements, trimmed the description down to 112 numbers per protein.
Those numbers trained a two-layer arrangement called stacking. Six classifiers of different kinds sit in the first layer, among them random forests and gradient-boosted trees, and their verdicts become the input to a seventh, a support vector machine, which issues the final label. Settings were chosen by grid search. The result was checked by five-fold cross-validation on the training data and then on the held-back 92 sequences, where it reached an accuracy of 0.967 and an AUC of 0.997. Two earlier predictors were compared using figures from their own papers; one of them, prPred, reported higher precision.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record
StackRPred classifies plant protein sequences as resistance (R) proteins or not. Each sequence was encoded through a residue pairwise energy content matrix, from which discrete wavelet transform and PseRECM features were extracted and then reduced to 112 dimensions by SVM-RFE + CBR. The features were used to train a two-layer stacking classifier whose base layer holds six classifiers (KNN, GBDT, SVM, XGBoost, LightGBM, RF) and whose meta-layer is an SVM. On the 92-sample independent test set the authors report accuracy 0.967, precision 0.980, recall 0.968, F1-score 0.980 and AUC 0.997; the compared prPred model is reported with accuracy 0.935 and a precision of 1, higher than StackRPred's precision.
How AI was used
Machine learning supplied the predictor itself. Protein sequences from a curated set of plant R proteins and non-R proteins, deduplicated with CD-HIT at 30% similarity and split 8:2 into training and independent test partitions, were converted into numeric features by treating a published 20x20 residue pairwise energy content matrix as the encoding, extracting five-level discrete wavelet transform statistics and discrete cosine coefficients from the resulting two-dimensional signal, and computing PseRECM descriptors analogous to PsePSSM. An SVM-RFE + CBR procedure, which repeatedly fits SVMs to rank and eliminate features, reduced the feature space to 112 dimensions. These features trained a stacking ensemble: a base layer of KNN, GBDT, SVM, XGBoost, LightGBM and random forest, whose outputs became the input to an SVM meta-classifier. Grid search tuned the SVM cost and RBF kernel parameters, tree counts for XGBoost and random forest, and n_estimators, max_depth and learning_rate for LightGBM; KNN and GBDT used default parameters. The model was assessed by five-fold cross-validation on the training set and by scoring the independent test partition, with metrics compared against values published for prPred and prPred-DRLF.
The shape of the work
Structural · the record, drawn
no AI
Assemble plant R protein and non-R protein sequences
Obtaining raw data, whether by measurement, download or retrieval.
R proteins of 35 plant species were obtained from the PRGdb databasewhere the paper describes this · verbatim
no AI
Remove redundancy and split into training and test sets
Cleaning, filtering, normalising or labelling data already obtained.
proteins with sequence similarity greater than 30% were excluded from the non-R protein dataset using CD-HITwhere the paper describes this · verbatim
no AI
Encode sequences as residue pairwise energy features
Encoding data into features, descriptors, embeddings or graphs.
extract the PsePSSM and DWT characteristics of each Plant R protein based on the RECM matrixwhere the paper describes this · verbatim
AI
Select feature subset with SVM-RFE + CBR
Cleaning, filtering, normalising or labelling data already obtained.
we employ the SVM-RFE + CBR algorithm to select the best feature subsetwhere the paper describes this · verbatim
AI
Train two-layer stacking classifier
Fitting model parameters, including fine-tuning an existing model. The AI stood in for conventional algorithm.
the 112-dimensional feature information was fed into the constructed stacking model for trainingwhere the paper describes this · verbatim
AI
Evaluate by cross-validation and independent test
Testing outputs against ground truth.
To further compare the performance of our proposed method with other methods in independent testswhere the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
The paper's result is the trained classifier itself; the reported finding is the predictor's performance on plant R protein identification
we selected six classification algorithms as the base classifier for the first layerwhere the paper describes this · verbatim
the number of training samples is 364, and the number of independent test samples is 92where the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- CodeWhether the code is available is not stated.
- Trained model weightsWhether the trained model is available is not stated.
- DataWhether the data are available is not stated.
- ComputeThe hardware or time used is not stated.
- Version of XGBoostWhich version of the model was used is not stated.
- Version of LightGBMWhich version of the model was used is not stated.
- Version of Gradient Boosting Decision Tree (GBDT)Which version of the model was used is not stated.
- Version of Random ForestWhich version of the model was used is not stated.
- Version of K-Nearest NeighborWhich version of the model was used is not stated.
- Version of Support Vector Machine (RBF kernel; base classifier and meta-classifier)Which version of the model was used is not stated.
- What step 4 replacedThe paper gives no basis for what the AI stood in for.
- What step 6 replacedThe paper gives no basis for what the AI stood in for.
About this article
Record aix-00107, version 2, checked by a person on 2026-10-08. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error