~/aixsci
200 records · all checked

structural-biology/ai produced the result/Journal of Chemical Information and Modeling 2025 · v2

Machine learning tool flags protein helices likely to bind DNA or RNA

Researchers built a web server that scans a protein sequence for short helical stretches and predicts which ones bind nucleic acids. Eight machine learning classifiers, trained on physicochemical features, vote on each segment to produce the prediction.

1. Assemble labelled training sequences2. Scan sequences and extract helical segments3. Calculate physicochemical features4. Train and tune classifiers5. Select the eight core models6. Classify query helices with the ensemble7. Compute NABh consensus index and threshold8. Check recovery of known nucleic acid-binding proteins

spectrum · one line per step, placed by what the step does · bright lines used AI

NABhClassifier Server: A Tool for the Identification of Helical Nucleic Acid-Binding Sequences in Proteins
Journal of Chemical Information and Modeling, 2025

doi:10.1021/acs.jcim.4c02244 · record aix-00050 v2 · checked 2026-10-08

ai-resultrole of AI
AI was for
Classification
Model family
Gradient-boosted trees, Support vector machine, Random forest, Linear model, Multilayer perceptron, Recurrent neural network, Probabilistic graphical model
Checked by
Held-out474 tested, 454 worked
Code
not reported

The finding the paper is about came from the AI.

read as

The science is explained before the AI appears. Switch to field specialist to go straight to the method.

Assumes the discipline and goes straight to the method.

The diagram, the record and what the paper did not report are identical in both modes. Only the framing changes — never the evidence.

Introduction by AIxSci · plain language

What this research was about

Proteins do much of their work by touching DNA and RNA. Often the contact is made by a helix, a short stretch of the protein chain coiled into a spiral. Working out which helices in a protein grip nucleic acids normally means laboratory work, and there is no simple rule to read the answer off the sequence of amino acid letters. The chemistry that matters is spread across the segment: its electric charge, how readily it coils, how it behaves in water. Many different combinations of letters can produce a surface that binds, which makes the pattern hard to spot by eye.

The researchers set out to build a tool that takes a protein sequence a user submits, picks out its candidate helices, and reports which of them look like nucleic acid binders. They made it available as a web server called NABhClassifier.

Where AI came in

A rule-based step first marks stretches of six or more amino acids as candidate helices, and then twenty numerical descriptors are calculated for each one, including its charge at pH 7.4, its tendency to coil, its isoelectric point and a redox potential. The machine learning comes next. Fifteen supervised classifiers were trained from scratch on these descriptors, using 22-residue helices labelled as binders or non-binders: helices from a bacterial protein called KhpB as the positive examples, and albumin helices plus randomly generated sequences as the negatives. Settings were chosen by grid search, with repeated cross-validation.

Eight of the fifteen models, those with accuracy above 0.99, were kept and run together as an ensemble. Each returns a simple true or false for a segment; the share of models saying true becomes a consensus score, the NABh index, and segments scoring 0.75 or above are reported as candidates. The models stand in for experimental testing of each helix. On published sets of RNA-binding proteins the server picked out about 96% (454 of 474) of the double-stranded RNA binders.

Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.

The work

Technical · from the record

NABhClassifier is a web server that finds short helical segments in a protein sequence and classifies them as nucleic acid-binding or not. Fifteen machine learning classifiers were fitted to 20 physicochemical descriptors of 22-residue helices, using KhpB carboxy-terminal helices from 2560 bacterial species as positives and albumin helices plus randomly generated sequences as negatives; eight models with accuracy above 0.99 were kept and their votes combined into a consensus NABh index. Applied to published RNA-binding protein sets, the server recovered about 96% (454/474) of double-stranded RNA-binding proteins and about 95% (434/458) of single-stranded RNA-binding proteins at an index threshold of 0.75, and recovered 93–99% of keyword-annotated nucleic acid-binding proteins across five organism proteomes at the same threshold.

How AI was used

A rule-based scan first marks six-residue or longer windows enriched in alanine, glutamate, histidine, lysine, leucine, methionine, glutamine and arginine as candidate helices, and 20 sequence-derived features (charge at pH 7.4, residue-class frequencies, helix, sheet and turn propensities, hydrophobicity and hydropathy indices, instability index, isoelectric point, RNA ligation and redox potential) are computed for each segment. These features were used to fit fifteen supervised classifiers — gradient boosting, AdaBoost, bagging, stacking, random forest, support vector machine, logistic regression, ridge, linear and quadratic discriminant analysis, naive Bayes, Bernoulli naive Bayes, k-nearest neighbours, a multilayer perceptron and a recurrent neural network — on a 70/30 train/test split, with grid-searched hyperparameters and repeated cross-validation; models were also retrained with feature sets ranging from three to 20 features and feature relevance was averaged across the selected models. Eight models were retained and run as an ensemble over query sequences, and their individual TRUE calls are summed and divided by the number of models to give the NABh consensus index used to report candidate helices in the server.

The shape of the work

Structural · the record, drawn

ACQUISITIONPREPARATIONREPRESENTATIONTRAININGSCREENINGINFERENCESCREENINGVALIDATION12345678AIAIAssemble labelledtrainingsequencesScan sequencesand extracthelical segmentsCalculatephysicochemicalfeaturesTrain and tuneclassifiersSelect the eightcore modelsClassify queryhelices with theensembleCompute NABhconsensus indexand thresholdCheck recovery ofknown nucleicacid-binding pro…↤ physical experiment
AI stepNo AI↤ what the AI stood in for
1Acquisition
no AI

Assemble labelled training sequences

Obtaining raw data, whether by measurement, download or retrieval.

The positive set was formed by segments extracted from the KhpB carboxy-terminal α helix from 2560 different bacterial specieswhere the paper describes this · verbatim
in the paper
2Preparation
no AI

Scan sequences and extract helical segments

Cleaning, filtering, normalising or labelling data already obtained.

only helices consisting of six or more residues are extracted and advanced to the next step in the pipelinewhere the paper describes this · verbatim
in the paper
3Representation
no AI

Calculate physicochemical features

Encoding data into features, descriptors, embeddings or graphs.

For each input sequence, a total of 20 features are calculated.where the paper describes this · verbatim
in the paper
4Training
AI

Train and tune classifiers

Fitting model parameters, including fine-tuning an existing model.

Hyperparameters were set using Grid search technique, and each model underwent cross-validation through 50 iterations.where the paper describes this · verbatim
in the paper
5Screening
no AI

Select the eight core models

Reducing a candidate set by filtering or ranking, in a single pass.

Among the 15 ML models tested during training, eight demonstrated accuracy above 0.99 and were selected to form the NABhClassifier corewhere the paper describes this · verbatim
in the paper
6Inference
AI

Classify query helices with the ensemble

Running a trained model over new data to predict, classify or score. The AI stood in for physical experiment.

NABhClassifier was challenged with a data set consisting of 458 single-stranded RNA-binding proteins (ssRBPs) and 474 double-stranded RNA-binding proteins (dsRBPs)where the paper describes this · verbatim
in the paper
7Screening
no AI

Compute NABh consensus index and threshold

Reducing a candidate set by filtering or ranking, in a single pass.

Helical sequences with an NABh index of 0.75 or higher are displayed on the “Results” pagewhere the paper describes this · verbatim
in the paper
8Validation
no AI

Check recovery of known nucleic acid-binding proteins

Testing outputs against ground truth.

To validate the NABhClassifier server, we tested it on two case studies.where the paper describes this · verbatim
in the paper

What the record says

Technical · every part carries its own basis

+ in the paper~ our reading− not reported

How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.

~Role of AI
AI produced the resultour reading

The paper's object is a classifier server whose reported output — identification of nucleic acid-binding helices and the consensus NABh index — is produced entirely by the eight machine learning models

+What the AI was for
Classificationin the paper
models were trained using features calculated from 22 amino acid long helical sequenceswhere the paper describes this · verbatim
+How it was taught
Supervisedin the paper
+Models named
Gradient Boosting Classifier (GBC) · Trained from scratchSupport Vector Machine (SVM) · Trained from scratchAda Boosting · Trained from scratchLogistic Regression (LogReg) · Trained from scratchRidge · Trained from scratchStacking · Trained from scratchRandom Forest (RF) · Trained from scratchBagging · Trained from scratchMulti-Layer Perceptron (MLP) · Trained from scratchNaive Bayes (NB) · Trained from scratchBernoulli Naive Bayes (BNB) · Trained from scratchLinear Discriminant Analysis (LDA) · Trained from scratchQuadratic Discriminant Analysis (QDA) · Trained from scratchk-Nearest Neighbors (KNN) · Trained from scratchRecurrent Neural Network (RNN) · Trained from scratchin the paper
+How results were checked
Held-out474 tested, 454 workedin the paper
a recovery rate of about 96% (454/474) for dsRBPwhere the paper describes this · verbatim
+Code · weights · data
code not reportedweights not reporteddata availablein the paper
Sequences of all proteins are provided in Supplementary File 1.where the paper describes this · verbatim
+Compute
Hosted on a 12-core virtual machine; a typical submission takes less than 2 s to processin the paper

What this paper did not report

Technical · absence is published deliberately

Reported as not stated — 18 items
  • CodeWhether the code is available is not stated.
  • Trained model weightsWhether the trained model is available is not stated.
  • Version of Gradient Boosting Classifier (GBC)Which version of the model was used is not stated.
  • Version of Support Vector Machine (SVM)Which version of the model was used is not stated.
  • Version of Ada BoostingWhich version of the model was used is not stated.
  • Version of Logistic Regression (LogReg)Which version of the model was used is not stated.
  • Version of RidgeWhich version of the model was used is not stated.
  • Version of StackingWhich version of the model was used is not stated.
  • Version of Random Forest (RF)Which version of the model was used is not stated.
  • Version of BaggingWhich version of the model was used is not stated.
  • Version of Multi-Layer Perceptron (MLP)Which version of the model was used is not stated.
  • Version of Naive Bayes (NB)Which version of the model was used is not stated.
  • Version of Bernoulli Naive Bayes (BNB)Which version of the model was used is not stated.
  • Version of Linear Discriminant Analysis (LDA)Which version of the model was used is not stated.
  • Version of Quadratic Discriminant Analysis (QDA)Which version of the model was used is not stated.
  • Version of k-Nearest Neighbors (KNN)Which version of the model was used is not stated.
  • Version of Recurrent Neural Network (RNN)Which version of the model was used is not stated.
  • What step 4 replacedThe paper gives no basis for what the AI stood in for.

About this article

Record aix-00050, version 2, checked by a person on 2026-10-08. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error