~/aixsci
200 records · all checked

structural-biology/ai produced the result/PLoS ONE 2026 · v2

Machine learning sorts antibody heavy chains by the germ they target

Researchers turned 1111 antibody heavy chain sequences into numerical descriptors and trained tree-based classifiers, an ensemble and a Transformer to predict which of five antigens each antibody targets, in place of laboratory binding tests.

1. Retrieve antigen-specific heavy chain sequences2. Preprocess and deduplicate sequences3. Encode sequences as engineered feature vectors4. Train classifiers and tune hyperparameters5. Classify held-out test sequences6. External validation on an independent dataset7. Attribute predictions to sequence features

spectrum · one line per step, placed by what the step does · bright lines used AI

Computational models for the classification of antibody specificity using heavy chain features
PLoS ONE, 2026

doi:10.1371/journal.pone.0349143 · record aix-00188 v2 · checked 2026-10-09

ai-resultrole of AI
AI was for
Classification
Model family
Gradient-boosted trees, Random forest, Transformer, Support vector machine, Linear model
Checked by
Held-out
Code
available

The finding the paper is about came from the AI.

read as

The science is explained before the AI appears. Switch to field specialist to go straight to the method.

Assumes the discipline and goes straight to the method.

The diagram, the record and what the paper did not report are identical in both modes. Only the framing changes — never the evidence.

Introduction by AIxSci · plain language

What this research was about

Antibodies are proteins the immune system makes to latch onto a specific invader. Each one is built from two kinds of chain, and the heavier of the two carries much of the region that does the gripping. Which target an antibody binds is normally established in the laboratory, by testing it against candidate molecules. That is slow and costly, and databases now hold far more antibody sequences than anyone could test. So a question arises: does the amino acid sequence of the heavy chain alone carry enough signal to say what the antibody was raised against? The chains that bind very different germs are not obviously different to the eye.

The authors gathered antibody heavy chain sequences from a public protein database for five targets: the dengue virus, influenza, tetanus, SARS-CoV-2 and the tuberculosis bacterium. After removing near-duplicates, 1111 sequences remained. Each was summarised as a list of 81 numbers describing its amino acid composition, the order of its residues, and chemical properties such as charge and bulk. The task was then to see whether a computer could read those numbers back to the correct antigen class, and which features mattered.

Where AI came in

The machine learning did the classifying. Several kinds of model were trained from scratch on the numerical descriptors, with the correct antigen supplied as the answer during training: five tree-based methods including CatBoost and XGBoost, a stacked arrangement in which simpler models feed a further one, and a Transformer, the architecture behind recent language models, here reading feature lists rather than text. Settings were tuned by repeatedly holding back portions of the training data. The models were then asked about a fifth of the sequences they had never seen, and about a separate published dataset.

Those predictions are the result. On the held-out sequences the stacked model was correct 0.7803 of the time and CatBoost 0.7713, while the Transformer reached 0.7399. On the external data the stacked model fell to 0.7240 and XGBoost led at 0.7729. A method called SHAP, which apportions credit among the inputs, ranked sequence-order and evolutionary descriptors highest, with the amino acid cysteine contributing most. The models stand in for the binding experiments that would otherwise assign each antibody to a target.

Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.

The work

Technical · from the record

The authors assembled 1111 low-redundancy antibody heavy chain sequences from the NCBI Protein database across five antigen classes (anti-dengue, anti-influenza, anti-tetanus, anti-SARS-CoV-2 and anti-Mycobacterium tuberculosis) and encoded each sequence as an 81-dimensional vector of evolutionary, sequence-order and physicochemical descriptors. Tree-based classifiers, a stacking ensemble and a feature-based Transformer were trained to predict the antigen class from these features; on the held-out test set CatBoost reached an accuracy of 0.7713, the stacking ensemble 0.7803, and the Transformer an accuracy of 0.7399 with an F1-score of 0.6761. On an external dataset of translated heavy chain sequences the stacking model's accuracy fell to 0.7240 while XGBoost reached the highest accuracy of 0.7729. SHAP attribution ranked sequence-order (PseAAC) and evolutionary (AAC-PSSM) features highest, with cysteine the largest single contributor.

How AI was used

Heavy chain amino acid sequences were retrieved from NCBI with per-antigen search strings, clustered with CD-HIT at a 40% identity threshold, length-standardised and mapped to standard residues, then converted by web servers (POSSUM, the PseAAC server and iFeature) into an 81-dimensional vector of 20 AAC-PSSM, 22 PseAAC and 39 CTDC descriptors. These vectors were z-score normalised and used to fit supervised multiclass classifiers: XGBoost, LightGBM, Random Forest, CatBoost and AdaBoost as single models, a scikit-learn StackingClassifier with logistic regression, SVM and KNN base learners feeding a Random Forest meta-learner trained on out-of-fold predictions, and a feature-based Transformer with eight-head self-attention, a 256-dimensional feedforward block, six encoder layers, L2 regularisation and dropout. Hyperparameters were chosen by Bayesian optimisation with 5-fold cross-validation on an 80/20 stratified split, the trained models were then run over the held-out split and over an independently published, translated heavy chain dataset processed through the same pipeline, and SHAP was applied to the CatBoost, Stacking and Transformer models to rank feature contributions overall and per antibody class.

The shape of the work

Structural · the record, drawn

ACQUISITIONPREPARATIONREPRESENTATIONTRAININGINFERENCEVALIDATIONINTERPRETATION1234567AIAIAIRetrieveantigen-specificheavy chain sequ…Preprocess anddeduplicatesequencesEncode sequencesas engineeredfeature vectorsTrain classifiersand tunehyperparametersClassify held-outtest sequencesExternalvalidation on anindependent data…Attributepredictions tosequence features↤ physical experiment↤ physical experiment
AI stepNo AI↤ what the AI stood in for
1Acquisition
no AI

Retrieve antigen-specific heavy chain sequences

Obtaining raw data, whether by measurement, download or retrieval.

We collected antigen-specific immunoglobulin sequences from the National Center for Biotechnology Information (NCBI) Protein databasewhere the paper describes this · verbatim
in the paper
2Preparation
no AI

Preprocess and deduplicate sequences

Cleaning, filtering, normalising or labelling data already obtained.

we utilized CD-HIT(Cluster Database at High Identity with Tolerance) to cluster sequences at a 40% identity thresholdwhere the paper describes this · verbatim
in the paper
3Representation
no AI

Encode sequences as engineered feature vectors

Encoding data into features, descriptors, embeddings or graphs.

an 81-dimensional feature vector was generated for each antibody sequence using three complementary encoding strategieswhere the paper describes this · verbatim
in the paper
4Training
AI

Train classifiers and tune hyperparameters

Fitting model parameters, including fine-tuning an existing model.

Hyperparameter optimization was performed using Bayesian optimization with 5-fold cross-validation.where the paper describes this · verbatim
in the paper
5Inference
AI

Classify held-out test sequences

Running a trained model over new data to predict, classify or score. The AI stood in for physical experiment.

The dataset was randomly split into training (80%) and test (20%) sets, maintaining class distribution.where the paper describes this · verbatim
in the paper
6Validation
AI

External validation on an independent dataset

Testing outputs against ground truth. The AI stood in for physical experiment.

we conducted rigorous external validation using an independent dataset from Wang et al. (2022)where the paper describes this · verbatim
in the paper
7Interpretation
no AI

Attribute predictions to sequence features

Extracting understanding from model behaviour.

We utilized SHAP (SHapley Additive exPlanations) to explain our model outputs by assigning importance values to features.where the paper describes this · verbatim
in the paper

What the record says

Technical · every part carries its own basis

+ in the paper~ our reading− not reported

How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.

~Role of AI
AI produced the resultour reading

The paper's result is the classification performance of the trained models themselves; every reported finding, including the SHAP feature rankings, comes from model output

+What the AI was for
Classificationin the paper
Our approach begins with five robust tree-based algorithms as primary classifierswhere the paper describes this · verbatim
+How it was taught
Supervisedin the paper
+Models named
CatBoost · Trained from scratchXGBoost · Trained from scratchLightGBM · Trained from scratchRandom Forest · Trained from scratchAdaBoost · Trained from scratchStacking ensemble (logistic regression, SVM, KNN base learners with Random Forest meta-learner) · Trained from scratchFeature-Based Transformer · Trained from scratchin the paper
+How results were checked
Held-outin the paper
To assess the generalizability of our models, we conducted rigorous external validation using an independent datasetwhere the paper describes this · verbatim
+Code · weights · data
code availableweights not reporteddata availablein the paper
The source code is available on GitHub (https://github.com/LJxp22/AbClass-Classifier).where the paper describes this · verbatim
+Compute
not reportedin the paper

What this paper did not report

Technical · absence is published deliberately

Reported as not stated — 11 items
  • Trained model weightsWhether the trained model is available is not stated.
  • ComputeThe hardware or time used is not stated.
  • How many were testedThe paper gives no count of what was tested.
  • Version of CatBoostWhich version of the model was used is not stated.
  • Version of XGBoostWhich version of the model was used is not stated.
  • Version of LightGBMWhich version of the model was used is not stated.
  • Version of Random ForestWhich version of the model was used is not stated.
  • Version of AdaBoostWhich version of the model was used is not stated.
  • Version of Stacking ensemble (logistic regression, SVM, KNN base learners with Random Forest meta-learner)Which version of the model was used is not stated.
  • Version of Feature-Based TransformerWhich version of the model was used is not stated.
  • What step 4 replacedThe paper gives no basis for what the AI stood in for.

About this article

Record aix-00188, version 2, checked by a person on 2026-10-09. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error