~/aixsci
200 records · all checked

structural-biology/ai produced the result/Frontiers in Microbiology 2026 · v2

Protein language model embeddings tested against classical descriptors for spotting resistance genes

Researchers compared two ways of turning bacterial protein sequences into numbers for predicting antimicrobial resistance. One used hand-computed chemical descriptors; the other used frozen embeddings from the ESM-2 protein language model, fed to five classifiers.

1. Collect AMR-positive and candidate negative sequences2. Filter, de-duplicate, balance and partition dataset3. Compute classical sequence descriptors4. Generate ESM-2 protein embeddings5. Train and tune classifiers on each feature space6. Apply locked models to test, control and external sets7. Benchmark against existing AMR detection tools8. Structural and phylogenetic error analysis of persistent misclassifications

spectrum · one line per step, placed by what the step does · bright lines used AI

Computational mapping of resistance-relevant signals in proteins using deep and classical feature spaces
Frontiers in Microbiology, 2026

doi:10.3389/fmicb.2026.1899081 · record aix-00112 v2 · checked 2026-10-08

ai-resultrole of AI
AI was for
Classification, Structure determination
Model family
Transformer, Protein language model, Linear model, Support vector machine, Random forest, Multilayer perceptron
Checked by
Held-out908 tested
Code
available

The finding the paper is about came from the AI.

read as

The science is explained before the AI appears. Switch to field specialist to go straight to the method.

Assumes the discipline and goes straight to the method.

The diagram, the record and what the paper did not report are identical in both modes. Only the framing changes — never the evidence.

Introduction by AIxSci · plain language

What this research was about

Bacteria can shrug off the drugs meant to kill them, and the instructions for doing so are written in their genes. Each resistance gene makes a protein — an enzyme that chops up an antibiotic, a pump that throws it out, a target reshaped so the drug no longer fits. Scanning a bacterium's proteins for these is routine public-health work. The usual method compares a new sequence letter by letter against catalogues of known resistance proteins. That works well when the new protein closely resembles something already catalogued, and less well when it does not, since resistance can arise in proteins that look unfamiliar.

An alternative is to skip the comparison entirely and instead describe each protein as a list of numbers, then let a classifier learn which patterns of numbers go with resistance. The question is which numbers to use. This study put two recipes side by side: classical descriptors, counting amino acids and short runs of them along with physical and chemical properties, and embeddings from ESM-2, a model trained on vast numbers of protein sequences to produce a numerical summary of each one.

Where AI came in

AI appears twice over. ESM-2 is a transformer, the same broad design behind language models for text, but trained on protein sequences instead of sentences. It learns by predicting masked-out amino acids, which forces it to pick up regularities of protein structure and function. Here it was used with its weights frozen and no further training: each sequence went in, and a single averaged vector came out. That vector stood in for the hand-designed descriptor list, and for sequence alignment against reference databases.

Five classifiers — logistic regression, a support vector machine, a random forest, a feed-forward neural network, and a stacked combination of the four — were then trained on each set of numbers and tested on 908 held-out proteins and 527 further sequences. Existing resistance-detection tools were run on the same test set for comparison. AlphaFold3, which predicts a protein's three-dimensional shape from its sequence, was used to model the 17 proteins that every version got wrong, standing in for laboratory structure determination. The study's findings are entirely classifier performance figures; no non-AI result is reported.

Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.

The work

Technical · from the record

The study compared two alignment-free ways of representing bacterial proteins for antimicrobial resistance detection: hand-computed compositional and physicochemical descriptors, and frozen embeddings from the ESM-2 protein language model. Logistic regression, an SVM, a random forest, a feed-forward neural network and a stacked fusion of the four were fitted on a balanced, redundancy-filtered dataset and evaluated on a held-out test set of 908 proteins and an external cohort of 527 sequences. On the held-out test set, AUROC ran from 0.817 to 0.850 for the classical descriptors, while the fusion model on deep embeddings reached 0.966; on the external cohort the neural network on deep embeddings reached an AUROC of 0.9717, and concatenating the two feature sets scored lower than embeddings alone across matching models. Label-permutation and sequence-scrambling controls dropped performance to AUROC 0.492 and to a range of 0.57 to 0.65 respectively, and AlphaFold3 structures plus phylogenetics were used to examine 17 proteins misclassified across all three feature spaces.

How AI was used

AMR-positive proteins were pooled from CARD, AMRFinderPlus and ResFinder, length- and residue-filtered, clustered with CD-HIT at 95% identity, and balanced against RefSeq proteins screened for latent resistance homology, then split 70/15/15. Each sequence was encoded twice: as classical amino acid, dipeptide, tripeptide and physicochemical descriptors, and as a mean-pooled protein-level vector from a pretrained ESM-2 transformer used with fixed weights and no fine-tuning; the two were also concatenated. For each of the three feature spaces, logistic regression, an RBF support vector machine, a random forest, a feed-forward neural network and a stacking fusion model were fitted on the training partition inside scikit-learn pipelines, with hyperparameters and decision thresholds tuned on the validation partition, 10-fold cross-validation on the pooled training-validation space, repetition across 10 random seeds, and a 50-iteration randomized re-partitioning procedure. The locked models were then run without retraining on the internal test partition, on label-shuffled and window-scrambled control matrices built through the same feature pipeline, and on an external beta-lactamase and RefSeq cohort. Existing tools (DeepARG, PLM-ARG, ProtAlign-ARG, AMRFinderPlus) were run on the same test set for comparison, and AlphaFold3 was used to model structures of the proteins misclassified across all feature spaces.

The shape of the work

Structural · the record, drawn

ACQUISITIONPREPARATIONREPRESENTATIONREPRESENTATIONTRAININGINFERENCEVALIDATIONINTERPRETATION12345678AIAIAIAIAICollectAMR-positive andcandidate negati…Filter,de-duplicate,balance and part…Compute classicalsequencedescriptorsGenerate ESM-2proteinembeddingsTrain and tuneclassifiers oneach feature spa…Apply lockedmodels to test,control and exte…Benchmark againstexisting AMRdetection toolsStructural andphylogeneticerror analysis o…↤ conventional algorithm↤ conventional algorithm↤ conventional algorithm↤ physical experiment
AI stepNo AI↤ what the AI stood in for
1Acquisition
no AI

Collect AMR-positive and candidate negative sequences

Obtaining raw data, whether by measurement, download or retrieval.

AMR-associated proteins were collected from three established databases.where the paper describes this · verbatim
in the paper
2Preparation
no AI

Filter, de-duplicate, balance and partition dataset

Cleaning, filtering, normalising or labelling data already obtained.

Redundancy was reduced by using CD-HIT clustering at 95% sequence identity to remove near-identical sequences while preserving functional diversity.where the paper describes this · verbatim
in the paper
3Representation
no AI

Compute classical sequence descriptors

Encoding data into features, descriptors, embeddings or graphs.

Proteins were encoded into fixed-length vectors using established descriptors.where the paper describes this · verbatim
in the paper
4Representation
AI

Generate ESM-2 protein embeddings

Encoding data into features, descriptors, embeddings or graphs. The AI stood in for conventional algorithm.

Protein sequences were processed using a pretrained ESM-2 model with fixed weights without any fine-tuningwhere the paper describes this · verbatim
in the paper
5Training
AI

Train and tune classifiers on each feature space

Fitting model parameters, including fine-tuning an existing model. The AI stood in for conventional algorithm.

All models were trained and 10-fold cross validated using the Training dataset (n = 4,269 proteins)where the paper describes this · verbatim
in the paper
6Inference
AI

Apply locked models to test, control and external sets

Running a trained model over new data to predict, classify or score. The AI stood in for conventional algorithm.

the pre-trained, locked machine learning models were evaluated on external dataset (n = 527 sequences)where the paper describes this · verbatim
in the paper
7Validation
AI

Benchmark against existing AMR detection tools

Testing outputs against ground truth.

we benchmarked it against diverse homology-, machine learning- and deep learning-based AMR detection toolswhere the paper describes this · verbatim
in the paper
8Interpretation
AI

Structural and phylogenetic error analysis of persistent misclassifications

Extracting understanding from model behaviour. The AI stood in for physical experiment.

we performed phylogenetic analysis and extracted high-resolution structural metrics using a standalone implementation of AlphaFold3where the paper describes this · verbatim
in the paper

What the record says

Technical · every part carries its own basis

+ in the paper~ our reading− not reported

How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.

~Role of AI
AI produced the resultour reading

The paper's reported findings are the discriminative performance of trained classifiers over classical and protein language model feature spaces; no non-AI result is reported

+What the AI was for
Four supervised classifiers representing distinct learning paradigms were evaluated: Logistic Regression (linear), Support Vector Machine (kernel-based), Random Forest (ensemble)where the paper describes this · verbatim
+How it was taught
SupervisedSelf-supervisedin the paper
+Models named
ESM-2 · Off the shelfLogistic Regression · Trained from scratchSupport Vector Machine (RBF kernel) · Trained from scratchRandom Forest · Trained from scratchFeed-forward Neural Network · Trained from scratchFusion Model (stacked LR, SVM, RF, NN) · Trained from scratchAlphaFold3 · Off the shelfDeepARG · Off the shelfPLM-ARG · Off the shelfProtAlign-ARG · Off the shelfin the paper
+How results were checked
Held-out908 testedin the paper
final performance was evaluated on a strictly held-out Test Dataset (n = 908, 451 AMR-positive and 457 AMR-negative)where the paper describes this · verbatim
+Code · weights · data
code availableweights availabledata availablein the paper
All the scripts, datasets, and prediction models are made publicly available through GitHubwhere the paper describes this · verbatim
+Compute
High-performance computing environment with CPU and GPU nodes; no accelerator hours or run times givenin the paper

What this paper did not report

Technical · absence is published deliberately

Reported as not stated — 11 items
  • Version of ESM-2Which version of the model was used is not stated.
  • Version of Logistic RegressionWhich version of the model was used is not stated.
  • Version of Support Vector Machine (RBF kernel)Which version of the model was used is not stated.
  • Version of Random ForestWhich version of the model was used is not stated.
  • Version of Feed-forward Neural NetworkWhich version of the model was used is not stated.
  • Version of Fusion Model (stacked LR, SVM, RF, NN)Which version of the model was used is not stated.
  • Version of AlphaFold3Which version of the model was used is not stated.
  • Version of DeepARGWhich version of the model was used is not stated.
  • Version of PLM-ARGWhich version of the model was used is not stated.
  • Version of ProtAlign-ARGWhich version of the model was used is not stated.
  • What step 7 replacedThe paper gives no basis for what the AI stood in for.

About this article

Record aix-00112, version 2, checked by a person on 2026-10-08. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error