~/aixsci
200 records · all checked

structural-biology/ai produced the result/International Journal of Molecular Sciences 2026 · v2

Protein language model and two classifiers predict where ATP binds on proteins

Researchers built a tool that reads a protein's amino acid sequence and marks which positions bind ATP. A pre-trained protein language model encoded each position, and two learned classifiers, combined by weighted sum, made the calls.

1. Assemble benchmark residue datasets2. Encode residues with ESM-2 embeddings3. Train local–global dual-attention Transformer classifier4. Train DQN reinforcement-learning classifier5. Select window size and ensemble weights by cross-validation6. Predict binding probability per residue on test sets7. Evaluate against published predictors and a case-study protein8. Profile residue and protein-family prediction patterns

spectrum · one line per step, placed by what the step does · bright lines used AI

A Novel Weighted Ensemble Framework of Transformer and Deep Q-Network for ATP-Binding Site Prediction Using Protein Language Model Features
International Journal of Molecular Sciences, 2026

doi:10.3390/ijms27073097 · record aix-00156 v2 · checked 2026-10-09

ai-resultrole of AI
AI was for
Classification
Model family
Transformer, Protein language model, Multilayer perceptron
Checked by
Benchmark
Code
available

The finding the paper is about came from the AI.

read as

The science is explained before the AI appears. Switch to field specialist to go straight to the method.

Assumes the discipline and goes straight to the method.

The diagram, the record and what the paper did not report are identical in both modes. Only the framing changes — never the evidence.

Introduction by AIxSci · plain language

What this research was about

Proteins are long chains of amino acids, often called residues, that fold into shapes. Many of them do their job only by gripping a small partner molecule. One of the commonest partners is ATP, the molecule cells use to carry energy. ATP usually sits in a pocket formed by a handful of residues scattered along the chain, so neighbours in the folded pocket can be far apart in the sequence. Working out which residues form that pocket normally means solving the protein's structure in the laboratory, which is slow and does not work for every protein. Predicting the pocket from sequence alone is therefore a long-standing goal, and a hard one: binding residues are heavily outnumbered by non-binding ones.

The researchers set out to build a predictor that takes only the sequence and labels each residue as binding or not. They trained and tested it on four published benchmark sets of protein chains whose ATP-binding residues had been annotated experimentally, compared it with other sequence-based methods, and then looked at which amino acids and which protein families it handled well or badly.

Where AI came in

Every stage of the prediction is machine learning. First, a pre-trained protein language model called ESM-2 read each sequence and turned every residue into a list of 1,280 numbers. Such models are trained on large collections of sequences without being told anything about binding; they learn patterns of amino acid usage, and the numbers they emit stand in for the handcrafted sequence descriptors earlier methods had to design by hand.

Two classifiers were then trained on those numbers. One was a transformer, a network that weighs how much each position should attend to others; here it combined attention within a fixed local window with attention across the whole sequence. The other was a deep Q-network, borrowed from reinforcement learning, which treats each residue decision as an action and learns from rewards. Five-fold cross-validation fixed the window size at 32 and the blend at 0.6 transformer to 0.4 deep Q-network. Their combined scores replace laboratory measurement of where ATP sits.

Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.

The work

Technical · from the record

The study builds a sequence-only predictor of ATP-binding residues in proteins. Each residue is encoded as a 1280-dimensional vector by the pre-trained ESM-2 protein language model, and two classifiers trained on those vectors — a Transformer with combined local-window and global self-attention plus a contrastive learning branch, and a deep Q-network that treats each residue call as a reinforcement-learning action — are combined by a weighted sum whose weights were set by five-fold cross-validation to 0.6 and 0.4. On the ATP-41 independent testing set the authors report an AUC of 0.9188 and an MCC of 0.6564, and on ATP-17 an AUC of 0.9353 and an MCC of 0.6625, which they state are higher than the sequence-based methods they compared against on both sets. Residue-level and Pfam-family breakdowns report the highest recognition rates for glycine and lysine and the lowest for cysteine and proline, and the highest family-level MCC for AAA+ ATPases.

How AI was used

Protein chains with annotated ATP-binding residues were taken from four published benchmark sets (ATP-388/ATP-41 and ATP-227/ATP-17). The pre-trained ESM-2 model esm2_t33_650M_UR50D was run over each FASTA sequence to produce a 1280-dimensional embedding per residue, replacing handcrafted sequence features. Two classifiers were then fitted on these embeddings: a Transformer with learnable position embeddings, an eight-head global self-attention module, a local attention module over windows of fixed size, a learnable weight fusing the two attention outputs, a contrastive projection head with temperature-scaled loss, and a hybrid Focal Loss plus MCC-penalty objective optimised with Adam and cosine annealing; and a deep Q-network whose feedforward policy and target networks (512 and 256 hidden units) map a residue state to Q-values for the binary binding/non-binding action, trained with experience replay, an epsilon-greedy policy decaying from 1.0 to 0.01, discount factor 0.95 and target updates every ten episodes. Five-fold cross-validation on the training sets selected the local window size among six candidates and the ensemble weight split; the two classifiers' probability scores were then weighted and summed per residue for the independent testing sets, with the decision threshold chosen by maximising MCC. Ablation variants (traditional, local-only and global-only Transformers, and models with and without the contrastive branch) were trained for comparison, and t-SNE was used to visualise the learned features.

The shape of the work

Structural · the record, drawn

ACQUISITIONREPRESENTATIONTRAININGTRAININGOPTIMISATIONINFERENCEVALIDATIONINTERPRETATION12345678AIAIAIAIAssemblebenchmark residuedatasetsEncode residueswith ESM-2embeddingsTrainlocal–globaldual-attention T…Train DQNreinforcement-learningclassifierSelect windowsize and ensembleweights by cross…Predict bindingprobability perresidue on test …Evaluate againstpublishedpredictors and a…Profile residueandprotein-family p…↤ conventional algorithm↤ physical experimentloops back
AI stepNo AI↤ what the AI stood in for
1Acquisition
no AI

Assemble benchmark residue datasets

Obtaining raw data, whether by measurement, download or retrieval.

ATP-388 includes 388 protein chains deposited before 5 November 2014, containing 5657 binding residues and 142,086 non-binding residueswhere the paper describes this · verbatim
in the paper
2Representation
AI

Encode residues with ESM-2 embeddings

Encoding data into features, descriptors, embeddings or graphs. The AI stood in for conventional algorithm.

the input protein sequence (in FASTA format) is processed by the ESM-2 protein language modelwhere the paper describes this · verbatim
in the paper
3Training
AI

Train local–global dual-attention Transformer classifier

Fitting model parameters, including fine-tuning an existing model.

the model is trained using the Adam optimizer with a cosine annealing learning rate schedulewhere the paper describes this · verbatim
in the paper
4Training
AI

Train DQN reinforcement-learning classifier

Fitting model parameters, including fine-tuning an existing model.

we parallelly constructed a reinforcement learning-based DQN model for ATP-binding site predictionwhere the paper describes this · verbatim
in the paper
5Optimisation
no AI

Select window size and ensemble weights by cross-validation

Iterative search over a space. Its result feeds back into an earlier step.

the optimal weight combination converged to 0.6 for the Transformer classifier and 0.4 for the DQN classifier on both datasetswhere the paper describes this · verbatim
in the paper
6Inference
AI

Predict binding probability per residue on test sets

Running a trained model over new data to predict, classify or score. The AI stood in for physical experiment.

Each model independently outputs a probability score for each residue, which indicates the likelihood of it being an ATP-binding sitewhere the paper describes this · verbatim
in the paper
7Validation
no AI

Evaluate against published predictors and a case-study protein

Testing outputs against ground truth.

we compared them with some state-of-the-art sequence-based prediction methods on two independent testing sets, ATP-41 and ATP-17where the paper describes this · verbatim
in the paper
8Interpretation
no AI

Profile residue and protein-family prediction patterns

Extracting understanding from model behaviour.

we calculated the recognition rate for each amino acid in true ATP-binding siteswhere the paper describes this · verbatim
in the paper

What the record says

Technical · every part carries its own basis

+ in the paper~ our reading− not reported

How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.

~Role of AI
AI produced the resultour reading

The paper's result is the predictor itself; all reported binding-site predictions are produced by the trained models

+What the AI was for
Classificationin the paper
the ESM-2 protein language model is used to extract deep features rich in biological semantics from the original sequenceswhere the paper describes this · verbatim
+How it was taught
SupervisedSelf-supervisedReinforcementin the paper
+Models named
ESM-2 esm2_t33_650M_UR50D · Off the shelfLocal–Global dual-attention Transformer classifier with contrastive learning branch · Trained from scratchDQN classifier (feedforward policy and target networks) · Trained from scratchTraditional Transformer (ablation baseline) · Trained from scratchLocal-only Transformer (ablation variant) · Trained from scratchGlobal-only Transformer (ablation variant) · Trained from scratchZhang's Method (adapted baseline) · Trained from scratchin the paper
+How results were checked
Benchmarkin the paper
we compared them with some state-of-the-art sequence-based prediction methods on two independent testing sets, ATP-41 and ATP-17where the paper describes this · verbatim
+Code · weights · data
code availableweights not reporteddata availablein the paper
The source code and datasets of proposed method is available at https://github.com/tlsjz/ATPbindingEnsemblewhere the paper describes this · verbatim
+Compute
not reportedin the paper

What this paper did not report

Technical · absence is published deliberately

Reported as not stated — 11 items
  • Trained model weightsWhether the trained model is available is not stated.
  • ComputeThe hardware or time used is not stated.
  • How many were testedThe paper gives no count of what was tested.
  • Version of Local–Global dual-attention Transformer classifier with contrastive learning branchWhich version of the model was used is not stated.
  • Version of DQN classifier (feedforward policy and target networks)Which version of the model was used is not stated.
  • Version of Traditional Transformer (ablation baseline)Which version of the model was used is not stated.
  • Version of Local-only Transformer (ablation variant)Which version of the model was used is not stated.
  • Version of Global-only Transformer (ablation variant)Which version of the model was used is not stated.
  • Version of Zhang's Method (adapted baseline)Which version of the model was used is not stated.
  • What step 3 replacedThe paper gives no basis for what the AI stood in for.
  • What step 4 replacedThe paper gives no basis for what the AI stood in for.

About this article

Record aix-00156, version 2, checked by a person on 2026-10-09. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error