structural-biology/ai produced the result/International Journal of Molecular Sciences 2026 · v2
Protein language model and two classifiers predict where ATP binds on proteins
Researchers built a tool that reads a protein's amino acid sequence and marks which positions bind ATP. A pre-trained protein language model encoded each position, and two learned classifiers, combined by weighted sum, made the calls.
spectrum · one line per step, placed by what the step does · bright lines used AI
A Novel Weighted Ensemble Framework of Transformer and Deep Q-Network for ATP-Binding Site Prediction Using Protein Language Model Features
International Journal of Molecular Sciences, 2026
doi:10.3390/ijms27073097 · record aix-00156 v2 · checked 2026-10-09
- AI was for
- Classification
- Model family
- Transformer, Protein language model, Multilayer perceptron
- Checked by
- Benchmark
- Code
- available
The finding the paper is about came from the AI.
What this research was about
Proteins are long chains of amino acids, often called residues, that fold into shapes. Many of them do their job only by gripping a small partner molecule. One of the commonest partners is ATP, the molecule cells use to carry energy. ATP usually sits in a pocket formed by a handful of residues scattered along the chain, so neighbours in the folded pocket can be far apart in the sequence. Working out which residues form that pocket normally means solving the protein's structure in the laboratory, which is slow and does not work for every protein. Predicting the pocket from sequence alone is therefore a long-standing goal, and a hard one: binding residues are heavily outnumbered by non-binding ones.
The researchers set out to build a predictor that takes only the sequence and labels each residue as binding or not. They trained and tested it on four published benchmark sets of protein chains whose ATP-binding residues had been annotated experimentally, compared it with other sequence-based methods, and then looked at which amino acids and which protein families it handled well or badly.
Where AI came in
Every stage of the prediction is machine learning. First, a pre-trained protein language model called ESM-2 read each sequence and turned every residue into a list of 1,280 numbers. Such models are trained on large collections of sequences without being told anything about binding; they learn patterns of amino acid usage, and the numbers they emit stand in for the handcrafted sequence descriptors earlier methods had to design by hand.
Two classifiers were then trained on those numbers. One was a transformer, a network that weighs how much each position should attend to others; here it combined attention within a fixed local window with attention across the whole sequence. The other was a deep Q-network, borrowed from reinforcement learning, which treats each residue decision as an action and learns from rewards. Five-fold cross-validation fixed the window size at 32 and the blend at 0.6 transformer to 0.4 deep Q-network. Their combined scores replace laboratory measurement of where ATP sits.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record
The study builds a sequence-only predictor of ATP-binding residues in proteins. Each residue is encoded as a 1280-dimensional vector by the pre-trained ESM-2 protein language model, and two classifiers trained on those vectors — a Transformer with combined local-window and global self-attention plus a contrastive learning branch, and a deep Q-network that treats each residue call as a reinforcement-learning action — are combined by a weighted sum whose weights were set by five-fold cross-validation to 0.6 and 0.4. On the ATP-41 independent testing set the authors report an AUC of 0.9188 and an MCC of 0.6564, and on ATP-17 an AUC of 0.9353 and an MCC of 0.6625, which they state are higher than the sequence-based methods they compared against on both sets. Residue-level and Pfam-family breakdowns report the highest recognition rates for glycine and lysine and the lowest for cysteine and proline, and the highest family-level MCC for AAA+ ATPases.
How AI was used
Protein chains with annotated ATP-binding residues were taken from four published benchmark sets (ATP-388/ATP-41 and ATP-227/ATP-17). The pre-trained ESM-2 model esm2_t33_650M_UR50D was run over each FASTA sequence to produce a 1280-dimensional embedding per residue, replacing handcrafted sequence features. Two classifiers were then fitted on these embeddings: a Transformer with learnable position embeddings, an eight-head global self-attention module, a local attention module over windows of fixed size, a learnable weight fusing the two attention outputs, a contrastive projection head with temperature-scaled loss, and a hybrid Focal Loss plus MCC-penalty objective optimised with Adam and cosine annealing; and a deep Q-network whose feedforward policy and target networks (512 and 256 hidden units) map a residue state to Q-values for the binary binding/non-binding action, trained with experience replay, an epsilon-greedy policy decaying from 1.0 to 0.01, discount factor 0.95 and target updates every ten episodes. Five-fold cross-validation on the training sets selected the local window size among six candidates and the ensemble weight split; the two classifiers' probability scores were then weighted and summed per residue for the independent testing sets, with the decision threshold chosen by maximising MCC. Ablation variants (traditional, local-only and global-only Transformers, and models with and without the contrastive branch) were trained for comparison, and t-SNE was used to visualise the learned features.
The shape of the work
Structural · the record, drawn
no AI
Assemble benchmark residue datasets
Obtaining raw data, whether by measurement, download or retrieval.
ATP-388 includes 388 protein chains deposited before 5 November 2014, containing 5657 binding residues and 142,086 non-binding residueswhere the paper describes this · verbatim
AI
Encode residues with ESM-2 embeddings
Encoding data into features, descriptors, embeddings or graphs. The AI stood in for conventional algorithm.
the input protein sequence (in FASTA format) is processed by the ESM-2 protein language modelwhere the paper describes this · verbatim
AI
Train local–global dual-attention Transformer classifier
Fitting model parameters, including fine-tuning an existing model.
the model is trained using the Adam optimizer with a cosine annealing learning rate schedulewhere the paper describes this · verbatim
AI
Train DQN reinforcement-learning classifier
Fitting model parameters, including fine-tuning an existing model.
we parallelly constructed a reinforcement learning-based DQN model for ATP-binding site predictionwhere the paper describes this · verbatim
no AI
Select window size and ensemble weights by cross-validation
Iterative search over a space. Its result feeds back into an earlier step.
the optimal weight combination converged to 0.6 for the Transformer classifier and 0.4 for the DQN classifier on both datasetswhere the paper describes this · verbatim
AI
Predict binding probability per residue on test sets
Running a trained model over new data to predict, classify or score. The AI stood in for physical experiment.
Each model independently outputs a probability score for each residue, which indicates the likelihood of it being an ATP-binding sitewhere the paper describes this · verbatim
no AI
Evaluate against published predictors and a case-study protein
Testing outputs against ground truth.
we compared them with some state-of-the-art sequence-based prediction methods on two independent testing sets, ATP-41 and ATP-17where the paper describes this · verbatim
no AI
Profile residue and protein-family prediction patterns
Extracting understanding from model behaviour.
we calculated the recognition rate for each amino acid in true ATP-binding siteswhere the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
The paper's result is the predictor itself; all reported binding-site predictions are produced by the trained models
the ESM-2 protein language model is used to extract deep features rich in biological semantics from the original sequenceswhere the paper describes this · verbatim
we compared them with some state-of-the-art sequence-based prediction methods on two independent testing sets, ATP-41 and ATP-17where the paper describes this · verbatim
The source code and datasets of proposed method is available at https://github.com/tlsjz/ATPbindingEnsemblewhere the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- Trained model weightsWhether the trained model is available is not stated.
- ComputeThe hardware or time used is not stated.
- How many were testedThe paper gives no count of what was tested.
- Version of Local–Global dual-attention Transformer classifier with contrastive learning branchWhich version of the model was used is not stated.
- Version of DQN classifier (feedforward policy and target networks)Which version of the model was used is not stated.
- Version of Traditional Transformer (ablation baseline)Which version of the model was used is not stated.
- Version of Local-only Transformer (ablation variant)Which version of the model was used is not stated.
- Version of Global-only Transformer (ablation variant)Which version of the model was used is not stated.
- Version of Zhang's Method (adapted baseline)Which version of the model was used is not stated.
- What step 3 replacedThe paper gives no basis for what the AI stood in for.
- What step 4 replacedThe paper gives no basis for what the AI stood in for.
About this article
Record aix-00156, version 2, checked by a person on 2026-10-09. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error