structural-biology/ai produced the result/Biophysical Reviews 2024 · v2
Neural network reads protein surfaces to predict DNA and RNA binding
Researchers built PNAbind, a graph neural network trained on protein surface shapes and chemistry, to say whether a protein binds DNA or RNA and which of its residues do the binding.
spectrum · one line per step, placed by what the step does · bright lines used AI
Structure-based prediction of protein-nucleic acid binding using graph neural networks
Biophysical Reviews, 2024
doi:10.1007/s12551-024-01201-w · record aix-00116 v2 · checked 2026-10-08
- AI was for
- Classification, Segmentation
- Model family
- Graph neural network, Multilayer perceptron
- Checked by
- Benchmark
- Code
- not reported
The finding the paper is about came from the AI.
What this research was about
Proteins do much of their work by sticking to other molecules. A large group of them latch onto DNA or RNA, the long chains that carry genetic information, and in doing so switch genes on and off, copy them, cut them or carry them about. Knowing whether a given protein binds these nucleic acids, and exactly which parts of it make contact, matters for understanding what the protein does. Working that out in the laboratory is slow. Structures of a protein caught in the act of gripping DNA or RNA are harder to obtain than structures of the protein alone, so for most proteins the contact points are simply unknown.
The clues are in the shape. A binding site tends to sit on the protein's outer surface, in a pocket or groove of a particular curvature, lined with chemical groups whose electrical charge suits the strongly negative backbone of DNA or RNA. The authors set out to learn those surface signatures directly. They represented each protein as a mesh covering its outer surface, with every point on the mesh tagged with local geometry, chemistry, calculated electrical potential, and a record of how that part of the protein has varied across related species.
Where AI came in
The whole result is the output of two neural networks the authors trained from scratch. Both are graph neural networks: they treat the surface mesh as a web of connected points and pass information between neighbours, so each point is judged in the context of its surroundings. One network pools everything into a single verdict on the protein as a whole, DNA-binding, RNA-binding or neither. The other works point by point, assigning each patch of surface a probability of being part of a binding site; these are then combined into a score for each individual residue, the building blocks of the protein chain.
In place of the structures of proteins bound to DNA or RNA that would otherwise have to be solved experimentally, the networks were run on structures alone, including ones predicted by AlphaFold2, and on a set of proteins with no known nucleic acid binding role as a control. The authors also probed the trained classifiers, shuffling input features to see which mattered and using an attribution method to see which surface regions drove each verdict. Applied to the protein APOBEC3G, the models pointed to an RNA binding region present when two copies pair up but not in the single molecule.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record
The authors built PNAbind, a graph neural network that represents a protein as a molecular surface mesh carrying geometric, chemical, electrostatic and evolutionary features, and trained it both to classify a protein's overall nucleic acid binding function and to label individual residues as binding or non-binding. Models trained on AlphaFold2-predicted structures reached AUROC values of 0.943 for DNA-binding versus non-binding, 0.945 for RNA-binding versus non-binding, and 0.920 for DNA- versus RNA-binding. Binding site models evaluated on four published benchmark test sets gave AUROC values from 0.916 to 0.953, with the full biological assembly scoring higher than the monomer on these sets, and predictions on AlphaFold2 unbound structures correlated with those on native bound structures (Pearson correlation 0.76). Applied to APOBEC3G, the models predicted an RNA binding region spanning the CD1-CD1 dimerization interface in the dimer but not the monomer, plus a second dimerization-independent region.
How AI was used
Protein structures were converted to solvent-excluded molecular surface meshes with NanoShaper, with electrostatic potential from a boundary integral Poisson-Boltzmann solver, geometric descriptors such as curvature, circular variance and heat kernel signatures, chemical descriptors, and PSSM and profile HMM features from PSI-BLAST and HHblits mapped onto vertices; four rotation-invariant edge features described mesh geometry. Two graph neural network variants built from crystal-graph convolutions, farthest-point-sampling pooling and KNN unpooling were trained from scratch by minimising cross-entropy with ADAM under fivefold cross-validation with early stopping: a graph U-Net segmentation network producing per-vertex binding probabilities, max-pooled to residue level and ensembled with Platt scaling and F1-optimised thresholds, and a classification network with global attention pooling producing a single binding-function probability per structure. Classification models were trained on annotation-derived datasets of AlphaFold2 structures from Swiss-Prot, and segmentation models on benchmark datasets of experimentally determined protein-nucleic acid complexes. The trained models were then run on benchmark test sets, on AlphaFold2 unbound structures, on a negative control set of proteins without known nucleic acid binding function, and on APOBEC3G monomer and dimer structures. Feature permutation and Grad-CAM attribution were applied to the trained classification models to score feature groups and surface regions.
The shape of the work
Structural · the record, drawn
no AI
Assemble annotation-based and benchmark datasets
Obtaining raw data, whether by measurement, download or retrieval.
Three sets of proteins were identified in the Swiss-Prot knowledgebase (Bateman et al. ) using functional annotationswhere the paper describes this · verbatim
no AI
Filter, quality-screen and cluster protein sets
Cleaning, filtering, normalising or labelling data already obtained.
Low-confidence regions of the predicted structures were removed (confidence < 0.65), and only structures which remained as non-disjoint structure were kept.where the paper describes this · verbatim
no AI
Build surface meshes and map vertex/edge features
Encoding data into features, descriptors, embeddings or graphs.
We used NanoShaper (Decherchi and Rocchia ) for generating the molecular surface meshwhere the paper describes this · verbatim
AI
Train GNN classification and segmentation models
Fitting model parameters, including fine-tuning an existing model.
Model parameters were determined using the ADAM optimizer (Kingma and Ba ) to minimize the cross-entropy between predicted probabilities and the ground-truth labelswhere the paper describes this · verbatim
AI
Predict overall DNA/RNA binding function
Running a trained model over new data to predict, classify or score. The AI stood in for expert judgement.
we report on DNA and RNA binding function prediction using models trained on AlphaFold2-predicted protein structureswhere the paper describes this · verbatim
AI
Predict residue-level NA binding sites
Running a trained model over new data to predict, classify or score. The AI stood in for physical experiment.
Residue-level probabilities are obtained by max pooling over all vertices that correspond to the solvent-excluded surface of each residue.where the paper describes this · verbatim
AI
Attribute predictions to features and surface regions
Extracting understanding from model behaviour.
We computed the spatial attribution via Grad-CAM (Selvaraju et al. ), an attribution method where gradients for a target class probability are computedwhere the paper describes this · verbatim
no AI
Evaluate against benchmarks, negative control and published experiments
Testing outputs against ground truth.
Trained models were validated on four test sets, two containing native DBP, and two native RBPwhere the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
The paper's results are the outputs of the authors' trained GNN models: binding function classifications, residue-level binding site predictions, and attribution-based interpretation.
We present PNAbind, a GNN-based method for predicting DNA and RNA binding function and binding sites from protein structure.where the paper describes this · verbatim
PNAbind achieves the highest AUROC value (ranging from 0.916 to 0.953), and the highest Matthews correlation coefficient (MCC)where the paper describes this · verbatim
median wall time to evaluate our segmentation model (inference) is less than one second per protein surface meshwhere the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- CodeWhether the code is available is not stated.
- Trained model weightsWhether the trained model is available is not stated.
- DataWhether the data are available is not stated.
- How many were testedThe paper gives no count of what was tested.
- Version of PNAbind classification model (binding function prediction)Which version of the model was used is not stated.
- Version of PNAbind segmentation model (binding site prediction)Which version of the model was used is not stated.
- What step 4 replacedThe paper gives no basis for what the AI stood in for.
- What step 7 replacedThe paper gives no basis for what the AI stood in for.
About this article
Record aix-00116, version 2, checked by a person on 2026-10-08. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error