~/aixsci
200 records · all checked

structural-biology/ai produced the result/Nature Communications 2024 · v2

Neural network predicts protein pair binding strength from simplified molecular models

Researchers built MCGLPPI, a framework that turns protein complexes into coarse-grained graphs. Graph neural networks learned from these graphs to predict binding strength and to tell real biological interfaces from artefacts of crystal packing.

1. Curate downstream PPI benchmark datasets2. Curate 3DID domain-domain interaction pre-training set3. Generate MARTINI coarse-grained structures and force field parameters4. Build and crop CG-scale multi-relational complex graphs5. Self-supervised denoising pre-training of the CG graph encoder6. Train from scratch and fine-tune task models7. Predict complex binding affinity and interface type8. Evaluate accuracy and computational cost against baselines

spectrum · one line per step, placed by what the step does · bright lines used AI

Integration of molecular coarse-grained model into geometric representation learning framework for protein-protein complex property prediction
Nature Communications, 2024

doi:10.1038/s41467-024-53583-w · record aix-00179 v2 · checked 2026-10-09

ai-resultrole of AI
AI was for
Property prediction, Classification
Model family
Graph neural network, Diffusion model, Multilayer perceptron
Checked by
Benchmark161 tested
Code
available

The finding the paper is about came from the AI.

read as

The science is explained before the AI appears. Switch to field specialist to go straight to the method.

Assumes the discipline and goes straight to the method.

The diagram, the record and what the paper did not report are identical in both modes. Only the framing changes — never the evidence.

Introduction by AIxSci · plain language

What this research was about

Proteins rarely work alone. They latch onto one another, and how tightly they hold determines much of what happens inside a cell. That grip strength is usually written as a binding free energy, often shortened to delta G. Measuring it in the laboratory is slow, so there is long-standing interest in predicting it from a complex's three-dimensional shape instead. The difficulty is scale. A pair of proteins can contain tens of thousands of atoms, and tracking every one is costly. Separately, structures solved by crystallography contain contacts between protein chains that are real in the crystal but not in the living cell, and telling the two apart is its own problem.

One way to cut the cost is coarse-graining: lumping several atoms into a single bead, so the model keeps the overall shape while carrying far fewer points. The MARTINI force field is a widely used recipe for doing this, and it also supplies the bead types and the bonds, angles and twists between them. The authors set out to feed that coarse-grained description directly into a learning system, and to test the result on binding strength prediction and on interface classification.

Where AI came in

The coarse-grained beads became the nodes of a graph, with MARTINI bond types and nearby contacts as seven kinds of connecting edge, and the graph was cropped down to the region where the two proteins meet. A graph neural network, GearNet-Edge, read these graphs and condensed each complex into a single numerical summary. A three-layer network then turned that summary into either a predicted delta G value or a judgement about the interface type. The learned model stands in here for laboratory measurement of binding strength.

The encoder was also pre-trained without labels, in a denoising scheme: coordinates and sequences of domain-domain interaction graphs were deliberately scrambled with noise, and the network learned to undo it, before being fine-tuned on each task. Atom-scale and residue-scale versions of the same encoder, plus GVP-GNN, were trained as comparisons. The coarse-grained model used roughly five times and three times less GPU memory than those two, and ran three times faster than each.

Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.

The work

Technical · from the record

MCGLPPI is a graph neural network framework that represents protein-protein complexes at the MARTINI coarse-grained scale, mapping CG beads to graph nodes and MARTINI bonded parameters to edges and node features, and predicts overall complex properties from the cropped interaction region. The authors evaluated it with tenfold cross-validation on a curated 1270-sample PDBbind strict-dimer set and a 531-structure ATLAS TCR-pMHC set for binding affinity regression, and on the MANY/DC sets for biological-versus-crystal interface classification. Against atom- and residue-scale versions of the same GearNet-Edge encoder on the 915-sample PDBbind subset, the CG model reduced GPU consumption by approximately 5x and 3x and total elapsed time by 3x and 3x at the same batch size. Self-supervised denoising pre-training on 3DID domain-domain interaction structures raised Pearson correlation for MCGLPPI-M2 from 0.597 to 0.606 on PDBbind and 0.825 to 0.830 on ATLAS, while AUPR on MANY/DC fell from 0.880 to 0.866.

How AI was used

Atomistic complex structures were repaired with pdbfixer and converted to MARTINI22 (martinize.py) or MARTINI3 (Martinize2/Vermouth) coarse-grained structures and force field parameters. Each complex was then encoded as a multi-relational graph in which beads are nodes, MARTINI bond types plus intra- and inter-residue contact edges within 5 Angstroms form seven edge types, and bead types, angles and dihedrals are node and edge features; a backbone-distance rule cropped the graph to a core interaction region within 8.5 Angstroms plus adjacent residues within 10 Angstroms. A multi-relational heterogeneous GNN encoder (GearNet-Edge) operating on these CG graphs produced a graph-level representation that a three-layer MLP mapped to either a binding affinity value or an interface class. The encoder was additionally pre-trained in a self-supervised diffusion denoising scheme that adds noise to CG bead coordinates and sequences of 3DID domain-domain interaction graphs, then fine-tuned on each downstream task. Atom- and residue-scale GearNet-Edge models and GVP-GNN were trained by the authors under matched cropping and hyper-parameter settings as baselines, all with PyTorch and TorchDrug, Adam at learning rate 0.0001, and a single A100 GPU.

The shape of the work

Structural · the record, drawn

ACQUISITIONACQUISITIONPREPARATIONREPRESENTATIONTRAININGTRAININGINFERENCEVALIDATION12345678AIAIAICurate downstreamPPI benchmarkdatasetsCurate 3DIDdomain-domaininteraction pre-…Generate MARTINIcoarse-grainedstructures and f…Build and cropCG-scalemulti-relational…Self-superviseddenoisingpre-training of …Train fromscratch andfine-tune task m…Predict complexbinding affinityand interface ty…Evaluate accuracyand computationalcost against bas…↤ physical experiment
AI stepNo AI↤ what the AI stood in for
1Acquisition
no AI

Curate downstream PPI benchmark datasets

Obtaining raw data, whether by measurement, download or retrieval.

we obtained 1270 dimer samples with binding affinity labels △G, referred to as the PDBbind-strict-dimer datasetwhere the paper describes this · verbatim
in the paper
2Acquisition
no AI

Curate 3DID domain-domain interaction pre-training set

Obtaining raw data, whether by measurement, download or retrieval.

we obtained a pre-training dataset which provides 41,663 DDI structure samples in totalwhere the paper describes this · verbatim
in the paper
3Preparation
no AI

Generate MARTINI coarse-grained structures and force field parameters

Cleaning, filtering, normalising or labelling data already obtained.

was used to generate MARTINI22-based CG structure and force field parameters for each protein complexwhere the paper describes this · verbatim
in the paper
4Representation
no AI

Build and crop CG-scale multi-relational complex graphs

Encoding data into features, descriptors, embeddings or graphs.

First, an edge will be wired if any two bead nodes have the Euclidean distance smaller than 5Åwhere the paper describes this · verbatim
in the paper
5Training
AI

Self-supervised denoising pre-training of the CG graph encoder

Fitting model parameters, including fine-tuning an existing model.

a CG-scale complex pre-training technique was developed, which adds noise with changing magnitudes into 3D coordinates and sequences of MARTINI-based CG bead nodeswhere the paper describes this · verbatim
in the paper
6Training
AI

Train from scratch and fine-tune task models

Fitting model parameters, including fine-tuning an existing model.

we fine-tuned the CG graph encoder that had undergone pre-training for each respective downstream taskwhere the paper describes this · verbatim
in the paper
7Inference
AI

Predict complex binding affinity and interface type

Running a trained model over new data to predict, classify or score. The AI stood in for physical experiment.

the generated representation was further learnt by a three-layer task-specific multi-layer perception (MLP) to give the final property prediction resultswhere the paper describes this · verbatim
in the paper
8Validation
no AI

Evaluate accuracy and computational cost against baselines

Testing outputs against ground truth.

The standard tenfold cross-validation (CV) strategy was used to evaluate the modelwhere the paper describes this · verbatim
in the paper

What the record says

Technical · every part carries its own basis

+ in the paper~ our reading− not reported

How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.

~Role of AI
AI produced the resultour reading

The paper's result is the predictive framework itself and the property predictions it produces; every reported finding comes from training and running learned graph neural network models.

+What the AI was for
we introduce MCGLPPI, a geometric representation learning framework that combines graph neural networks (GNNs) with MARTINI molecular coarse-grained (CG) modelswhere the paper describes this · verbatim
+How it was taught
SupervisedSelf-supervisedTransfer / fine-tuningin the paper
+Models named
MCGLPPI-M2 (GearNet-Edge encoder on MARTINI22 CG graphs) · Trained from scratchMCGLPPI-M3 (GearNet-Edge encoder on MARTINI3 CG graphs) · Trained from scratchGearNet-Edge (atom-scale and residue-scale baselines) · Trained from scratchGVP-GNN · Trained from scratchin the paper
+How results were checked
Benchmark161 testedin the paper
the MANY and DC datasets were utilized, containing 5739 and 161 dimers respectivelywhere the paper describes this · verbatim
+Code · weights · data
code availableweights not reporteddata availablein the paper
The source code of MCGLPPI (Version 1.0) can be downloaded from https://github.com/arantir123/MCGLPPI.where the paper describes this · verbatim
+Compute
All experiments were run on one NVIDIA A100 GPU 40GB; reported GPU memory and total tenfold cross-validation run times include 11,560 MB and 14,312 s for MCGLPPI-M2 on the complete PDBbind-strict-dimer dataset. Computations used the Baskerville Tier 2 HPC service.in the paper

What this paper did not report

Technical · absence is published deliberately

Reported as not stated — 7 items
  • Trained model weightsWhether the trained model is available is not stated.
  • Version of MCGLPPI-M2 (GearNet-Edge encoder on MARTINI22 CG graphs)Which version of the model was used is not stated.
  • Version of MCGLPPI-M3 (GearNet-Edge encoder on MARTINI3 CG graphs)Which version of the model was used is not stated.
  • Version of GearNet-Edge (atom-scale and residue-scale baselines)Which version of the model was used is not stated.
  • Version of GVP-GNNWhich version of the model was used is not stated.
  • What step 5 replacedThe paper gives no basis for what the AI stood in for.
  • What step 6 replacedThe paper gives no basis for what the AI stood in for.

About this article

Record aix-00179, version 2, checked by a person on 2026-10-09. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error