~/aixsci
200 records · all checked

structural-biology/ai produced the result/Nucleic Acids Research 2024 · v2

Deep learning tool predicts how single mutations weaken protein partnerships

Researchers built DDMut-PPI, a neural network that estimates how much a change to one amino acid alters the strength with which two proteins stick together. The model was trained on measured binding energies and tested on mutations it had not seen.

1. Assemble mutation and structure datasets2. Augment training data with hypothetical reverse mutations3. Build mutant complex structures4. Encode features and interface graph5. Train Siamese graph convolutional network6. Tune hyperparameters under protein-level cross-validation7. Predict mutation effects on held-out and multiple-mutation sets8. Evaluate against measurements, baselines and ablations

spectrum · one line per step, placed by what the step does · bright lines used AI

DDMut-PPI: predicting effects of mutations on protein–protein interactions using graph-based deep learning
Nucleic Acids Research, 2024

doi:10.1093/nar/gkae412 · record aix-00056 v2 · checked 2026-10-08

ai-resultrole of AI
AI was for
Property prediction
Model family
Graph neural network, Protein language model, Transformer, Convolutional neural network, Multilayer perceptron
Checked by
Benchmark645 tested
Code
not reported

The finding the paper is about came from the AI.

read as

The science is explained before the AI appears. Switch to field specialist to go straight to the method.

Assumes the discipline and goes straight to the method.

The diagram, the record and what the paper did not report are identical in both modes. Only the framing changes — never the evidence.

Introduction by AIxSci · plain language

What this research was about

Proteins rarely work alone. Much of what happens inside a cell depends on two proteins gripping each other at a shared surface, called an interface. How tightly they grip can be described by a quantity known as binding free energy. Swap a single amino acid, one link in a protein's chain, at or near that interface, and the grip may tighten, loosen, or fail altogether. This matters for understanding disease-causing variants and for designing antibodies, which work by latching onto a target. Measuring the effect of each possible swap in the laboratory is slow, so only a small fraction of conceivable mutations has ever been tested.

The researchers set out to predict that change in binding energy computationally, for single swaps and for combinations of several at once.

Where AI came in

The prediction itself is the AI. The team assembled mutations with laboratory-measured binding energy changes, took the matching three-dimensional structures of the protein pairs from the Protein Data Bank, and modelled the mutant structures with existing software. They then described each interface as a network of its contacting amino acids, with the chemical contacts between them as the links. Each amino acid carried a numerical fingerprint produced by ProtT5, a language model trained on protein sequences, which reads a sequence much as a text model reads words.

A neural network was trained on these descriptions to output the predicted energy change. It processed the original and mutated forms side by side, and was taught that reversing a mutation should flip the answer's sign. Effects of multiple mutations were estimated by adding up the single-mutation predictions. The trained model was then run on held-out mutations, including 645 in antibody-antigen pairs, and its outputs compared with the measurements and with other predictors. Here the model stands in for the laboratory experiment that would otherwise supply each number.

Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.

The work

Technical · from the record

DDMut-PPI is a deep learning model that predicts the change in protein–protein binding free energy caused by single or multiple point mutations. It uses a Siamese network in which each sub-network contains a graph convolutional network over the protein–protein interface graph, with ProtT5 language model embeddings as node features and Arpeggio-derived interaction types as edge features. The authors report a Pearson correlation of 0.75 and an RMSE of 1.33 kcal/mol under leave-one-complex-out cross-validation on the S8338 training set, and 0.83 with an RMSE of 1.51 kcal/mol on the SM1124 double- and triple-mutation set. The model is distributed as a web server and API.

How AI was used

Mutation records with measured binding free energy changes were collected from SKEMPI 2.0 and related sources, wild-type structures were taken from the Protein Data Bank, and mutant structures were modelled with MODELLER. The single-mutation training set was augmented with hypothetical reverse mutations so that each forward mutation was paired with its reverse. Sequence features from substitution matrices and PSI-BLAST position-specific scoring matrices, structural features from Biopython and FoldX, atomic interaction changes from Arpeggio, mCSM graph-based signatures and python-igraph network metrics were computed and normalised. The interface was additionally encoded as a graph whose nodes are residues within 5.0 Å of the opposing chain, carrying 1024-dimensional ProtT5 embeddings, and whose edges are ten Arpeggio interaction types. A Siamese network in TensorFlow, with a graph convolutional network per sub-network, a convolutional layer and transformer encoder for the graph signatures, and dense layers for other features, was trained with a modified contrastive loss enforcing anti-symmetry between forward and reverse mutations; hyperparameters and layers were tuned under leave-one-binding-site-out cross-validation. The trained model was then run on blind test sets, with multiple-mutation effects obtained by summing predictions for the constituent single mutations, and assessed by correlation and error metrics, component ablation and feature shuffling.

The shape of the work

Structural · the record, drawn

ACQUISITIONPREPARATIONPREPARATIONREPRESENTATIONTRAININGOPTIMISATIONINFERENCEVALIDATION12345678AIAIAIAIAIAssemble mutationand structuredatasetsAugment trainingdata withhypothetical rev…Build mutantcomplexstructuresEncode featuresand interfacegraphTrain Siamesegraphconvolutional ne…Tunehyperparametersunder protein-le…Predict mutationeffects onheld-out and mul…Evaluate againstmeasurements,baselines and ab…↤ expert judgement↤ physical experimentloops back
AI stepNo AI↤ what the AI stood in for
1Acquisition
no AI

Assemble mutation and structure datasets

Obtaining raw data, whether by measurement, download or retrieval.

our study utilized the S4169 dataset, derived from SKEMPI 2.0, for model trainingwhere the paper describes this · verbatim
in the paper
2Preparation
no AI

Augment training data with hypothetical reverse mutations

Cleaning, filtering, normalising or labelling data already obtained.

This adjustment expanded the original S4169 dataset by incorporating an additional 4169 hypothetical reverse mutationswhere the paper describes this · verbatim
in the paper
3Preparation
no AI

Build mutant complex structures

Cleaning, filtering, normalising or labelling data already obtained.

corresponding mutant structures were generated from these wild-type structures using the MODELLER software (version 10.4)where the paper describes this · verbatim
in the paper
4Representation
AI

Encode features and interface graph

Encoding data into features, descriptors, embeddings or graphs.

Each of these residues was represented as a graph node, enriched with ProtT5 embeddings from the protein language modelwhere the paper describes this · verbatim
in the paper
5Training
AI

Train Siamese graph convolutional network

Fitting model parameters, including fine-tuning an existing model.

we incorporated a graph convolutional network (GCN) within each sub-network to analyse the PPI interface graphwhere the paper describes this · verbatim
in the paper
6Optimisation
AI

Tune hyperparameters under protein-level cross-validation

Iterative search over a space. The AI stood in for expert judgement. Its result feeds back into an earlier step.

the model’s hyperparameters were optimized to achieve optimal performance, guided by a leave-one-binding-site-out cross-validation strategy on the training datasetwhere the paper describes this · verbatim
in the paper
7Inference
AI

Predict mutation effects on held-out and multiple-mutation sets

Running a trained model over new data to predict, classify or score. The AI stood in for physical experiment.

DDMut-PPI employs an additive approach, aggregating the effects of individual single point mutations to predict the overall outcomewhere the paper describes this · verbatim
in the paper
8Validation
AI

Evaluate against measurements, baselines and ablations

Testing outputs against ground truth.

DDMut-PPI achieved competitive performance against the top methods across these diverse datasetswhere the paper describes this · verbatim
in the paper

What the record says

Technical · every part carries its own basis

+ in the paper~ our reading− not reported

How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.

~Role of AI
AI produced the resultour reading

The paper's result is the predictor itself: a trained deep learning model whose predicted binding free energy changes are the reported output

+What the AI was for
the DDMut-PPI model was enhanced with a graph convolutional network operated on the protein interaction interfacewhere the paper describes this · verbatim
+How it was taught
Supervisedin the paper
+Models named
DDMut-PPI · Trained from scratchProtT5 · Off the shelfmCSM-PPI2 · Off the shelfDGCddG · Off the shelfin the paper
+How results were checked
Benchmark645 testedin the paper
DDMut-PPI was then evaluated on three different blind test sets, including 645 mutations in antibody–antigen complexeswhere the paper describes this · verbatim
−Code · weights · data
code not reportedweights not reporteddata not reportednot reported
−Compute
not reportednot reported

What this paper did not report

Technical · absence is published deliberately

Reported as not stated — 11 items
  • CodeWhether the code is available is not stated.
  • Trained model weightsWhether the trained model is available is not stated.
  • DataWhether the data are available is not stated.
  • ComputeThe hardware or time used is not stated.
  • Version of DDMut-PPIWhich version of the model was used is not stated.
  • Version of ProtT5Which version of the model was used is not stated.
  • Version of mCSM-PPI2Which version of the model was used is not stated.
  • Version of DGCddGWhich version of the model was used is not stated.
  • What step 4 replacedThe paper gives no basis for what the AI stood in for.
  • What step 5 replacedThe paper gives no basis for what the AI stood in for.
  • What step 8 replacedThe paper gives no basis for what the AI stood in for.

About this article

Record aix-00056, version 2, checked by a person on 2026-10-08. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error