~/aixsci
200 records · all checked

structural-biology/ai in a supporting role/Nature 2023 · v2

Researchers measure folding stability for hundreds of thousands of protein variants

A new laboratory assay produced around 776,000 measurements of how firmly protein domains hold their shape. Learned and fitted models chose the domains, generated some of the test sequences, turned sequencing counts into stability numbers and helped interpret the results.

1. Curate small monomeric domains from the PDB2. Model domain structures and trim flexible termini3. Design de novo and redesigned sequences4. cDNA display proteolysis assay and deep sequencing5. Infer K50 and folding stability, then filter for quality6. Principal component analysis of per-site amino acid stabilities7. Fit classifier predicting wild-type amino acid from stabilities8. Identify functional sites using evolutionary sensitivity scores

spectrum · one line per step, placed by what the step does · bright lines used AI

Mega-scale experimental analysis of protein folding stability in biology and design
Nature, 2023

doi:10.1038/s41586-023-06328-6 · record aix-00001 v2 · checked 2026-10-07

ai-supportingrole of AI
AI was for
Structure determination, Candidate generation, Property prediction, Classification
Model family
Transformer, Convolutional neural network, Linear model, Clustering
Checked by
Replication1188 tested
Code
available

AI processed or interpreted data, but the main finding does not rest on it.

read as

The science is explained before the AI appears. Switch to field specialist to go straight to the method.

Assumes the discipline and goes straight to the method.

The diagram, the record and what the paper did not report are identical in both modes. Only the framing changes — never the evidence.

Introduction by AIxSci · plain language

What this research was about

A protein is a chain of amino acids that folds into a particular shape, and the shape is what lets it work. Folding stability is a measure of how strongly the folded form is favoured over the loose, unfolded chain. Change a single amino acid and the stability may barely shift, or the protein may stop holding its shape at all. Measuring this has traditionally meant purifying one protein at a time, so the catalogue of measurements has stayed small relative to the number of possible variants. That makes it hard to say which general rules govern stability, and how evolution has worked within them.

The researchers built a method they call cDNA display proteolysis, which tests many sequences at once by exposing them to protein-cutting enzymes and reading out which survive by DNA sequencing. They applied it to all single amino acid changes, plus selected pairs of changes, across 331 natural protein domains and 148 designed ones, and used the resulting dataset to look at what stability depends on at each site.

Where AI came in

Computation entered at several points without making the measurements themselves. AlphaFold, a model that predicts a protein's three-dimensional shape from its sequence, was run on candidate domains so that floppy, loosely attached ends could be trimmed before the assay; the same predicted structures later supplied information on how buried each site is and what local shape it sits in. Some of the designed sequences were generated by a trRosetta hallucination protocol, which invents new backbones and sequences rather than copying existing ones. Other designs came from Rosetta and from PROSS redesign of natural domains.

Turning raw sequencing counts into stability values also relied on fitted models: one describing the enzyme cutting kinetics, and one, trained on 64,238 scrambled sequences, predicting how readily an unfolded chain of a given sequence is cut. That second model stood in for a separate measurement of the unfolded reference state for every sequence. Afterwards, principal component analysis summarised the per-site patterns, a fitted classifier tried to recover which amino acid nature had chosen at a site from the stabilities of its variants, and GEMME evolutionary sensitivity scores were compared with measured stability effects to flag sites that look functional rather than structural.

Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.

The work

Technical · from the record

The authors developed cDNA display proteolysis, a high-throughput proteolysis assay read out by deep sequencing, and curated around 776,000 folding stability measurements covering all single amino acid variants and selected double mutants of 331 natural and 148 designed protein domains. Computation entered at several points: AlphaFold models were used to trim flexible termini when choosing domains and for later structural analysis, trRosetta hallucination generated part of the designed set, and a Bayesian kinetic model with a fitted unfolded-state scoring matrix converted sequencing counts into absolute stabilities. Fitted statistical models were then used to analyse the dataset, including a principal component analysis of per-site amino acid stabilities, a classifier that predicts the wild-type amino acid at a site from the stabilities of its variants, and a comparison of stability effects with GEMME evolutionary sensitivity scores to flag functional sites. The measurements agreed with published purified-protein data for 1,188 variants of 10 proteins, and the dataset was also used to characterise thermodynamic couplings between residue pairs and to evaluate the design tool PROSS.

How AI was used

AlphaFold was run with default parameters, keeping the highest-pLDDT model, to predict structures for candidate PDB sequences so that low-contact terminal segments could be trimmed before assay design, and the same models supplied the burial, contact-count and secondary-structure features used in later analyses. Part of the designed domain set was produced by a trRosetta hallucination protocol that generates backbones and sequences by maximising the Kullback-Leibler divergence between predicted and background distance and angle distributions, with other designs built by Rosetta blueprint-based design and by PROSS redesign of wild-type domains. Sequencing counts from the proteolysis assay were converted into stabilities by Bayesian inference in Numpyro using two models: a K50 model fitted to the count data under single-turnover kinetics with a universal maximum cleavage rate, and an unfolded-state model, a position-specific scoring matrix parameterised on 64,238 scrambled sequences, that predicts each sequence's unfolded-state K50. Downstream, principal component analysis implemented in scikit-learn was fitted to the per-site stabilities of the 20 amino acids, a classifier with a shared monotonic weighting function and amino acid-specific offsets was fitted to predict wild-type amino acids from variant stabilities and evaluated on a held-out set of domains with no similarity to the training set, and GEMME was run on each natural sequence with default parameters over Jackhmmer alignments to score per-site evolutionary sensitivity for comparison with measured stability effects.

The shape of the work

Structural · the record, drawn

ACQUISITIONINFERENCEGENERATIONEXPERIMENTPREPARATIONINTERPRETATIONINTERPRETATIONINFERENCE12345678AIAIAIAIAIAICurate smallmonomeric domainsfrom the PDBModel domainstructures andtrim flexible te…Design de novoand redesignedsequencescDNA displayproteolysis assayand deep sequenc…Infer K50 andfoldingstability, then …Principalcomponentanalysis of per-…Fit classifierpredictingwild-type amino …Identifyfunctional sitesusing evolutiona…
AI stepNo AI↤ what the AI stood in for
1Acquisition
no AI

Curate small monomeric domains from the PDB

Obtaining raw data, whether by measurement, download or retrieval.

We first collected all monomeric proteins in the PDB in the 30–100 amino acid length range in June 2021.where the paper describes this · verbatim
in the paper
2Inference
AI

Model domain structures and trim flexible termini

Running a trained model over new data to predict, classify or score.

We then predicted the structures of these PDB sequences using AlphaFold (even though the PDB structures were known)where the paper describes this · verbatim
in the paper
3Generation
AI

Design de novo and redesigned sequences

Producing candidate objects that did not previously exist.

to unconditionally generate protein backbones and sequences with lengths ranging from 46 to 69 amino acidswhere the paper describes this · verbatim
in the paper
4Experiment
no AI

cDNA display proteolysis assay and deep sequencing

Physical execution, by hand or by robot.

We then incubate the protein–cDNA complexes with different concentrations of protease, quench the reactions, and pull down the proteins using an N-terminal PA tagwhere the paper describes this · verbatim
in the paper
5Preparation
AI

Infer K50 and folding stability, then filter for quality

Cleaning, filtering, normalising or labelling data already obtained.

We inferred K50,U for each sequence using a position-specific scoring matrix model parameterized using measurements from 64,238 scrambled sequenceswhere the paper describes this · verbatim
in the paper
6Interpretation
AI

Principal component analysis of per-site amino acid stabilities

Extracting understanding from model behaviour.

we performed principal component analysis using 325,132 ΔG measurements at 17,093 sites in 365 domains (dataset 3)where the paper describes this · verbatim
in the paper
7Interpretation
AI

Fit classifier predicting wild-type amino acid from stabilities

Extracting understanding from model behaviour.

we created a simple classification model to predict the wild-type amino acid at any site in a natural protein based on the folding stabilitieswhere the paper describes this · verbatim
in the paper
8Inference
AI

Identify functional sites using evolutionary sensitivity scores

Running a trained model over new data to predict, classify or score.

We identified functional sites in 104 diverse protein domains by comparing each site’s average ΔΔG of substitutions with its normalized GEMME scorewhere the paper describes this · verbatim
in the paper

What the record says

Technical · every part carries its own basis

+ in the paper~ our reading− not reported

How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.

~Role of AI
AI in a supporting roleour reading

The central result is an experimental dataset of folding stabilities from cDNA display proteolysis. Learned models (AlphaFold, trRosetta hallucination, GEMME, fitted statistical models) were used to select and trim domains, generate some of the assayed designs, convert sequencing counts into stabilities and interpret the data, but the measurements themselves are experimental.

~How it was taught
SupervisedUnsupervisedZero-shotour reading
~Models named
AlphaFold · Off the shelftrRosetta (hallucination protocol) · Off the shelfGEMME · Off the shelfUnfolded state model (position-specific scoring matrix for K50,U) · Trained from scratchK50 model (Bayesian single-turnover kinetic model) · Trained from scratchPrincipal component analysis of per-site stabilities (scikit-learn) · Trained from scratchWild-type amino acid classifier model · Trained from scratchour reading
+How results were checked
Replication1188 testedin the paper
Our cDNA display proteolysis measurements are highly consistent with published studies using purified protein samples for 1,188 variants of 10 proteinswhere the paper describes this · verbatim
+Code · weights · data
code availableweights not reporteddata availablein the paper
The code for the analyses can be found at https://github.com/Rocklin-Lab/cdna-display-proteolysis-pipelinewhere the paper describes this · verbatim
+Compute
Quest high performance computing facility at Northwestern University (acknowledged computational resources); no accelerator time or run time statedin the paper

What this paper did not report

Technical · absence is published deliberately

Reported as not stated — 14 items
  • Trained model weightsWhether the trained model is available is not stated.
  • Version of AlphaFoldWhich version of the model was used is not stated.
  • Version of trRosetta (hallucination protocol)Which version of the model was used is not stated.
  • Version of GEMMEWhich version of the model was used is not stated.
  • Version of Unfolded state model (position-specific scoring matrix for K50,U)Which version of the model was used is not stated.
  • Version of K50 model (Bayesian single-turnover kinetic model)Which version of the model was used is not stated.
  • Version of Principal component analysis of per-site stabilities (scikit-learn)Which version of the model was used is not stated.
  • Version of Wild-type amino acid classifier modelWhich version of the model was used is not stated.
  • What step 2 replacedThe paper gives no basis for what the AI stood in for.
  • What step 3 replacedThe paper gives no basis for what the AI stood in for.
  • What step 5 replacedThe paper gives no basis for what the AI stood in for.
  • What step 6 replacedThe paper gives no basis for what the AI stood in for.
  • What step 7 replacedThe paper gives no basis for what the AI stood in for.
  • What step 8 replacedThe paper gives no basis for what the AI stood in for.

About this article

Record aix-00001, version 2, checked by a person on 2026-10-07. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY-4.0; quotations are at most 25 words. How we work · Report an error