~/aixsci
200 records · all checked

structural-biology/ai produced the result/Nature Communications 2024 · v2

Teaching protein language models to rank mutants from a few dozen lab measurements

Researchers built a training strategy, FSFP, that adapts large protein language models using only tens of measured mutants. The models then ranked mutants across a public benchmark and picked candidates for a DNA-copying enzyme tested in the lab.

1. Assemble benchmark and few-shot splits2. Embed wild-type proteins and retrieve similar datasets3. Generate MSA-based pseudo labels4. Meta-train PLM on auxiliary tasks with LoRA5. Fine-tune on target data with listwise ranking loss6. Score held-out mutants and benchmark against baselines7. Select Phi29 single-site mutants for testing8. Express, purify and measure Tm of mutants

spectrum · one line per step, placed by what the step does · bright lines used AI

Enhancing efficiency of protein language models with minimal wet-lab data through few-shot learning
Nature Communications, 2024

doi:10.1038/s41467-024-49798-6 · record aix-00180 v2 · checked 2026-10-09

ai-resultrole of AI
AI was for
Property prediction, Experimental design
Model family
Protein language model, Transformer, Linear model
Checked by
Experimental20 tested
Code
available

The finding the paper is about came from the AI.

read as

The science is explained before the AI appears. Switch to field specialist to go straight to the method.

Assumes the discipline and goes straight to the method.

The diagram, the record and what the paper did not report are identical in both modes. Only the framing changes — never the evidence.

Introduction by AIxSci · plain language

What this research was about

Proteins are chains of amino acids, and changing a single one can make the protein more stable, less stable, or stop it working. Biologists call the useful measure of how well a variant performs its fitness. Working out which single changes help is hard: a modest protein has thousands of possible single swaps, and measuring each one in the laboratory costs time and money. Methods that learn from measurements usually need many of them, which is exactly what is in short supply when a new protein is being engineered.

The authors set out to make large models of protein sequence useful when only a handful of laboratory measurements exist. Their strategy, FSFP, adapts an already-trained model using tens of measured single-amino-acid variants of the protein in hand, rather than hundreds or thousands.

Where AI came in

Three pre-trained protein language models — ESM-1v, ESM-2 and SaProt, programs trained on large collections of protein sequences much as text models are trained on text — did the predicting. The model first compared the target protein with proteins in the ProteinGym collection of mutation measurements and borrowed the two most similar sets as practice material. A third practice set came from GEMME, a method based on alignments of related sequences, which supplied stand-in labels. The model was then primed on these practice tasks and finally tuned on the target protein's own labelled mutants, learning to put them in order rather than predict exact values.

The trained models were scored on held-out mutants across the 87 benchmark datasets, and compared with their untrained versions and with a ridge regression baseline. For Phi29 DNA polymerase, an enzyme that copies DNA, ESM-1v chose the top 20 single-amino-acid variants for laboratory melting-temperature measurements, standing in for screening every possible variant at the bench; those measurements then fed a further round of training.

Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.

The work

Technical · from the record

The authors built FSFP, a training strategy that adapts pre-trained protein language models to predict the fitness of protein mutants using only tens of labelled single-site mutants from the target protein. It combines retrieval of similar mutational-scanning datasets, pseudo labels from the alignment-based method GEMME, meta-training with MAML, low-rank adaptation, and a listwise ranking loss. Across the 87 deep mutational scanning datasets of the ProteinGym substitution benchmark, FSFP-trained ESM-1v, ESM-2 and SaProt scored higher on average Spearman correlation than their zero-shot versions and than the ridge regression baseline at every training size tested. Applied to Phi29 DNA polymerase, the top 20 single-site mutants predicted by ESM-1v after FSFP training had an average melting temperature more than 1 °C higher and a positive rate 25% higher than the top 20 from the zero-shot model.

How AI was used

Three pre-trained protein language models — ESM-1v, ESM-2 and SaProt, each at 650 M parameters — were adapted to rank mutant fitness. The model to be trained first embedded wild-type sequences (or structures, for SaProt) of the target protein and of the proteins in ProteinGym, and the two datasets with highest cosine similarity became auxiliary meta-learning tasks; a third task was built by scoring candidate mutants of the target protein with the alignment-based method GEMME to produce pseudo labels. MAML was then used to meta-train the model on these tasks, with updates confined to LoRA rank-decomposition matrices injected into the self-attention and feed-forward weights (rank 16) while pre-trained weights stayed frozen, and a first-order approximation used for the outer gradient. The meta-trained initialisation was fine-tuned on the target protein's labelled mutants with a ListMLE listwise ranking loss, mutant scores being computed from the model's residue probabilities relative to the wild type; Monte Carlo cross-validation on the training data set the number of steps and early-stopped meta-training. Trained models were then run over held-out mutants of the benchmark and, for Phi29 DNA polymerase, over saturated single-site mutants to pick the top 20 for wet-lab melting-temperature assays, whose labels fed a further round of training. A ridge regression approach on one-hot plus density features, GEMME, and ridge-augmented variants were run as baselines.

The shape of the work

Structural · the record, drawn

PREPARATIONREPRESENTATIONINFERENCETRAININGTRAININGVALIDATIONSCREENINGEXPERIMENT12345678AIAIAIAIAIAIAssemblebenchmark andfew-shot splitsEmbed wild-typeproteins andretrieve similar…GenerateMSA-based pseudolabelsMeta-train PLM onauxiliary taskswith LoRAFine-tune ontarget data withlistwise ranking…Score held-outmutants andbenchmark agains…Select Phi29single-sitemutants for test…Express, purifyand measure Tm ofmutants↤ conventional algorithm↤ statistical model↤ physical experimentloops back
AI stepNo AI↤ what the AI stood in for
1Preparation
no AI

Assemble benchmark and few-shot splits

Cleaning, filtering, normalising or labelling data already obtained.

we first randomly sample 20 single-site mutants as an initial training setwhere the paper describes this · verbatim
in the paper
2Representation
AI

Embed wild-type proteins and retrieve similar datasets

Encoding data into features, descriptors, embeddings or graphs. The AI stood in for conventional algorithm.

the relevance between the target protein and a candidate protein is measured by the cosine similarity between their embeddings p and qwhere the paper describes this · verbatim
in the paper
3Inference
AI

Generate MSA-based pseudo labels

Running a trained model over new data to predict, classify or score.

Utilizing GEMME, we score the candidate mutated sequences of the target protein and build the dataset of the third task.where the paper describes this · verbatim
in the paper
4Training
AI

Meta-train PLM on auxiliary tasks with LoRA

Fitting model parameters, including fine-tuning an existing model.

We apply MAML, a state-of-the-art meta-learning algorithm, to enable PLMs to better utilize the few-shot training data.where the paper describes this · verbatim
in the paper
5Training
AI

Fine-tune on target data with listwise ranking loss

Fitting model parameters, including fine-tuning an existing model. The AI stood in for statistical model.

We use a listwise LTR approach, namely ListMLE to train PLMs on few-shot training data.where the paper describes this · verbatim
in the paper
6Validation
AI

Score held-out mutants and benchmark against baselines

Testing outputs against ground truth.

the predictive performance is measured by two metrics: Spearman rank correlation and normalized discounted cumulative gain (NDCG) with the fitness labels as ground truthwhere the paper describes this · verbatim
in the paper
7Screening
AI

Select Phi29 single-site mutants for testing

Reducing a candidate set by filtering or ranking, in a single pass. The AI stood in for physical experiment.

Then, 20 mutants with the highest predicted scores are chosen to measure experimental Tm values.where the paper describes this · verbatim
in the paper
8Experiment
no AI

Express, purify and measure Tm of mutants

Physical execution, by hand or by robot. Its result feeds back into an earlier step.

The Tm values are determined by differential scanning fluorimetry (DSF) method using the Protein Thermal Shift Dye Kit (Thermo Fisher).where the paper describes this · verbatim
in the paper

What the record says

Technical · every part carries its own basis

+ in the paper~ our reading− not reported

How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.

~Role of AI
AI produced the resultour reading

The paper's result is the trained-model strategy itself and the mutants it selected for wet-lab testing; both the benchmark finding and the Phi29 engineering outcome depend entirely on the models

+What the AI was for
By combining meta-transfer learning, learning to rank, and parameter-efficient fine-tuning, FSFP can significantly boost the performance of various protein language modelswhere the paper describes this · verbatim
+How it was taught
SupervisedTransfer / fine-tuningZero-shotSemi-supervisedin the paper
+Models named
ESM-1v 650 M parameters, first checkpoint of the five-model ensemble · Fine-tunedESM-2 650 M · Fine-tunedSaProt 650 M, version continuously pre-trained on PDB structures · Fine-tunedGEMME · Off the shelfridge regression (Hsu et al. approach) official implementation · Trained from scratchin the paper
+How results were checked
Experimental20 testedin the paper
The top 20 predictions from the trained model for single-site mutants are selected for the next iteration of wet-lab experiments.where the paper describes this · verbatim
+Code · weights · data
code availableweights not reporteddata availablein the paper
The source code of FSFP is available at https://github.com/ai4protein/FSFP.where the paper describes this · verbatim
+Compute
not reportedin the paper

What this paper did not report

Technical · absence is published deliberately

Reported as not stated — 6 items
  • Trained model weightsWhether the trained model is available is not stated.
  • ComputeThe hardware or time used is not stated.
  • Version of GEMMEWhich version of the model was used is not stated.
  • What step 3 replacedThe paper gives no basis for what the AI stood in for.
  • What step 4 replacedThe paper gives no basis for what the AI stood in for.
  • What step 6 replacedThe paper gives no basis for what the AI stood in for.

About this article

Record aix-00180, version 2, checked by a person on 2026-10-09. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error