structural-biology/ai produced the result/Nature 2021 · v2
Neural network predicts protein atom positions from amino acid sequence alone
AlphaFold, a neural network trained on known protein structures, predicts where every heavy atom in a protein sits using only its amino acid sequence and alignments of related sequences. The predicted structures are the result itself.
spectrum · one line per step, placed by what the step does · bright lines used AI
Highly accurate protein structure prediction with AlphaFold
Nature, 2021
doi:10.1038/s41586-021-03819-2 · record aix-00007 v2 · checked 2026-10-07
- AI was for
- Structure determination, Property prediction
- Model family
- Transformer, Multilayer perceptron
- Checked by
- Benchmark87 tested
- Code
- available
The finding the paper is about came from the AI.
What this research was about
Proteins are built as chains of amino acids, strung together in an order written in the gene. But a protein only works once that chain has folded into a particular three-dimensional shape. The shape decides what the protein can bind to and what it does, so biologists badly want to know it. Working it out by experiment is slow: it usually means crystallising the protein, or freezing it, and then reading the atom positions off the resulting images. Predicting the shape from the sequence instead has been a long-standing problem, because the number of ways a chain can fold is vast and the forces that pick out the real one are subtle.
One useful clue comes from evolution. The same protein exists, slightly altered, in many species, and when two positions in the chain sit close together in the folded shape they tend to change in step across those species. Collecting the related sequences into a stack, lined up position by position, is called a multiple sequence alignment. The researchers set out to turn those alignments, plus the sequence itself and any already-solved structures of similar proteins, directly into atomic coordinates.
Where AI came in
The AlphaFold network does the prediction. It takes the amino acid sequence, an alignment of related sequences built with standard search tools, and template structures where available, and outputs the three-dimensional coordinates of all the protein's heavy atoms, along with a per-residue score saying how confident it is. It was trained on structures from the Protein Data Bank released up to 30 April 2018. The trained network was then used to predict structures for around 350,000 further sequences, and its own high-confidence predictions were fed back in to train the network again from scratch.
What the network stands in for is the experimental work of determining a structure. In the blind CASP14 assessment, across 87 protein domains, its predictions matched the real backbone to a median of 0.96 Å, against 2.8 Å for the next best method. Afterwards the coordinates were tidied by conventional physics simulation, which fixes small chemical implausibilities but is not itself learned.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record
AlphaFold is a neural network that predicts the three-dimensional coordinates of a protein's heavy atoms from its amino acid sequence together with multiple sequence alignments of homologues and, where available, template structures. It was trained on PDB structures released up to 30 April 2018, then retrained using a self-distillation dataset of its own high-confidence predictions for around 350,000 Uniclust30 sequences. In the blind CASP14 assessment (87 protein domains), its predictions had a median backbone accuracy of 0.96 Å r.m.s.d.95 compared with 2.8 Å for the next best performing method. The authors also report that its per-residue pLDDT confidence score correlates with measured lDDT-Cα accuracy on 10,795 recent PDB chains.
How AI was used
Sequence databases (UniRef90, BFD, Uniclust30, MGnify) were searched with jackhmmer and HHBlits to build multiple sequence alignments, and PDB70 was searched with HHSearch for templates; these, plus the primary sequence, are the network inputs. The AlphaFold network has an Evoformer trunk that jointly updates an MSA representation and a pairwise residue representation through attention-based and triangle multiplicative updates, followed by a structure module that uses invariant point attention to iteratively update per-residue rigid frames and side-chain torsion angles, with outputs recycled back through the network. Training used the frame-aligned point error loss plus auxiliary distogram, BERT-style masked-MSA, per-residue lDDT, side-chain and structural violation losses, on 128 TPU v3 cores with 256-residue crops and later 384-residue fine-tuning. An initial model's high-confidence predictions for around 350,000 Uniclust30 sequences formed a distillation set used to retrain the architecture from scratch, sampling 75% from that set and 25% from clustered PDB. Five models with different random seeds were run at inference and selected per target by predicted confidence; predictions were then relaxed with OpenMM and the Amber99sb force field. Separate per-block structure modules were trained on the frozen network to read out intermediate structures.
The shape of the work
Structural · the record, drawn
no AI
Build sequence databases and search MSAs and templates
Cleaning, filtering, normalising or labelling data already obtained.
sequences from evolutionarily related proteins in the form of a MSA created by standard tools including jackhmmer and HHBlitswhere the paper describes this · verbatim
AI
Train AlphaFold on PDB structures
Fitting model parameters, including fine-tuning an existing model. The AI stood in for conventional algorithm.
To train, we use structures from the PDB with a maximum release date of 30 April 2018.where the paper describes this · verbatim
AI
Predict structures for unlabelled Uniclust sequences
Running a trained model over new data to predict, classify or score. The AI stood in for new capability.
we use a trained network to predict the structure of around 350,000 diverse sequences from Uniclust30 and make a new dataset of predicted structureswhere the paper describes this · verbatim
AI
Retrain with self-distillation and fine-tune
Fitting model parameters, including fine-tuning an existing model. The AI stood in for conventional algorithm.
These predictions were then used to train a final model with identical hyperparameters, except for sampling examples 75% of the timewhere the paper describes this · verbatim
AI
Predict structures and estimate confidence
Running a trained model over new data to predict, classify or score. The AI stood in for physical experiment.
We inference the five trained models and use the predicted confidence score to select the best model per target.where the paper describes this · verbatim
no AI
Relax structures in Amber force field
Numerical or physics simulation, including where a learned surrogate replaces it.
For constrained relaxation of structures, we used OpenMM v.7.3.1 with the Amber99sb force field.where the paper describes this · verbatim
no AI
Evaluate against CASP14 and recent PDB structures
Testing outputs against ground truth.
The predicted structure is compared to the true structure from the PDB in terms of lDDT metricwhere the paper describes this · verbatim
AI
Probe network behaviour with per-block structure modules
Extracting understanding from model behaviour. The AI stood in for new capability.
we trained a separate structure module for each of the 48 Evoformer blocks in the network while keeping all parameters of the main network frozenwhere the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
The paper's result is the predicted structures themselves, produced by the AlphaFold neural network
The AlphaFold network directly predicts the 3D coordinates of all heavy atoms for a given protein using the primary amino acid sequencewhere the paper describes this · verbatim
The performance of AlphaFold on the CASP14 dataset (n = 87 protein domains) relative to the top-15 entries (out of 146 entries)where the paper describes this · verbatim
These timings are measured using our open-source code, and the open-source code is notably faster than the version we ran in CASP14where the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- Trained model weightsWhether the trained model is available is not stated.
About this article
Record aix-00007, version 2, checked by a person on 2026-10-07. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY-4.0; quotations are at most 25 words. How we work · Report an error