~/aixsci
200 records · all checked

structural-biology/ai produced the result/PLoS ONE 2011 · v2

Protein shapes computed from patterns of change across families of related sequences

Researchers fitted a statistical model to the amino acid patterns in families of related protein sequences, letting it work out which residues are coupled. Those couplings became distance constraints that folded 15 test proteins into three-dimensional shapes.

1. Collect protein family sequence alignments2. Weight sequences to reduce sampling bias3. Fit maximum entropy model and infer residue couplings4. Convert top-ranked couplings into distance constraints5. Generate candidate all-atom structures6. Blind ranking and filtering of candidates7. Compare predictions with experimental structures8. Run alternative coupling methods through the same pipeline

spectrum · one line per step, placed by what the step does · bright lines used AI

Protein 3D Structure Computed from Evolutionary Sequence Variation
PLoS ONE, 2011

doi:10.1371/journal.pone.0028766 · record aix-00006 v2 · checked 2026-10-07

ai-resultrole of AI
AI was for
Structure determination
Model family
Probabilistic graphical model
Checked by
Held-out15 tested, 12 worked
Code
not reported

The finding the paper is about came from the AI.

read as

The science is explained before the AI appears. Switch to field specialist to go straight to the method.

Assumes the discipline and goes straight to the method.

The diagram, the record and what the paper did not report are identical in both modes. Only the framing changes — never the evidence.

Introduction by AIxSci · plain language

What this research was about

A protein is a chain of amino acids that folds into a specific three-dimensional shape, and that shape largely determines what the protein does. Working out the shape by experiment, usually by growing crystals and shining X-rays through them, is slow and does not work for every protein. So people have long wanted to compute the shape from the sequence alone. One clue lies in evolution. Proteins come in families of related sequences, and when two positions in the chain sit close together in the folded shape, a change at one often needs a compensating change at the other. Those paired changes leave a statistical trace in the family's sequences.

The trouble is that the trace is muddied. If position A is coupled to B, and B to C, then A and C will also look correlated even if they are nowhere near each other in space. Simple correlation counts cannot tell a real contact from a knock-on effect. The researchers set out to separate the two, and then to see whether the pairs that survived were enough to fold a protein. They assembled sequence alignments for 15 protein families, between 48 and 258 amino acids long, each with at least one member whose structure was already known from crystallography.

Where AI came in

The statistical model is where the computation sits. For each family, the researchers fitted a global probabilistic model to the alignment, matching how often each amino acid appears at each position and how often each pair appears together. Rather than scoring pairs one at a time, the model accounts for all positions at once, which lets it attribute an observed correlation to a direct coupling or to an indirect chain through other positions. It learned without being told any answers, from sequences alone, and was fitted fresh for each family. Its output was a ranked score for every pair of positions.

Everything downstream followed from that ranking. The top-scoring pairs were filtered by automated rules and turned into distance bounds, then an extended chain was folded to satisfy them using established physical modelling software, giving several hundred candidate shapes per protein. Candidates were ranked without reference to the known answers. For 12 of the 15 proteins the top-ranked shape came within 2.7 to 4.8 Angstrom of the crystal structure across at least three-quarters of the chain. The model stood in for the pairwise correlation measures used before it; shapes folded from those alternatives did not reach comparable accuracy.

Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.

The work

Technical · from the record

The authors fitted a maximum entropy statistical model to the amino acid frequencies and pair frequencies of a protein family's multiple sequence alignment, so that observed correlations between alignment columns were separated into direct residue-residue couplings and indirect ones. The top-ranked couplings were converted into distance constraints and used, together with predicted secondary structure, to fold extended polypeptide chains into all-atom structures by distance geometry and molecular dynamics. For 15 test proteins of 48 to 258 residues, including rhodopsin, the top blindly ranked structures of 12 proteins had coordinate errors of 2.7 to 4.8 Angstrom Ca-RMSD over at least 75% of residues. Structures folded from mutual information or SCA constraints did not reach reasonable accuracy, and constraints from a Bayesian network model gave structures below 5 Angstrom Ca-RMSD for 6 of 10 tested proteins.

How AI was used

For each protein family, a global maximum entropy (Potts-like, 21-state) model of full-length sequences was fitted to the single-column and column-pair amino acid frequencies of a PFAM alignment, after down-weighting sequences above 70% identity to family neighbours and requiring at least 1000 sequences per family. Couplings and single-residue terms were obtained by matrix inversion in a mean field approximation, and summed over amino acid pairs to give a direct information score per residue pair. The highest-scoring pairs were filtered by automated rules on predicted secondary structure, sequence separation, disulfide exclusivity and conservation, then converted into Ca, Cb and side-chain distance bounds and weighted restraints. Candidate all-atom structures were generated from an extended chain by distance geometry and simulated annealing molecular dynamics in CNS, for constraint counts incremented in steps of 10, with 20 structures per constraint bin; candidates were ranked by helix and strand handedness dihedral scores and filtered for knots. The same pipeline was also run on couplings from a Bayesian network model and on mutual information and SCA scores.

The shape of the work

Structural · the record, drawn

ACQUISITIONPREPARATIONTRAININGPREPARATIONGENERATIONSCREENINGVALIDATIONVALIDATION12345678AIAICollect proteinfamily sequencealignmentsWeight sequencesto reducesampling biasFit maximumentropy model andinfer residue co…Converttop-rankedcouplings into d…Generatecandidateall-atom structu…Blind ranking andfiltering ofcandidatesComparepredictions withexperimental str…Run alternativecoupling methodsthrough the same…↤ statistical model↤ statistical model
AI stepNo AI↤ what the AI stood in for
1Acquisition
no AI

Collect protein family sequence alignments

Obtaining raw data, whether by measurement, download or retrieval.

We identified a set of PFAM protein family sequence alignments with known crystal structure for at least one family memberwhere the paper describes this · verbatim
in the paper
2Preparation
no AI

Weight sequences to reduce sampling bias

Cleaning, filtering, normalising or labelling data already obtained.

sequences with over 70% residue identity to family neighbors are down-weightedwhere the paper describes this · verbatim
in the paper
3Training
AI

Fit maximum entropy model and infer residue couplings

Fitting model parameters, including fine-tuning an existing model. The AI stood in for statistical model.

once these parameters are determined by matrix inversion (Equations M4, M5), one can directly compute the effective pair probabilitieswhere the paper describes this · verbatim
in the paper
4Preparation
no AI

Convert top-ranked couplings into distance constraints

Cleaning, filtering, normalising or labelling data already obtained.

The first NC inferred EIC pairs, ranked according to their DCA coupling scores, are then translated to distance constraintswhere the paper describes this · verbatim
in the paper
5Generation
no AI

Generate candidate all-atom structures

Producing candidate objects that did not previously exist.

The protein polymers are folded from a fully extended amino acid sequence of the protein of interest using standard distance geometry techniqueswhere the paper describes this · verbatim
in the paper
6Screening
no AI

Blind ranking and filtering of candidates

Reducing a candidate set by filtering or ranking, in a single pass.

The elimination of mirror topologies and ranking of candidate structures is achieved by computing virtual dihedral angleswhere the paper describes this · verbatim
in the paper
7Validation
no AI

Compare predictions with experimental structures

Testing outputs against ground truth.

evaluation of prediction accuracy by computation of structural error of predicted contacts and predicted 3D structures relative to the reference crystal structurewhere the paper describes this · verbatim
in the paper
8Validation
AI

Run alternative coupling methods through the same pipeline

Testing outputs against ground truth. The AI stood in for statistical model.

We tested all three methods for their ability to generate protein folds for a number of families, using exactly the same pipelinewhere the paper describes this · verbatim
in the paper

What the record says

Technical · every part carries its own basis

+ in the paper~ our reading− not reported

How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.

~Role of AI
AI produced the resultour reading

The predicted 3D structures that the paper reports are produced by the fitted maximum entropy sequence model; the couplings it infers are the sole source of the distance constraints used to fold each protein

+What the AI was for
we use a global statistical model to compute a set of direct residue couplingswhere the paper describes this · verbatim
+Model families
+How it was taught
Unsupervisedin the paper
+Models named
EVfold mean field direct coupling analysis (DI) maximum entropy model · Trained from scratchBayesian network model (BNM) · Trained from scratchin the paper
+How results were checked
Held-out15 tested, 12 workedin the paper
For 12 out of the set of 15 protein families (Table 1), the top blindly ranked structures have coordinate errors from 2.7 Å–4.8 Åwhere the paper describes this · verbatim
+Code · weights · data
code not reportedweights not reporteddata availablein the paper
detailed 3D coordinates and Pymol session files for interactive inspection in Appendices A3 and A4, http://cbio.mskcc.org/foldingproteinswhere the paper describes this · verbatim
+Compute
Stated to run in well under an hour on a standard laptop computer for a medium-size protein, without high-performance computingin the paper

What this paper did not report

Technical · absence is published deliberately

Reported as not stated — 4 items
  • CodeWhether the code is available is not stated.
  • Trained model weightsWhether the trained model is available is not stated.
  • Version of EVfold mean field direct coupling analysis (DI) maximum entropy modelWhich version of the model was used is not stated.
  • Version of Bayesian network model (BNM)Which version of the model was used is not stated.

About this article

Record aix-00006, version 2, checked by a person on 2026-10-07. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY-4.0; quotations are at most 25 words. How we work · Report an error