~/aixsci
200 records · all checked

structural-biology/ai produced the result/Frontiers in Molecular Biosciences 2022 · v2

Machine learning tested against three definitions of protein shape-shifting

Researchers trained random forest classifiers on sequence-derived biophysical predictions to tell ordered, disordered and 'ambiguous' protein residues apart. The models' scores and rules were the evidence used to compare rival definitions of protein order.

1. Assemble labelled residue datasets2. Assemble proteome, PTM and mutation annotation sets3. Predict per-residue biophysical features from sequence4. Train and test random forest residue classifiers5. Classify MFIB residues with the combined model6. Derive surrogate rule models for interpretation7. Relate classes to pLDDT, PTMs and variants

spectrum · one line per step, placed by what the step does · bright lines used AI

Challenges in describing the conformation and dynamics of proteins with ambiguous behavior
Frontiers in Molecular Biosciences, 2022

doi:10.3389/fmolb.2022.959956 · record aix-00172 v2 · checked 2026-10-09

ai-resultrole of AI
AI was for
Classification, Property prediction
Model family
Random forest
Checked by
Held-out
Code
not reported

The finding the paper is about came from the AI.

read as

The science is explained before the AI appears. Switch to field specialist to go straight to the method.

Assumes the discipline and goes straight to the method.

The diagram, the record and what the paper did not report are identical in both modes. Only the framing changes — never the evidence.

Introduction by AIxSci · plain language

What this research was about

Proteins are chains of amino acids, and much of biology textbook teaching assumes each chain folds into one fixed shape. Many do not. Some stay floppy until they meet a partner molecule, then settle into a shape only once bound. Others flip between two different folds. Both behaviours sit awkwardly between the neat categories of 'ordered' and 'disordered'. That matters because the categories are not just labels: they feed the curated databases and prediction tools that researchers use to reason about what a protein does. If a residue's behaviour is ambiguous, it is not obvious which bin it belongs in, or whether the two kinds of ambiguity belong in the same bin at all.

The authors assembled residue-by-residue labels from curated structural resources: disorder-to-order transitions, conformationally stable residues, secondary-structure switches annotated from pairs of structures, and a set of mutually folding proteins kept aside. They then asked how separable these classes really are from sequence alone, and how the answer changes when folding-upon-binding and fold-switching are merged into a single ambiguous class.

Where AI came in

Seven per-residue properties were predicted from each sequence by existing tools — backbone and side-chain dynamics, helix, sheet and coil propensities, early folding propensity and disorder. Those predictions alone, with the amino acid identities withheld, were the input to random forest classifiers: one trained on the folding-upon-binding labels, one on the fold-switching labels, one on the merged set. A random forest is a crowd of simple decision trees that vote on an answer. The models were built for interpretability rather than top scores, so the authors read off which features mattered most and summarised each forest as a set of plain if-then rules.

The classifiers stood in for the judgement a curator would otherwise make residue by residue. Their performance became the measurement: where a class scored poorly, that was taken as evidence about the definition rather than only about the model. The merged model was then applied to the held-out mutually folding set, and its per-residue calls across the human proteome were compared with AlphaFold2's per-residue confidence scores, with sites of chemical modification, and with disease-linked and benign mutations.

Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.

The work

Technical · from the record

The authors assembled residue-level datasets of three kinds of protein conformational behaviour — ordered, disordered, and 'ambiguous' residues that either fold upon binding or switch secondary structure — and trained random forest classifiers on seven sequence-based biophysical predictions to separate these classes. Sequence-predicted disorder, early folding propensity and backbone dynamics were the most important features; the model trained on the folding-upon-binding set had its lowest F1 score for the disorder class, and the fold-switching model reached an F1 of 0.36 for converting residues, with a recall of 0.26. Merging the folding-upon-binding and fold-switching categories into one combined model lowered F1 performance for the merged ordered class relative to the separate ordered and same classes. Applying the combined model to a held-out MFIB set of mutually folding proteins predicted 79.6% of residues as ordered, 20% as ambiguous and under 1% as disordered, and comparisons with AlphaFold2 pLDDT values placed ambiguous residues between the ordered and disordered classes.

How AI was used

Residues were labelled from curated structural resources: disorder-to-order transitions from DisProt, conformationally stable residues from CoDNaS clusters filtered at a 2 A maximum pairwise RMSD, secondary-structure switches from DSSP annotations of fold-switcher structure pairs, and a redundancy-filtered MFIB set held out of training. For every sequence, seven per-residue features were predicted with the b2bTools suite — backbone and side-chain dynamics and conformational propensities from DynaMine, early folding propensity from EFoldMine, and disorder from DisoMine — and these features alone, with no amino acid codes, were the inputs to random forest classifiers built with scikit-learn: one on the DisProt/CoDNaS labels split 90/10 into train and test, one on the fold-switching labels, and one on the merged combined set split 70/30, each with hyperparameters chosen by 3-fold cross-validation. The stated aim was interpretability rather than best performance, so feature importances were extracted and each forest was additionally summarised by a surrogate rule model induced with the Weka implementation of the Ripper algorithm over the forest's predictions. The combined model was then run over the MFIB set, and the per-residue predictions for the human proteome were cross-tabulated against AlphaFold2 pLDDT values and DSSP categories from downloaded AlphaFold2 models, against PTM sites compiled from four databases, and against deleterious, benign, somatic and germline missense variants.

The shape of the work

Structural · the record, drawn

ACQUISITIONACQUISITIONREPRESENTATIONTRAININGINFERENCEINTERPRETATIONINTERPRETATION1234567AIAIAIAIAssemble labelledresidue datasetsAssembleproteome, PTM andmutation annotat…Predictper-residuebiophysical feat…Train and testrandom forestresidue classifi…Classify MFIBresidues with thecombined modelDerive surrogaterule models forinterpretationRelate classes topLDDT, PTMs andvariants↤ manual curation↤ manual curation↤ expert judgement
AI stepNo AI↤ what the AI stood in for
1Acquisition
no AI

Assemble labelled residue datasets

Obtaining raw data, whether by measurement, download or retrieval.

we downloaded a custom set of human proteins with manually curated disorder-to-order structural transitions, resulting in 138 different proteinswhere the paper describes this · verbatim
in the paper
2Acquisition
no AI

Assemble proteome, PTM and mutation annotation sets

Obtaining raw data, whether by measurement, download or retrieval.

AlphaFold 2’s mmCIF files for the human proteome were downloaded on 2 September 2021, from the AlphaFold protein structure database.where the paper describes this · verbatim
in the paper
3Representation
AI

Predict per-residue biophysical features from sequence

Encoding data into features, descriptors, embeddings or graphs.

seven biophysical features were predicted at the residue level using the following methods: backbone dynamics (DynaMine)where the paper describes this · verbatim
in the paper
4Training
AI

Train and test random forest residue classifiers

Fitting model parameters, including fine-tuning an existing model. The AI stood in for manual curation.

We used a combination of these datasets (disprot_codnas_set) to train a random forest (RF) predictorwhere the paper describes this · verbatim
in the paper
5Inference
AI

Classify MFIB residues with the combined model

Running a trained model over new data to predict, classify or score. The AI stood in for manual curation.

the disordered class, without ambiguous folding propensity, was shown to be depleted in the output of the combined_RF predictorwhere the paper describes this · verbatim
in the paper
6Interpretation
AI

Derive surrogate rule models for interpretation

Extracting understanding from model behaviour. The AI stood in for expert judgement.

The RF models were interpreted using a surrogate model trained over the predictions for each of the models.where the paper describes this · verbatim
in the paper
7Interpretation
no AI

Relate classes to pLDDT, PTMs and variants

Extracting understanding from model behaviour.

related the key biophysical predictions of the selected_human_set with the respective pLDDT values of the AlphaFold2 modelswhere the paper describes this · verbatim
in the paper

What the record says

Technical · every part carries its own basis

+ in the paper~ our reading− not reported

How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.

~Role of AI
AI produced the resultour reading

The paper's comparisons between definitions of order, disorder and ambiguity are read off the performance, feature importance and surrogate rules of random forest classifiers the authors trained, and off sequence-based predictor outputs; without those models there is no result.

+What the AI was for
The classification model was trained using seven predicted biophysical features at the residue levelwhere the paper describes this · verbatim
+Model families
Random forestin the paper
+How it was taught
Supervisedin the paper
+Models named
folding_upon_binding_RF (random forest) · Trained from scratchfold_switching_RF (random forest) · Trained from scratchcombined_RF (random forest) · Trained from scratchRipper (Weka JRip) surrogate rule model · Trained from scratchDynaMine · Off the shelfDisoMine · Off the shelfEFoldMine · Off the shelfAlphaFold2 · Off the shelfin the paper
+How results were checked
Held-outin the paper
The RF model is trained using those hyperparameters and finally tested on the remaining 10% of the data (test set)where the paper describes this · verbatim
+Code · weights · data
code not reportedweights not reporteddata availablein the paper
The complete dataset is available at https://bitbucket.org/bio2byte/protein_ambiguity/.where the paper describes this · verbatim
+Compute
not reportedin the paper

What this paper did not report

Technical · absence is published deliberately

Reported as not stated — 13 items
  • CodeWhether the code is available is not stated.
  • Trained model weightsWhether the trained model is available is not stated.
  • ComputeThe hardware or time used is not stated.
  • How many were testedThe paper gives no count of what was tested.
  • Version of folding_upon_binding_RF (random forest)Which version of the model was used is not stated.
  • Version of fold_switching_RF (random forest)Which version of the model was used is not stated.
  • Version of combined_RF (random forest)Which version of the model was used is not stated.
  • Version of Ripper (Weka JRip) surrogate rule modelWhich version of the model was used is not stated.
  • Version of DynaMineWhich version of the model was used is not stated.
  • Version of DisoMineWhich version of the model was used is not stated.
  • Version of EFoldMineWhich version of the model was used is not stated.
  • Version of AlphaFold2Which version of the model was used is not stated.
  • What step 3 replacedThe paper gives no basis for what the AI stood in for.

About this article

Record aix-00172, version 2, checked by a person on 2026-10-09. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error