~/aixsci
200 records · all checked

structural-biology/ai produced the result/Nature Communications 2022 · v2

Deep learning turns raw NMR spectra into protein structures without human intervention

Researchers built ARTINA, a workflow that takes NMR spectra and a protein's sequence and returns a structure. Five trained models read the spectra, fill in missing measurements and choose between candidate structures, doing work that normally needs an expert's eye.

1. Collect and standardise NMR spectrum benchmark2. Build training sets (annotated, synthetic and graph data)3. Train peak picking, deconvolution, shift prediction, ranking and density models4. Visual spectrum analysis: pick and deconvolve cross-peaks5. Unalias folded peaks to true resonance frequencies6. Assign chemical shifts and refine with GNN predictions7. Calculate structure proposals and select with GBT8. Compare automated results with reference structures and assignments

spectrum · one line per step, placed by what the step does · bright lines used AI

Rapid protein assignments and structures from raw NMR spectra with the deep learning technique ARTINA
Nature Communications, 2022

doi:10.1038/s41467-022-33879-5 · record aix-00002 v2 · checked 2026-10-07

ai-resultrole of AI
AI was for
Detection, Property prediction, Structure determination
Model family
Convolutional neural network, Graph neural network, Gradient-boosted trees, Clustering
Checked by
Held-out100 tested
Code
available

The finding the paper is about came from the AI.

read as

The science is explained before the AI appears. Switch to field specialist to go straight to the method.

Assumes the discipline and goes straight to the method.

The diagram, the record and what the paper did not report are identical in both modes. Only the framing changes — never the evidence.

Introduction by AIxSci · plain language

What this research was about

Proteins do their jobs by folding into particular three-dimensional shapes, so knowing the shape helps explain what a protein does. One way to work that shape out is nuclear magnetic resonance, or NMR: a protein in solution is placed in a strong magnetic field, and its atomic nuclei respond at frequencies that depend on their chemical surroundings. The result is a spectrum, a landscape of peaks. Each genuine peak reports that two or more atoms are close together or chemically linked. The trouble is that peaks overlap, noise and artefacts look much like real signals, and peaks can be 'folded' so they appear at the wrong frequency.

Interpreting those spectra has traditionally meant weeks of careful work by a spectroscopist, deciding which bumps are real, matching each one to a specific atom in the protein chain, and then judging which of several computed structures to trust. The researchers set out to automate the whole chain, from the raw spectra and the amino acid sequence through to a finished structure, with no human step in between. They assembled a benchmark of 1329 spectra for 100 proteins, with previously published assignments and structures to compare against.

Where AI came in

Five learned components sit inside the pipeline. One residual network, a type of image-recognition model, scores every bump in a spectrum as real signal or artefact; a second separates overlapping peaks into their components. A density-estimation model works out which peaks are folded and where they truly belong. A graph neural network, which learns from data arranged as networks of connected atoms, predicts the measurements the assignment routine could not pin down confidently, and those predictions guide a second pass. A gradient boosted trees model then compares ten candidate structures and picks one, which is fed back into the next assignment cycle.

Together these models stand in for the expert judgement usually applied at each of those stages: deciding what is a peak, what frequency it belongs at, what the unassigned values should be, and which structure to keep. The authors report a median difference of 1.44 ångström in the protein backbone between their structures and the reference ones, and on average 90.39% of measurements matching the manually prepared assignments, taking 4 to 20 hours per protein. Code and trained models are available.

Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.

The work

Technical · from the record

ARTINA is an automated workflow that takes only NMR spectra and a protein sequence and returns cross-peak positions, chemical shift assignments and a protein structure without human intervention. Deep residual networks pick and deconvolve cross-peaks, a kernel density estimator unfolds folded peaks, a graph neural network predicts missing chemical shifts during assignment refinement, and a gradient boosted trees model ranks the ten structure proposals produced by CYANA calculations. On a benchmark of 1329 2D/3D/4D spectra for 100 proteins, evaluated by 5-fold cross-validation, the workflow produced structures with a median backbone RMSD of 1.44 Å to the reference PDB structures and, on average, 90.39% of chemical shifts matching the manually prepared assignments. Reported computation time was 4–20 h per protein.

How AI was used

Five learned components were trained and then run inside a single automated pipeline. A residual network (pp-ResNet, 8 residual blocks, trained on 675,423 manually annotated 2D spectrum fragments) scores every signal extremum in a spectrum as true signal or artefact; a second residual network (deconv-ResNet), trained with a Chamfer distance loss on 110,000 synthetic fragments from a generator based on NMR physics, regresses the coordinates of up to three deconvolved components around each selected extremum. Per-experiment-type kernel density estimators, fitted on chemical shift lists from BMRB depositions excluding the benchmark proteins, define a discrete optimisation that unfolds aliased peaks to true resonance frequencies. The resulting peak lists are passed to the FLYA genetic-algorithm assignment routine; a DeepGCN graph neural network, trained on 28,400 graphs from 2840 referenced BMRB records, then predicts expected values of shifts FLYA did not confidently assign, and those predictions constrain a second FLYA call. Ten NOESY peak list variants at ResNet confidence thresholds from 0.05 to 0.5 drive ten independent CYANA structure calculations with TALOS-N torsion angle restraints, and a gradient boosted trees model trained on pairwise comparisons of CYANA runs, using 96 features per proposal, ranks the proposals; the selected structure is fed back into the next assignment cycle. Both refinement cycles are executed three times. pp-ResNet and GBT were trained in 5-fold cross-validation over the 100-protein benchmark and deployed as ensembles of the five fold models.

The shape of the work

Structural · the record, drawn

ACQUISITIONPREPARATIONTRAININGINFERENCEPREPARATIONINFERENCEOPTIMISATIONVALIDATION12345678AIAIAIAIAICollect andstandardise NMRspectrum benchma…Build trainingsets (annotated,synthetic and gr…Train peakpicking,deconvolution, s…Visual spectrumanalysis: pickand deconvolve c…Unalias foldedpeaks to trueresonance freque…Assign chemicalshifts and refinewith GNN predict…Calculatestructureproposals and se…Compare automatedresults withreference struct…↤ expert judgement↤ expert judgement↤ expert judgement↤ expert judgement↤ expert judgementloops back · 3 rounds
AI stepNo AI↤ what the AI stood in for
1Acquisition
no AI

Collect and standardise NMR spectrum benchmark

Obtaining raw data, whether by measurement, download or retrieval.

we implemented a crawler software, which systematically scanned the FTP server of the BMRB data bank, identifying data files relevant to our studywhere the paper describes this · verbatim
in the paper
2Preparation
no AI

Build training sets (annotated, synthetic and graph data)

Cleaning, filtering, normalising or labelling data already obtained.

675,423 diverse 2D fragments of size 256 × 32 × 1 were extracted from the normalized spectra and manually annotatedwhere the paper describes this · verbatim
in the paper
3Training
AI

Train peak picking, deconvolution, shift prediction, ranking and density models

Fitting model parameters, including fine-tuning an existing model. The AI stood in for expert judgement.

Five instances of pp-ResNet and GBT were trained, each one using data from about 80% of the proteins for trainingwhere the paper describes this · verbatim
in the paper
4Inference
AI

Visual spectrum analysis: pick and deconvolve cross-peaks

Running a trained model over new data to predict, classify or score. The AI stood in for expert judgement.

cross-peak positions are identified in frequency-domain NMR spectra using deep residual neural networks (ResNet)where the paper describes this · verbatim
in the paper
5Preparation
AI

Unalias folded peaks to true resonance frequencies

Cleaning, filtering, normalising or labelling data already obtained. The AI stood in for expert judgement.

We addressed the problem of automated signal unfolding with the classical machine learning approach to density estimationwhere the paper describes this · verbatim
in the paper
6Inference
AI

Assign chemical shifts and refine with GNN predictions

Running a trained model over new data to predict, classify or score. The AI stood in for expert judgement.

we bring to the workflow contextual information about thousands of protein structures solved by NMR in the past using a deep GNNwhere the paper describes this · verbatim
in the paper
7Optimisation
AI

Calculate structure proposals and select with GBT

Iterative search over a space. The AI stood in for expert judgement. Its result feeds back into an earlier step.

The structure proposals are ranked in the intermediate structure selection step based on 96 features with a dedicated GBT modelwhere the paper describes this · verbatim
in the paper
8Validation
no AI

Compare automated results with reference structures and assignments

Testing outputs against ground truth.

ARTINA was able to reproduce the reference structures with a median backbone root-mean-square deviation (RMSD) of 1.44 Åwhere the paper describes this · verbatim
in the paper

What the record says

Technical · every part carries its own basis

+ in the paper~ our reading− not reported

How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.

~Role of AI
AI produced the resultour reading

The reported peak positions, resonance assignments and structures are produced by the learned models inside the ARTINA workflow; the paper's results are those outputs.

+What the AI was for
deep residual neural networks for visual spectrum analysis to identify peak positions (pp-ResNet) and to deconvolve overlapping signals (deconv-ResNet)where the paper describes this · verbatim
+How it was taught
SupervisedUnsupervisedin the paper
+Models named
pp-ResNet · Trained from scratchdeconv-ResNet · Trained from scratchGNN (DeepGCN chemical shift predictor) · Trained from scratchGBT structure ranking model · Trained from scratchKDE signal unaliasing model · Trained from scratchin the paper
+How results were checked
Held-out100 testedin the paper
The accuracy of protein structure determination with ARTINA was evaluated in a 5-fold cross-validation experiment with the aforementioned benchmark datasetwhere the paper describes this · verbatim
+Code · weights · data
code availableweights availabledata availablein the paper
pp-ResNet, deconv-ResNet, GNN, and GBT are available for download in binary formwhere the paper describes this · verbatim
+Compute
Execution times of 4–20 h per protein; analysis of a large 3D C-resolved NOESY spectrum in less than 5 min on a high-end desktop computerin the paper

What this paper did not report

Technical · absence is published deliberately

Reported as not stated — 5 items
  • Version of pp-ResNetWhich version of the model was used is not stated.
  • Version of deconv-ResNetWhich version of the model was used is not stated.
  • Version of GNN (DeepGCN chemical shift predictor)Which version of the model was used is not stated.
  • Version of GBT structure ranking modelWhich version of the model was used is not stated.
  • Version of KDE signal unaliasing modelWhich version of the model was used is not stated.

About this article

Record aix-00002, version 2, checked by a person on 2026-10-07. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY-4.0; quotations are at most 25 words. How we work · Report an error