~/aixsci
200 records · all checked

structural-biology/ai produced the result/Briefings in Bioinformatics 2023 · v2

A compact neural network learns to generate shape-shifting protein structures

Researchers trained a small generative model on simulations of floppy proteins, then used it to produce new backbone shapes. The model stood in for further simulation, turning out 50,000 conformations of one protein in about 50 seconds.

1. Run all-atom MD simulations of target proteins2. Shuffle and split trajectories into train/validation/test3. Encode conformations as atom-level graphs4. Train graph-encoder/transformer-decoder variational model per system5. Sample latent space to generate and interpolate backbones6. Refine generated backbones to all-atom conformations7. Evaluate ensembles against MD, REMD and experimental observables

spectrum · one line per step, placed by what the step does · bright lines used AI

Phanto-IDP: compact model for precise intrinsically disordered protein backbone generation and enhanced sampling
Briefings in Bioinformatics, 2023

doi:10.1093/bib/bbad429 · record aix-00083 v2 · checked 2026-10-08

ai-resultrole of AI
AI was for
Candidate generation, Simulation surrogate
Model family
Graph neural network, Transformer, Autoencoder
Checked by
Held-out12500 tested
Code
available

The finding the paper is about came from the AI.

read as

The science is explained before the AI appears. Switch to field specialist to go straight to the method.

Assumes the discipline and goes straight to the method.

The diagram, the record and what the paper did not report are identical in both modes. Only the framing changes — never the evidence.

Introduction by AIxSci · plain language

What this research was about

Most proteins fold into a fairly fixed shape, and that shape explains what they do. But a large group, called intrinsically disordered proteins, never settle. They wriggle between many shapes at once, so describing one of them means describing a whole crowd of structures, known as an ensemble, together with how often each shape appears. The usual way to obtain that crowd is molecular dynamics simulation, which computes the physical motion of every atom step by step. It is faithful but slow, and it needs large computing facilities to run long enough for the protein to explore its options.

The researchers set out to build a small model that could learn the shapes a given protein visits and then produce more of them without further simulation. They worked with ten disordered proteins and four ordered ones, all under 200 residues, each simulated for a microsecond beforehand to supply the training material.

Where AI came in

The model, called Phanto-IDP, is a variational autoencoder: one half squeezes a structure down to a compact numerical code, the other half expands a code back into a structure, and because the codes are arranged smoothly, fresh codes yield fresh structures. The squeezing half treats the protein backbone as a graph of linked atoms; the expanding half is a transformer, the architecture behind language models, which here outputs atomic coordinates. It was trained from scratch separately for each protein, using no labels beyond the simulated structures themselves, with 142,000 parameters.

In place of running more simulation, the trained model was fed random seeds to generate ensembles, and codes for two chosen shapes were blended step by step to trace paths between them. Training on one protein took 30.91 hours on a single graphics card; generating 50,000 of its shapes then took about 50 seconds. Accuracy was checked on 12,500 held-out structures and against simulated and published experimental measurements, with other generative models for comparison.

Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.

The work

Technical · from the record

Phanto-IDP is a variational autoencoder for protein backbones: a graph convolutional encoder reads backbone N, CA and C atoms as a graph, a variational layer compresses them, and a three-block transformer decoder outputs Cartesian coordinates. It was trained separately on each target system using 1 μs all-atom MD trajectories from which 50 000 conformations were extracted, covering 10 intrinsically disordered proteins and four ordered proteins under 200 residues. On a held-out set of 12 500 conformations, average backbone reconstruction RMSD was 0.511 Å for RS1, 0.885 Å for PaaA2 and 2.714 Å for α-synuclein, and the model generated 50 000 α-synuclein conformations in about 50 s on a Tesla V100 after 30.91 h of training. When trained on a short MD trajectory, 5.62% of the generated structures were in a folded state, compared with 18.3% sampled by replica exchange molecular dynamics, and latent-space interpolation between two selected conformations produced continuous transition paths.

How AI was used

Converged all-atom MD trajectories were simulated with the ESFF1 force field and OPC3 solvent model for 1 μs per protein, and 50 000 conformations were extracted at 0.02 ns intervals, shuffled and split into training, validation and test sets (50%/25%/25% for the main systems; 80% training for the short AAQAA3 trajectory). Each conformation was reduced to its N, CA and C backbone atoms and converted into an atom-level crystal graph with one-hot atom-type node features, a 30-nearest-neighbour adjacency and SE(3) equivariant edge features built from local frames computed with the mylddt toolset. A graph convolutional encoder with edge gating produced atom embeddings, a fully connected layer reduced dimensionality to residue-level features, and a variational layer produced mean and log-variance matrices from which latent codes were drawn by reparameterisation at sampling temperature T = 0.02; a decoder of three transformer blocks, each with a self-attention layer and an update module, emitted backbone Cartesian coordinates. Training used a frame-aligned point error reconstruction loss computed in Gram–Schmidt local frames plus a KL term whose weight was increased dynamically to avoid posterior collapse, for 400 epochs per system on one Tesla V100. The trained decoder was then run on seeds from a normal distribution to generate new conformation ensembles and on linear interpolations between the latent codes of two selected conformations; generated backbones were passed through a side-chain refinement step, and ensembles were compared with MD and REMD trajectories and with published experimental observables.

The shape of the work

Structural · the record, drawn

SIMULATIONPREPARATIONREPRESENTATIONTRAININGGENERATIONPREPARATIONVALIDATION1234567AIAIRun all-atom MDsimulations oftarget proteinsShuffle and splittrajectories intotrain/validation/testEncodeconformations asatom-level graphsTraingraph-encoder/transformer-decodervariational mode…Sample latentspace to generateand interpolate …Refine generatedbackbones toall-atom conform…Evaluateensembles againstMD, REMD and exp…↤ simulation↤ simulation
AI stepNo AI↤ what the AI stood in for
1Simulation
no AI

Run all-atom MD simulations of target proteins

Numerical or physics simulation, including where a learned surrogate replaces it.

We simulated the proteins with force field ESFF1 and solvent model OPC3 for 1 μswhere the paper describes this · verbatim
in the paper
2Preparation
no AI

Shuffle and split trajectories into train/validation/test

Cleaning, filtering, normalising or labelling data already obtained.

we shuffled the trajectories from MD simulation and split the dataset into the training set, evaluation set and test setwhere the paper describes this · verbatim
in the paper
3Representation
no AI

Encode conformations as atom-level graphs

Encoding data into features, descriptors, embeddings or graphs.

We preserve backbone atoms (N, CA, C) for each input conformation and construct a crystal graph at the atomic level.where the paper describes this · verbatim
in the paper
4Training
AI

Train graph-encoder/transformer-decoder variational model per system

Fitting model parameters, including fine-tuning an existing model. The AI stood in for simulation.

we trained Phanto-IDP for 400 epochs to ensure sufficient convergencewhere the paper describes this · verbatim
in the paper
5Generation
AI

Sample latent space to generate and interpolate backbones

Producing candidate objects that did not previously exist. The AI stood in for simulation.

the well-trained model is able to generate a large number of unseen protein conformations with seeds from normal distributionwhere the paper describes this · verbatim
in the paper
6Preparation
no AI

Refine generated backbones to all-atom conformations

Cleaning, filtering, normalising or labelling data already obtained.

the refinement process might not fully reconstruct the side-chain distributions of the corresponding conformationswhere the paper describes this · verbatim
in the paper
7Validation
no AI

Evaluate ensembles against MD, REMD and experimental observables

Testing outputs against ground truth.

we assessed the quality of conformation reconstruction using RMSD and dihedral angle distributions as indexeswhere the paper describes this · verbatim
in the paper

What the record says

Technical · every part carries its own basis

+ in the paper~ our reading− not reported

How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.

~Role of AI
AI produced the resultour reading

The conformational ensembles the paper reports and analyses are produced by the trained generative model; the findings are about the model's output

~What the AI was for
a graph-based encoder to extract protein features and a transformer-based decoder combined with variational samplingwhere the paper describes this · verbatim
~How it was taught
Self-supervisedour reading
~Models named
Phanto-IDP · Trained from scratchAE · Trained from scratchVAE · Trained from scratchFoldingDiff · Off the shelfEigenFold · Off the shelfour reading
+How results were checked
Held-out12500 testedin the paper
On the test set, which consists of 12 500 conformations that were not encountered during the model training processwhere the paper describes this · verbatim
+Code · weights · data
code availableweights availabledata not reportedin the paper
The code for training Phanto-IDP and weights for generating conformation ensembles are available at https://github.com/HFChenLab/PhantoIDP.where the paper describes this · verbatim
+Compute
One Tesla V100; training took 30.91 h for α-synuclein (400 epochs) and no longer than 2 days for proteins under 200 residues; 50 000 conformations generated in approximately 50 s for α-synucleinin the paper

What this paper did not report

Technical · absence is published deliberately

Reported as not stated — 6 items
  • DataWhether the data are available is not stated.
  • Version of Phanto-IDPWhich version of the model was used is not stated.
  • Version of AEWhich version of the model was used is not stated.
  • Version of VAEWhich version of the model was used is not stated.
  • Version of FoldingDiffWhich version of the model was used is not stated.
  • Version of EigenFoldWhich version of the model was used is not stated.

About this article

Record aix-00083, version 2, checked by a person on 2026-10-08. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error