structural-biology/ai produced the result/Briefings in Bioinformatics 2023 · v2
A compact neural network learns to generate shape-shifting protein structures
Researchers trained a small generative model on simulations of floppy proteins, then used it to produce new backbone shapes. The model stood in for further simulation, turning out 50,000 conformations of one protein in about 50 seconds.
spectrum · one line per step, placed by what the step does · bright lines used AI
Phanto-IDP: compact model for precise intrinsically disordered protein backbone generation and enhanced sampling
Briefings in Bioinformatics, 2023
doi:10.1093/bib/bbad429 · record aix-00083 v2 · checked 2026-10-08
- AI was for
- Candidate generation, Simulation surrogate
- Model family
- Graph neural network, Transformer, Autoencoder
- Checked by
- Held-out12500 tested
- Code
- available
The finding the paper is about came from the AI.
What this research was about
Most proteins fold into a fairly fixed shape, and that shape explains what they do. But a large group, called intrinsically disordered proteins, never settle. They wriggle between many shapes at once, so describing one of them means describing a whole crowd of structures, known as an ensemble, together with how often each shape appears. The usual way to obtain that crowd is molecular dynamics simulation, which computes the physical motion of every atom step by step. It is faithful but slow, and it needs large computing facilities to run long enough for the protein to explore its options.
The researchers set out to build a small model that could learn the shapes a given protein visits and then produce more of them without further simulation. They worked with ten disordered proteins and four ordered ones, all under 200 residues, each simulated for a microsecond beforehand to supply the training material.
Where AI came in
The model, called Phanto-IDP, is a variational autoencoder: one half squeezes a structure down to a compact numerical code, the other half expands a code back into a structure, and because the codes are arranged smoothly, fresh codes yield fresh structures. The squeezing half treats the protein backbone as a graph of linked atoms; the expanding half is a transformer, the architecture behind language models, which here outputs atomic coordinates. It was trained from scratch separately for each protein, using no labels beyond the simulated structures themselves, with 142,000 parameters.
In place of running more simulation, the trained model was fed random seeds to generate ensembles, and codes for two chosen shapes were blended step by step to trace paths between them. Training on one protein took 30.91 hours on a single graphics card; generating 50,000 of its shapes then took about 50 seconds. Accuracy was checked on 12,500 held-out structures and against simulated and published experimental measurements, with other generative models for comparison.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record
Phanto-IDP is a variational autoencoder for protein backbones: a graph convolutional encoder reads backbone N, CA and C atoms as a graph, a variational layer compresses them, and a three-block transformer decoder outputs Cartesian coordinates. It was trained separately on each target system using 1 μs all-atom MD trajectories from which 50 000 conformations were extracted, covering 10 intrinsically disordered proteins and four ordered proteins under 200 residues. On a held-out set of 12 500 conformations, average backbone reconstruction RMSD was 0.511 Å for RS1, 0.885 Å for PaaA2 and 2.714 Å for α-synuclein, and the model generated 50 000 α-synuclein conformations in about 50 s on a Tesla V100 after 30.91 h of training. When trained on a short MD trajectory, 5.62% of the generated structures were in a folded state, compared with 18.3% sampled by replica exchange molecular dynamics, and latent-space interpolation between two selected conformations produced continuous transition paths.
How AI was used
Converged all-atom MD trajectories were simulated with the ESFF1 force field and OPC3 solvent model for 1 μs per protein, and 50 000 conformations were extracted at 0.02 ns intervals, shuffled and split into training, validation and test sets (50%/25%/25% for the main systems; 80% training for the short AAQAA3 trajectory). Each conformation was reduced to its N, CA and C backbone atoms and converted into an atom-level crystal graph with one-hot atom-type node features, a 30-nearest-neighbour adjacency and SE(3) equivariant edge features built from local frames computed with the mylddt toolset. A graph convolutional encoder with edge gating produced atom embeddings, a fully connected layer reduced dimensionality to residue-level features, and a variational layer produced mean and log-variance matrices from which latent codes were drawn by reparameterisation at sampling temperature T = 0.02; a decoder of three transformer blocks, each with a self-attention layer and an update module, emitted backbone Cartesian coordinates. Training used a frame-aligned point error reconstruction loss computed in Gram–Schmidt local frames plus a KL term whose weight was increased dynamically to avoid posterior collapse, for 400 epochs per system on one Tesla V100. The trained decoder was then run on seeds from a normal distribution to generate new conformation ensembles and on linear interpolations between the latent codes of two selected conformations; generated backbones were passed through a side-chain refinement step, and ensembles were compared with MD and REMD trajectories and with published experimental observables.
The shape of the work
Structural · the record, drawn
no AI
Run all-atom MD simulations of target proteins
Numerical or physics simulation, including where a learned surrogate replaces it.
We simulated the proteins with force field ESFF1 and solvent model OPC3 for 1 μswhere the paper describes this · verbatim
no AI
Shuffle and split trajectories into train/validation/test
Cleaning, filtering, normalising or labelling data already obtained.
we shuffled the trajectories from MD simulation and split the dataset into the training set, evaluation set and test setwhere the paper describes this · verbatim
no AI
Encode conformations as atom-level graphs
Encoding data into features, descriptors, embeddings or graphs.
We preserve backbone atoms (N, CA, C) for each input conformation and construct a crystal graph at the atomic level.where the paper describes this · verbatim
AI
Train graph-encoder/transformer-decoder variational model per system
Fitting model parameters, including fine-tuning an existing model. The AI stood in for simulation.
we trained Phanto-IDP for 400 epochs to ensure sufficient convergencewhere the paper describes this · verbatim
AI
Sample latent space to generate and interpolate backbones
Producing candidate objects that did not previously exist. The AI stood in for simulation.
the well-trained model is able to generate a large number of unseen protein conformations with seeds from normal distributionwhere the paper describes this · verbatim
no AI
Refine generated backbones to all-atom conformations
Cleaning, filtering, normalising or labelling data already obtained.
the refinement process might not fully reconstruct the side-chain distributions of the corresponding conformationswhere the paper describes this · verbatim
no AI
Evaluate ensembles against MD, REMD and experimental observables
Testing outputs against ground truth.
we assessed the quality of conformation reconstruction using RMSD and dihedral angle distributions as indexeswhere the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
The conformational ensembles the paper reports and analyses are produced by the trained generative model; the findings are about the model's output
a graph-based encoder to extract protein features and a transformer-based decoder combined with variational samplingwhere the paper describes this · verbatim
On the test set, which consists of 12 500 conformations that were not encountered during the model training processwhere the paper describes this · verbatim
The code for training Phanto-IDP and weights for generating conformation ensembles are available at https://github.com/HFChenLab/PhantoIDP.where the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- DataWhether the data are available is not stated.
- Version of Phanto-IDPWhich version of the model was used is not stated.
- Version of AEWhich version of the model was used is not stated.
- Version of VAEWhich version of the model was used is not stated.
- Version of FoldingDiffWhich version of the model was used is not stated.
- Version of EigenFoldWhich version of the model was used is not stated.
About this article
Record aix-00083, version 2, checked by a person on 2026-10-08. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error