structural-biology/ai produced the result/The Journal of Physical Chemistry B 2024 · v2
Language model trained on protein sequences predicts shape and stability of disordered proteins
Researchers fine-tuned ProtBERT, a model pretrained on protein sequences, to predict three physical properties of intrinsically disordered proteins straight from their amino acid sequence, using values first computed by molecular simulation.
spectrum · one line per step, placed by what the step does · bright lines used AI
IDP-Bert: Predicting Properties of Intrinsically Disordered Proteins Using Large Language Models
The Journal of Physical Chemistry B, 2024
doi:10.1021/acs.jpcb.4c02507 · record aix-00186 v2 · checked 2026-10-09
- AI was for
- Property prediction, Simulation surrogate
- Model family
- Transformer, Protein language model, Multilayer perceptron, Clustering
- Checked by
- Held-out522 tested
- Code
- available
The finding the paper is about came from the AI.
What this research was about
Most proteins fold into a fixed shape, and that shape explains what they do. But a large group of proteins, called intrinsically disordered proteins, never settle into one form. They flicker between many loose arrangements. That makes them hard to study, because there is no single structure to measure or picture. Instead, scientists describe them with averaged physical quantities: how spread out the chain is on average, how long it takes for the two ends to forget their relative positions, and how much heat it takes to warm the molecule. These are usually obtained by simulating the chain's motion on a computer, atom group by atom group, which costs a great deal of computing time.
The researchers wanted to skip that cost. They asked whether those three properties could be read off the amino acid sequence alone, by a model that had already learned the statistical patterns of protein sequences.
Where AI came in
The team started from ProtBERT, a transformer model pretrained on a large collection of protein sequences. A transformer learns which parts of a sequence matter to which other parts, in the same way text models do. On top of it they attached two fully connected layers that output a number, and trained three separate versions, one for each property. The training targets were not laboratory measurements but values produced by coarse-grained molecular dynamics simulations for 2,585 disordered protein sequences taken from the DisProt database.
So the model stood in for the simulation. Once trained, it was given held-out sequences it had not seen and asked to produce the same property values directly, and the predictions were compared against the simulated values and against previously published models. Across five repeated splits, the mean R scores were 0.9881 for radius of gyration, 0.9713 for decorrelation time and 0.9645 for heat capacity. The researchers also examined the model's internal representations, reducing them to two dimensions and grouping them into 15 clusters, then retrained using those clusters to define the data splits.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record
The authors fine-tuned ProtBERT, a pretrained protein language model, to predict three properties of intrinsically disordered proteins — Radius of Gyration, end-to-end Decorrelation Time and Heat Capacity — directly from amino acid sequence, using labels computed by coarse-grained molecular dynamics simulations for 2585 IDPs. Across five repeated train-test splits the mean R scores were 0.9881 for Radius of Gyration, 0.9713 for Decorrelation Time and 0.9645 for Heat Capacity. For the same three properties the best models reported by Patel et al. gave mean R values of 0.9859, 0.7828 and 0.8823. The authors also examined penultimate-layer features with t-SNE and KMeans clustering, and retrained using cluster-based data splits.
How AI was used
Property labels came from coarse-grained molecular dynamics simulations in LAMMPS using an improved HPS model, run at 300 K for 2585 IDP sequences drawn from DisProt version 9.0; labels were Yeo–Johnson transformed and sequence redundancy was checked with MMSeqs2 at a 0.8 similarity threshold. Sequences were tokenized per amino acid and passed to ProtBERT, pretrained on the Big Fantastic Database, in a modified configuration (16 hidden layers of size 256, 16 attention heads, 0.15 hidden dropout) with a two-layer fully connected regression head using ReLU activations. Three separate models were trained, one per property, with MSE loss over 5 epochs using AdamW and a ReduceLROnPlateau scheduler. Training-set sequence-length buckets were oversampled by duplication to balance bucket sizes. The fine-tuned models were then run over held-out sequences to predict property values, and penultimate-layer features were extracted for t-SNE visualisation and KMeans clustering, with the resulting 15 clusters used to define alternative train/validation/test splits for retraining.
The shape of the work
Structural · the record, drawn
no AI
Molecular dynamics simulation of IDP properties
Numerical or physics simulation, including where a learned surrogate replaces it.
the MD simulations were conducted by Patel et al., employing the hydropathy scale (HPS) model via the LAMMPS simulation packagewhere the paper describes this · verbatim
no AI
Transform labels, split data, tokenize and oversample
Cleaning, filtering, normalising or labelling data already obtained.
The labels were subjected to a data transformation method known as Yeo–Johnson transformation to reduce skewness and non-normality in their distributionwhere the paper describes this · verbatim
AI
Fine-tune ProtBERT with a regression head per property
Fitting model parameters, including fine-tuning an existing model. The AI stood in for simulation.
These tokens were fed to IDP-Bert and three separate models were trained (one for each IDP property)where the paper describes this · verbatim
AI
Predict properties for held-out sequences
Running a trained model over new data to predict, classify or score. The AI stood in for simulation.
The time taken to obtain the predictions for the entire testing data set of 522 protein sequences was recordedwhere the paper describes this · verbatim
no AI
Evaluate against simulated ground truth and published baselines
Testing outputs against ground truth.
The average R values, along with their corresponding Standard Deviations (SD), were computed after training the model five times, each with a distinct train-test splitwhere the paper describes this · verbatim
AI
Inspect learned representations and re-split by cluster
Extracting understanding from model behaviour. Its result feeds back into an earlier step.
15 clusters, which were then randomly assigned to training, validation, or test setswhere the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
The reported result is the predictive model itself and the property values it produces; the paper's claims are entirely about the fine-tuned language model's predictions
IDP-Bert utilizes ProtBERT, a transformer-based model with 16 attention heads and 30 hidden layers as its backbonewhere the paper describes this · verbatim
The average R values, along with their corresponding Standard Deviations (SD), were computed after training the model five times, each with a distinct train-test splitwhere the paper describes this · verbatim
The necessary information containing the codes and data for downstream tasks used in this study is available herewhere the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- Trained model weightsWhether the trained model is available is not stated.
- ComputeThe hardware or time used is not stated.
- Version of IDP-Bert (ProtBERT backbone with two fully connected regression layers)Which version of the model was used is not stated.
- Version of KMeans clustering on t-SNE embeddingsWhich version of the model was used is not stated.
- What step 6 replacedThe paper gives no basis for what the AI stood in for.
About this article
Record aix-00186, version 2, checked by a person on 2026-10-09. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error