~/aixsci
200 records · all checked

materials-chemistry/ai produced the result/arXiv 2026 · v2

Deep learning model predicts polymer heat-softening across 48,208 designed candidates

Researchers built Periodic-TDL, a model that reads a polymer's repeating unit as a shape and predicts its properties. It forecast glass transition temperatures for an enumerated library of 48,208 polymers; three were then made and measured.

1. Assemble unlabelled and labelled polymer datasets2. Build periodic Vietoris–Rips representations3. Self-supervised pretraining of the HSMP encoder4. Fine-tune and benchmark on nine property tasks5. Enumerate systematically substituted polymer library6. Predict Tg across the library with a ten-model ensemble7. Matched-pair statistical analysis of predicted Tg shifts8. Synthesise and characterise polymers for trend comparison

spectrum · one line per step, placed by what the step does · bright lines used AI

Periodic Topological Deep Learning for Polymer Design and Discovery
arXiv, 2026

doi:10.48550/arxiv.2605.26833 · record aix-00117 v2 · checked 2026-10-08

ai-resultrole of AI
AI was for
Property prediction
Model family
Graph neural network, Transformer, Multilayer perceptron, Gradient-boosted trees, Linear model
Checked by
Experimental6 tested, 6 worked
Code
available

The finding the paper is about came from the AI.

read as

The science is explained before the AI appears. Switch to field specialist to go straight to the method.

Assumes the discipline and goes straight to the method.

The diagram, the record and what the paper did not report are identical in both modes. Only the framing changes — never the evidence.

Introduction by AIxSci · plain language

What this research was about

Polymers are long chains built by repeating a small chemical unit over and over. One of the most practical things to know about such a chain is its glass transition temperature, or Tg: the point at which a rigid, glassy plastic softens into something rubbery. Tg decides whether a material suits a hot car dashboard or a cold drinks bottle. Predicting it from chemistry alone is awkward. The property belongs to the whole tangled chain, not to any single molecule, yet the information a chemist starts with is usually just the repeating unit. Measuring Tg in the laboratory means synthesising each candidate first, which limits how many ideas can be tried.

The team set out to describe a repeating unit in a way that keeps its periodic, endlessly repeating nature, and to learn property predictions from that description. They then used the resulting model to sweep through a large set of related polymers, looking at how two specific chemical swaps shifted the predicted Tg, and checked the predicted directions against measurements.

Where AI came in

Each repeating unit, written as a text string, was turned into three-dimensional coordinates with standard chemistry software, then into a nested geometric construction that records which atoms fall within given distances of one another. A neural network passed messages across that structure. It was first trained on roughly a million unlabelled polymers using tasks invented from the data itself, such as guessing an atom's surroundings, so that it learned general chemical regularities without needing measured properties. It was then tuned on nine datasets of known properties and compared against a range of other published models on identical data splits.

For the main application, ten copies of the model were trained on separate slices of the experimental Tg data and their averaged output was applied to the enumerated library of 48,208 polymers, to polymers from the literature, and to three newly made ones. Here the AI stood in for laboratory synthesis and measurement: every Tg value across the library is a prediction, not an observation. The reported mean shifts of 57.0 °C and 54.6 °C for swapping an ester group for an amide, and 15.4 °C and 13.0 °C for adding a methyl group to the backbone, are statistics over those predictions. Experiment entered at the end, as six matched pairs whose measured directions were compared with the predicted ones.

Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.

The work

Technical · from the record

The authors built Periodic-TDL, a deep learning model that represents a polymer as a periodic Vietoris–Rips simplicial complex and encodes it with a hierarchical simplicial message-passing network, pretrained on about one million unlabelled polymers from PI1M and fine-tuned on nine property datasets. On five-fold splits shared with all baselines, Periodic-TDL had the lowest test RMSE on seven of the nine tasks, while PerioGT gave slightly lower errors on bandgap (chain) and Tg; adding chemical descriptors and a residual correction reduced Tg RMSE by 6.5% relative to the base model. Applied to an enumerated library of 48,208 substituted acrylate and acrylamide polymers, the model predicted mean Tg increases of 57.0 °C and 54.6 °C for the two ester-to-amide comparisons and 15.4 °C and 13.0 °C for the two backbone α-methylation comparisons. Three polymers were synthesised and measured by DSC and combined with four literature pairs; the predicted direction of ΔTg matched experiment in all six matched pairs.

How AI was used

Polymer repeating units given as pSMILES were converted to 3D coordinates with RDKit and UFF optimisation over all cyclic rearrangements of the repeating unit, from which a periodic distance matrix and a nested Vietoris–Rips filtration at 2.0, 3.0 and 4.0 Å were built; simplices carried RDKit atom and bond descriptors plus Forman–Ricci curvature features. A hierarchical simplicial message-passing encoder with multi-head updates and cross-scale refinement was pretrained on the PI1M corpus using three self-supervised objectives adapted from GROVER (atom context, bond context and multi-label functional-group prediction), then fine-tuned with a two-stage schedule and a two-layer regression head on nine property targets under five-fold cross-validation, with baselines retrained or fine-tuned on identical splits. Two further configurations stacked the fine-tuned predictions with predictors on frozen embeddings, Mordred descriptors and Morgan fingerprints, and added a similarity-based residual correction. For the Tg analysis, ten models were trained on disjoint folds of the experimental Tg data and their averaged predictions were applied to an enumerated library of substituted monomers, to literature polymers and to three newly synthesised polymers.

The shape of the work

Structural · the record, drawn

ACQUISITIONREPRESENTATIONTRAININGTRAININGGENERATIONINFERENCEINTERPRETATIONEXPERIMENT12345678AIAIAIAssembleunlabelled andlabelled polymer…Build periodicVietoris–RipsrepresentationsSelf-supervisedpretraining ofthe HSMP encoderFine-tune andbenchmark on nineproperty tasksEnumeratesystematicallysubstituted poly…Predict Tg acrossthe library witha ten-model ense…Matched-pairstatisticalanalysis of pred…Synthesise andcharacterisepolymers for tre…↤ simulation↤ physical experiment
AI stepNo AI↤ what the AI stood in for
1Acquisition
no AI

Assemble unlabelled and labelled polymer datasets

Obtaining raw data, whether by measurement, download or retrieval.

we used the PI1M dataset, a benchmark resource in polymer informatics comprising approximately one million polymers represented as pSMILES stringswhere the paper describes this · verbatim
in the paper
2Representation
no AI

Build periodic Vietoris–Rips representations

Encoding data into features, descriptors, embeddings or graphs.

we therefore generated 3D coordinates using RDKit by embedding the monomer and performing geometry optimization with the Universal Force Field (UFF)where the paper describes this · verbatim
in the paper
3Training
AI

Self-supervised pretraining of the HSMP encoder

Fitting model parameters, including fine-tuning an existing model.

We pretrained HSMP using three self-supervised tasks adapted from the GROVER framework, namely atom context prediction, bond context prediction, and functional group (FG) prediction.where the paper describes this · verbatim
in the paper
4Training
AI

Fine-tune and benchmark on nine property tasks

Fitting model parameters, including fine-tuning an existing model. The AI stood in for simulation.

the pretrained HSMP encoder was fine-tuned on nine supervised polymer property prediction taskswhere the paper describes this · verbatim
in the paper
5Generation
no AI

Enumerate systematically substituted polymer library

Producing candidate objects that did not previously exist.

This procedure yielded 12,052 unique monomers per family and 48,208 polymers total across the four families.where the paper describes this · verbatim
in the paper
6Inference
AI

Predict Tg across the library with a ten-model ensemble

Running a trained model over new data to predict, classify or score. The AI stood in for physical experiment.

each polymer pSMILES was processed through the HSMP encoder using all ten trained modelswhere the paper describes this · verbatim
in the paper
7Interpretation
no AI

Matched-pair statistical analysis of predicted Tg shifts

Extracting understanding from model behaviour.

We tested whether mean Δ​Tg deviated significantly from zero using two-sided one-sample t -tests.where the paper describes this · verbatim
in the paper
8Experiment
no AI

Synthesise and characterise polymers for trend comparison

Physical execution, by hand or by robot.

we synthesized three polymers that were entirely absent from the training dataset and had not been previously characterized in the experimental literaturewhere the paper describes this · verbatim
in the paper

What the record says

Technical · every part carries its own basis

+ in the paper~ our reading− not reported

How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.

~Role of AI
AI produced the resultour reading

The paper's quantitative claims about Tg shifts come from model predictions over a computationally enumerated library; the property-prediction results are themselves model outputs.

~What the AI was for
We pretrained HSMP using three self-supervised tasks adapted from the GROVER framework, namely atom context prediction, bond context prediction, and functional group (FG) prediction.where the paper describes this · verbatim
~How it was taught
Self-supervisedSupervisedTransfer / fine-tuningour reading
~Models named
Periodic-TDL (HSMP encoder) · Trained from scratchPerioGT · Fine-tunedpolyBERT · Fine-tunedTransPolymer · Fine-tunedMolCLR (GCN) · Fine-tunedMolCLR (GIN) · Fine-tunedTransChem · Fine-tunedMMPolymer · Fine-tunedpolyGNN · Trained from scratchMorgan (NN) · Trained from scratchXGBoost regressor on Mordred descriptors (Periodic-TDL +Chem component) · Trained from scratchour reading
+How results were checked
Experimental6 tested, 6 workedin the paper
Predicted directional trends were in agreement with experiment in all six caseswhere the paper describes this · verbatim
+Code · weights · data
code availableweights not reporteddata availablein the paper
and the pretraining and fine-tuning pipelines is publicly available at https://github.com/yasharthy/Periodic-TDL.where the paper describes this · verbatim
+Compute
not reportedin the paper

What this paper did not report

Technical · absence is published deliberately

Reported as not stated — 14 items
  • Trained model weightsWhether the trained model is available is not stated.
  • ComputeThe hardware or time used is not stated.
  • Version of Periodic-TDL (HSMP encoder)Which version of the model was used is not stated.
  • Version of PerioGTWhich version of the model was used is not stated.
  • Version of polyBERTWhich version of the model was used is not stated.
  • Version of TransPolymerWhich version of the model was used is not stated.
  • Version of MolCLR (GCN)Which version of the model was used is not stated.
  • Version of MolCLR (GIN)Which version of the model was used is not stated.
  • Version of TransChemWhich version of the model was used is not stated.
  • Version of MMPolymerWhich version of the model was used is not stated.
  • Version of polyGNNWhich version of the model was used is not stated.
  • Version of Morgan (NN)Which version of the model was used is not stated.
  • Version of XGBoost regressor on Mordred descriptors (Periodic-TDL +Chem component)Which version of the model was used is not stated.
  • What step 3 replacedThe paper gives no basis for what the AI stood in for.

About this article

Record aix-00117, version 2, checked by a person on 2026-10-08. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error