~/aixsci
200 records · all checked

materials-chemistry/ai produced the result/arXiv 2026 · v2

Zero-padding lets one neural network encode crystals with differing element counts

Researchers padded the symmetry description of inorganic crystals with zeros so a single variational autoencoder could handle compositions with different numbers of elements. The network learned to reconstruct and generate candidate structures; pretrained potentials then relaxed and scored them.

1. Assemble and partition benchmark datasets2. Encode structures as padded Wyckoff representation3. Train a single VAE on all compositions4. Sample latent space and decode candidates5. Discard crystallographically invalid decodings6. Build 3D crystal structures with Pyxtal7. Relax structures and predict energy with ML potentials8. Compute energy above hull and apply stability threshold

spectrum · one line per step, placed by what the step does · bright lines used AI

A Padding Method for Enhanced Encoding of Inorganic Structures with Varying Chemical Compositions
arXiv, 2026

doi:10.48550/arxiv.2605.30743 · record aix-00111 v2 · checked 2026-10-08

ai-resultrole of AI
AI was for
Candidate generation, Property prediction
Model family
Autoencoder, Graph neural network
Checked by
Held-out
Code
not reported

The finding the paper is about came from the AI.

read as

The science is explained before the AI appears. Switch to field specialist to go straight to the method.

Assumes the discipline and goes straight to the method.

The diagram, the record and what the paper did not report are identical in both modes. Only the framing changes — never the evidence.

Introduction by AIxSci · plain language

What this research was about

Inorganic crystals are built from a repeating arrangement of atoms. Crystallographers describe that arrangement compactly using a space group, which records the symmetries of the pattern, and Wyckoff positions, which say which symmetry-equivalent sites the atoms sit on. This shorthand is far smaller than listing every coordinate, which makes it attractive for machine learning. The trouble is that it is not a fixed size. A compound made of three elements needs fewer entries than one made of five, so a neural network expecting a fixed input cannot swallow both. One common workaround is to train a separate model for each number of elements, which splits the data and the effort.

The authors set out to remove that split. They append zeros to the Wyckoff matrix of materials with fewer elements, so that every structure in a dataset is written at the same size, and a single model can be trained on all of them. They tested this on perovskites, on a set of samples from the Materials Project database, and on a set of proton-conducting ceramic electrolytes, comparing against an unpadded version of the same model.

Where AI came in

The central model is a variational autoencoder, a network that squeezes each input into a compact numerical summary, called a latent space, and then tries to rebuild the original from it. Trained without labels on the padded descriptions, it learned both to reconstruct Wyckoff positions and space groups and, by sampling and nudging points in that latent space, to propose new combinations. The reconstruction scores were checked on a held-out fifth of the data against the unpadded baseline. Pre-computed elemental feature vectors from an existing model, CGCNN, supplied part of the input description.

Decoded proposals that broke crystallographic rules were thrown out, and the rest were expanded into full three-dimensional structures by a conventional software library. Two off-the-shelf machine learning potentials, CHGNet and M3GNet, then settled the atoms into relaxed positions and predicted each structure's energy. The record notes this stood in for density functional theory, the quantum-mechanical calculation normally used for the same job but far more costly. Those predicted energies fed a thermodynamic screen that kept structures close enough to the known stable phases to count as plausible.

Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.

The work

Technical · from the record

The authors modify the Wyckoff-position encoding used by a crystal variational autoencoder by padding the Wyckoff matrix with zeros so that materials with different numbers of chemical elements share one matrix size, which lets a single VAE be trained per dataset instead of one model per composition count. Trained on composition-balanced subsets of perov-5 (18,928 perovskites), mp-20 (45,231 Materials Project samples) and a proton-conductor dataset of more than 4,793 entries, the padded model reconstructed Wyckoff positions at 99.9%, 94.8% and 88% and space groups at 100%, 92.8% and 91.1% respectively, which the paper reports as up to 5.3% higher Wyckoff accuracy than the Wyckoff VAE baseline on proton-conductor data, although its space group accuracy there fell below the baseline's 5_chem figure. Structures sampled from the latent space were converted to 3D with Pyxtal, relaxed with pretrained CHGNet or M3GNet, and screened on energy above hull; for perov-5 3_chem at 0.08 eV/atom with CHGNet the padded model yielded 170 structures below threshold against 104 for the baseline, reported as a 63.5% increase, while at 0.5 eV/atom with CHGNet it returned 564 against the baseline's 574.

How AI was used

Chemical compositions, space group numbers and Wyckoff position dictionaries were converted into a fixed-dimension representation combining one-hot atomic numbers, normalised stoichiometric ratios, pre-computed CGCNN elemental feature vectors, a one-hot space group vector and a Wyckoff site/multiplicity matrix, with shorter sequences zero-padded to the longest Wyckoff position length in each batch. A single variational autoencoder was then trained per dataset on these representations using KL divergence, space group, reconstruction and Wyckoff position losses, with an 80/20 training-validation split, learning rate 2x10-4, batch size 256, 1,000 epochs and the RMSprop optimizer. For generation, latent vectors were perturbed with Gaussian noise and decoded into Wyckoff positions and space group encodings; decodings that violated crystallographic rules were discarded, the remainder were expanded into 3D structures with Pyxtal, relaxed and assigned total energies by the pretrained CHGNet or M3GNet machine learning potentials in place of DFT, and filtered by energy above the convex hull computed with pymatgen against Materials Project phase diagrams at thresholds of 0.08, 0.1 and 0.5 eV/atom. For each implementation, 1,000 randomly selected generated samples were put through the stability check. The Wyckoff VAE served as the comparison model for both reconstruction and generation.

The shape of the work

Structural · the record, drawn

PREPARATIONREPRESENTATIONTRAININGGENERATIONSCREENINGGENERATIONSIMULATIONSCREENING12345678AIAIAIAssemble andpartitionbenchmark datase…Encode structuresas padded WyckoffrepresentationTrain a singleVAE on allcompositionsSample latentspace and decodecandidatesDiscardcrystallographicallyinvalid decodingsBuild 3D crystalstructures withPyxtalRelax structuresand predictenergy with ML p…Compute energyabove hull andapply stability …↤ exhaustive search↤ simulation
AI stepNo AI↤ what the AI stood in for
1Preparation
no AI

Assemble and partition benchmark datasets

Cleaning, filtering, normalising or labelling data already obtained.

we only use data with balanced and sufficiently large amounts in each benchmark data set for training the VAE modelwhere the paper describes this · verbatim
in the paper
2Representation
no AI

Encode structures as padded Wyckoff representation

Encoding data into features, descriptors, embeddings or graphs.

This is achieved by appending "0" values to the Wyckoff matrix for material structures with fewer chemical elementswhere the paper describes this · verbatim
in the paper
3Training
AI

Train a single VAE on all compositions

Fitting model parameters, including fine-tuning an existing model.

The Wyckoff representations are used to train a single VAE model.where the paper describes this · verbatim
in the paper
4Generation
AI

Sample latent space and decode candidates

Producing candidate objects that did not previously exist. The AI stood in for exhaustive search.

To generate new candidate materials, we sample from the latent space of the trained VAE.where the paper describes this · verbatim
in the paper
5Screening
no AI

Discard crystallographically invalid decodings

Reducing a candidate set by filtering or ranking, in a single pass.

The decoded Wyckoff positions are validated for physical consistency, ensuring they adhere to crystallographic ruleswhere the paper describes this · verbatim
in the paper
6Generation
no AI

Build 3D crystal structures with Pyxtal

Producing candidate objects that did not previously exist.

Valid Wyckoff positions and space groups are converted into three-dimensional crystal structures using the Pyxtal library.where the paper describes this · verbatim
in the paper
7Simulation
AI

Relax structures and predict energy with ML potentials

Numerical or physics simulation, including where a learned surrogate replaces it. The AI stood in for simulation.

The generated 3D structures are relaxed using a pretrained machine learning potentialwhere the paper describes this · verbatim
in the paper
8Screening
no AI

Compute energy above hull and apply stability threshold

Reducing a candidate set by filtering or ranking, in a single pass.

Stability is assessed by calculating the energy above the convex hullwhere the paper describes this · verbatim
in the paper

What the record says

Technical · every part carries its own basis

+ in the paper~ our reading− not reported

How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.

~Role of AI
AI produced the resultour reading

The paper's reported outcomes are the structures produced by the VAE and the stability counts assigned to them by machine learning potentials; without the models there is no result.

+What the AI was for
A VAE framework is employed, comprising an encoder and a decoderwhere the paper describes this · verbatim
+Model families
+How it was taught
Self-supervisedin the paper
+Models named
Padded Wyckoff VAE (this work) · Trained from scratchWyckoff VAE (baseline) · Trained from scratchCGCNN (elemental embeddings) · Off the shelfCHGNet · Off the shelfM3GNet · Off the shelfin the paper
+How results were checked
Held-outin the paper
we split the dataset into 80% training and 20% validation, with random shuffling appliedwhere the paper describes this · verbatim
−Code · weights · data
code not reportedweights not reporteddata not reportednot reported
−Compute
not reportednot reported

What this paper did not report

Technical · absence is published deliberately

Reported as not stated — 11 items
  • CodeWhether the code is available is not stated.
  • Trained model weightsWhether the trained model is available is not stated.
  • DataWhether the data are available is not stated.
  • ComputeThe hardware or time used is not stated.
  • How many were testedThe paper gives no count of what was tested.
  • Version of Padded Wyckoff VAE (this work)Which version of the model was used is not stated.
  • Version of Wyckoff VAE (baseline)Which version of the model was used is not stated.
  • Version of CGCNN (elemental embeddings)Which version of the model was used is not stated.
  • Version of CHGNetWhich version of the model was used is not stated.
  • Version of M3GNetWhich version of the model was used is not stated.
  • What step 3 replacedThe paper gives no basis for what the AI stood in for.

About this article

Record aix-00111, version 2, checked by a person on 2026-10-08. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error