materials-chemistry/ai in a supporting role/RSC Advances 2025 · v2
Simulations map three sodium–bismuth compounds, then models predict their heat-to-electricity performance
Researchers used quantum-mechanical simulations to work out the structure, vibrations and heat-to-electricity behaviour of three sodium–bismuth compounds. A random forest and a small neural network were then trained on those simulated results to predict one compound's efficiency figure.
spectrum · one line per step, placed by what the step does · bright lines used AI
Data-driven exploration of Na–Bi compounds: a first-principles and machine learning approach to topological thermoelectrics
RSC Advances, 2025
doi:10.1039/d5ra05888k · record aix-00101 v2 · checked 2026-10-08
- AI was for
- Property prediction, Simulation surrogate
- Model family
- Random forest, Multilayer perceptron
- Checked by
- Held-out25 tested
- Code
- not reported
AI processed or interpreted data, but the main finding does not rest on it.
What this research was about

A thermoelectric material turns a difference in temperature into electricity, and it does so without moving parts. How well it manages this is summed up in a single number called the figure of merit, usually written ZT. That number depends on three things pulling against each other: how strong a voltage the material builds up across a temperature gap, how easily it carries electrical current, and how readily it leaks heat. A good material needs the first two high and the last low, and in most solids improving one spoils another. Which combinations of elements strike a useful balance is therefore hard to guess from chemical intuition alone.
The researchers looked at three compounds made from sodium and bismuth: tetragonal NaBi, hexagonal NaBi3 and cubic Na3Bi, each a different packing arrangement of the same two elements. Using density functional theory, a method that solves approximate quantum-mechanical equations for the electrons in a crystal, they computed the compounds' relaxed structures, electronic states, lattice vibrations and transport behaviour. They report a peak ZT of 0.53 for cubic Na3Bi at 500 K.
Where AI came in
The machine learning came after the physics, not before it. From the transport calculations on cubic Na3Bi, the team sampled 101 points at each of three temperatures and tabulated the voltage response, the electrical conductivity and the thermal conductivity as inputs, with ZT as the quantity to be predicted. Two models were fitted separately at each temperature on a three-quarters training, one-quarter testing split: a random forest of 100 decision trees, and a small fully connected neural network of two layers with 16 units each. The random forest was the more accurate of the two, especially at the lowest temperature.
What the models stood in for was the step that turns transport coefficients into ZT, which is otherwise a closed-form calculation, and, by extension, the expensive simulations that related compounds would require. The authors also applied SHAP, a technique that apportions a model's prediction among its inputs, to the random forest. It ranked the voltage response as the largest contributor at all three temperatures. None of the reported physical findings rests on the models.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record

Density functional theory with spin–orbit coupling, density functional perturbation theory and Boltzmann transport theory were used to compute the structural, electronic, vibrational and thermoelectric properties of tetragonal NaBi, hexagonal NaBi3 and cubic Na3Bi. A random forest and a two-layer neural network were then trained on the DFT-derived Seebeck coefficient, electrical conductivity and thermal conductivity of cubic Na3Bi to predict the figure of merit ZT at 100 K, 500 K and 950 K, with SHAP used to quantify each feature's contribution. The paper reports a peak ZT of 0.53 for cubic Na3Bi at 500 K, states that the random forest was more accurate than the neural network, especially at 100 K, and reports that SHAP ranked the Seebeck coefficient as the largest contributor to predicted ZT at all three temperatures.
How AI was used
Transport coefficients for cubic Na3Bi were taken from Boltzmann transport calculations on DFT band structures at 100 K, 500 K and 950 K, sampling 101 chemical potentials between −1.5 eV and +1.5 eV to give 101 data points per temperature. The Seebeck coefficient, electrical conductivity and thermal conductivity were used as input features and ZT as the regression target. Two supervised models were fitted separately at each temperature on a 75%/25% train–test split: a random forest with 100 trees selected by manual grid-based tuning under a mean squared error criterion, and a fully connected network of two dense layers of 16 neurons with ReLU activations, trained with the Adam optimiser and early stopping with patience of 20 epochs, with inputs z-score normalised for the network. The trained models were applied to the held-out split to predict ZT, and R and RMSE together with parity plots were computed for both models. SHAP values were then computed on the random forest to attribute predicted ZT to the three input features at each temperature.
The shape of the work
Structural · the record, drawn
no AI
Relax structures and compute electronic structure with DFT
Numerical or physics simulation, including where a learned surrogate replaces it.
Structural optimization was carried out using the Broyden–Fletcher–Goldfarb–Shanno (BFGS) approach, employing a 12 × 12 × 8 Monkhorst–Pack grid.where the paper describes this · verbatim
no AI
Compute thermoelectric transport coefficients and ZT
Numerical or physics simulation, including where a learned surrogate replaces it.
The thermoelectric characteristics are evaluated using the semi-classical Boltzmann transport theory, implemented via the BoltzTrap softwarewhere the paper describes this · verbatim
no AI
Build ZT dataset from transport coefficients and split it
Cleaning, filtering, normalising or labelling data already obtained.
By taking samples from 101 chemical potentials between −1.5 eV and +1.5 eV, we got 101 data points for each temperature.where the paper describes this · verbatim
AI
Train random forest and neural network ZT regressors per temperature
Fitting model parameters, including fine-tuning an existing model. The AI stood in for conventional algorithm.
Each model was trained independently for each temperature using a 75%/25% train-test split.where the paper describes this · verbatim
AI
Predict ZT for held-out test points
Running a trained model over new data to predict, classify or score. The AI stood in for conventional algorithm.
Parity plots (Fig. 10) also correspond to test predictions, showing the correlation between predicted and true ZT values.where the paper describes this · verbatim
no AI
Score predictions against DFT-derived ZT
Testing outputs against ground truth.
Performance metrics (R and RMSE) of Random Forest and Neural Network models in predicting ZT for cubic Na3Biwhere the paper describes this · verbatim
AI
Attribute ZT predictions to transport features with SHAP
Extracting understanding from model behaviour. The AI stood in for expert judgement.
We employed SHAP to interpret the RF model.where the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
The paper's physical findings (Dirac semimetal character, phonon stability, ZT values) come from DFT and Boltzmann transport; the ML models were trained on those DFT outputs as surrogate predictors of ZT and as a feature-importance analysis, so no reported finding depends on them.
we trained supervised ML models [Random Forest (RF) and Neural Network (NN)] on thermoelectric results from DFTwhere the paper describes this · verbatim
We divided the dataset into two parts: 75% for training (76 points) and 25% for testing (25 points).where the paper describes this · verbatim
All data are available in the data repository ZENODO at: 10.5281/zenodo.15499220where the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- CodeWhether the code is available is not stated.
- Trained model weightsWhether the trained model is available is not stated.
- ComputeThe hardware or time used is not stated.
- Version of Random Forest regressor (100 trees)Which version of the model was used is not stated.
- Version of Fully connected neural network (two dense layers, 16 neurons each, ReLU)Which version of the model was used is not stated.
About this article
Record aix-00101, version 2, checked by a person on 2026-10-08. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error