materials-chemistry/ai produced the result/Journal of Chemical Information and Modeling 2026 · v2
Statistical model of 5,441 perovskite solar cells predicts efficiency and suggests recipes
Researchers fitted a probability model to records of 5,441 perovskite solar cells drawn from a public database. The same fitted model grouped the devices, predicted their efficiency, invented plausible new recipes and worked backwards to synthesis conditions for a chosen efficiency.
spectrum · one line per step, placed by what the step does · bright lines used AI
Density Estimation Based on Mixtures of Gaussians for Perovskite Solar Cells Modeling
Journal of Chemical Information and Modeling, 2026
doi:10.1021/acs.jcim.5c02017 · record aix-00040 v2 · checked 2026-10-08
- AI was for
- Property prediction, Candidate generation, Experimental design
- Model family
- Clustering, Gradient-boosted trees
- Checked by
- Held-out1051 tested
- Code
- not reported
The finding the paper is about came from the AI.
What this research was about
Perovskite solar cells are made from a class of crystalline materials that can be deposited from liquid solution rather than cut from a silicon block. That makes them cheap to prepare, but it also means a finished cell depends on a long list of choices: which elements go into the crystal, the solvents used, the spinning and heating steps, the size of the cell. Each combination gives a different power conversion efficiency, the share of incoming sunlight the cell turns into electricity. The number of possible combinations is far larger than anyone can test in a laboratory, and published results are scattered, inconsistent and often incomplete.
The researchers worked instead from the Perovskite Database Project, an open collection of device records taken from the literature. They narrowed it to cells made by one common route, using the solvents DMF and DMSO, and kept the synthesis settings, the cell area, the band gap and the measured efficiency. Rather than build a separate tool for each question, they set out to describe all of these quantities together as one statistical distribution, and then ask different questions of that single description.
Where AI came in
The central piece of machine learning is a Gaussian mixture model: a way of describing a cloud of data as a handful of overlapping bell-shaped blobs, each with its own centre and spread. Fitted to the device records, it treats the synthesis settings and the resulting efficiency as one joint pattern rather than as inputs and an output. Because the chemical formula of a perovskite is not a number, compositions were first turned into element counts and then compressed into four coordinates by a technique called locally linear embedding, which arranges items so that near neighbours stay near. A separate method, t-SNE, was used to view the resulting space.
From the one fitted model the researchers read off five blob centres as typical device types, calculated expected efficiencies for devices held back from fitting, and drew fresh synthetic recipes with matching efficiencies. They also refitted it on data with values deliberately removed, using a variant that keeps incomplete records instead of discarding them. For working backwards, samples from the model supplied starting guesses which an optimiser refined against a separate predictor, XGBoost, trained to estimate efficiency from the same descriptors. On regression with those descriptors, the authors report that XGBoost did better than the mixture model.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record
The study fits a Gaussian mixture model to a 5441-device subset of the Perovskite Database Project, treating synthesis parameters, a four-dimensional embedding of the perovskite composition and power conversion efficiency as a single joint density. The one fitted density is then used for five tasks: identifying clusters, PCE regression, sampling synthetic device records, fitting when values are missing, and inferring synthesis conditions for a target PCE. For the inverse task, samples from the model's conditional distribution seed a derivative-free optimisation against an XGBoost forward model; against target PCEs this gave an RMSE of 1.52, compared with 3.32 for random-start optimisation. With the same input descriptors, the authors report that XGBoost outperformed the Gaussian mixture approach on regression.
How AI was used
Device records from the Perovskite Database Project were filtered to spin-coated DMF/DMSO cells with complete numeric entries and PCE above 10%, leaving a tabular set of synthesis variables, cell area, band gap and efficiency. Perovskite compositions were encoded as sum-pooled one-hot element vectors of dimension 133 and then reduced with Locally Linear Embedding using cosine distance to four components, with the number of components and neighbours chosen by cross-validation and the resulting space inspected with t-SNE. A Gaussian mixture model was fitted by expectation-maximisation over the nine descriptors plus PCE, with the number of components selected by ten-fold cross-validation on a development split under the one-standard-error rule, and an XGBoost ensemble was trained on the same descriptors as a forward predictor. The same fitted density supplied conditional expectations for PCE regression, component means as cluster prototypes, ancestral sampling for synthetic records, and conditional distributions of synthesis parameters given a material and a target PCE. For inverse design, conditional samples provided initial guesses that a derivative-free optimiser refined by minimising squared error between the target PCE and the XGBoost output, with the best configuration retained; a random-start variant served as the comparison. A separate experiment refitted the mixture on data with artificially introduced missing values using the ECM algorithm via an R package, alongside mean imputation, median imputation and listwise deletion.
The shape of the work
Structural · the record, drawn
no AI
Filter device records to a modelling subset
Cleaning, filtering, normalising or labelling data already obtained.
Once these filters were applied, we obtained a data set of 5441 observations.where the paper describes this · verbatim
AI
Embed perovskite composition in four dimensions
Encoding data into features, descriptors, embeddings or graphs. The AI stood in for conventional algorithm.
We then employed Locally Linear Embedding (LLE) with the cosine distance to reduce this dimensionality.where the paper describes this · verbatim
AI
Fit joint density and forward predictor
Fitting model parameters, including fine-tuning an existing model.
This approach favors a more parsimonious model, resulting in 5 Normal full distributions.where the paper describes this · verbatim
AI
Predict PCE from the conditional density
Running a trained model over new data to predict, classify or score. The AI stood in for physical experiment.
Once we have the distribution, we are able to perform tasks such as regressionwhere the paper describes this · verbatim
AI
Read cluster prototypes from mixture components
Extracting understanding from model behaviour.
Our analysis identified five main clusterswhere the paper describes this · verbatim
AI
Sample synthetic device records
Producing candidate objects that did not previously exist. The AI stood in for physical experiment.
Samples from a GMM, joint as well conditional, can be easily drawn by implementing ancestral sampling.where the paper describes this · verbatim
AI
Infer synthesis conditions for a target PCE
Iterative search over a space. The AI stood in for conventional algorithm.
From this, we sample three initial guesses (X 0) to seed the optimization.where the paper describes this · verbatim
AI
Refit the density on data with missing entries
Fitting model parameters, including fine-tuning an existing model. The AI stood in for conventional algorithm.
In this paper we employ the Expectation Conditional Maximization (ECM) algorithmwhere the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
Every reported finding — clusters, PCE regression, synthetic samples, missing-data robustness, inferred synthesis conditions — is an output of the fitted GMM or the XGBoost forward model; no non-AI result is reported.
We employ Gaussian Mixture Models (GMMs), a pragmatic and interpretable choice well-suited for the scarce, low-dimensional tabular datawhere the paper describes this · verbatim
The data set (5441 devices) was randomly split into a development set (80%) and a test set (20%).where the paper describes this · verbatim
In this work, we have used the open-access data set from the Perovsktie Database Project.where the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- CodeWhether the code is available is not stated.
- Trained model weightsWhether the trained model is available is not stated.
- ComputeThe hardware or time used is not stated.
- Version of Gaussian Mixture Model (GMM, fitted by EM)Which version of the model was used is not stated.
- Version of MGMM (GMM fitted by ECM for missing data, R package MGMM)Which version of the model was used is not stated.
- Version of XGBoostWhich version of the model was used is not stated.
- Version of Locally Linear Embedding (LLE)Which version of the model was used is not stated.
- Version of t-SNEWhich version of the model was used is not stated.
- What step 3 replacedThe paper gives no basis for what the AI stood in for.
- What step 5 replacedThe paper gives no basis for what the AI stood in for.
About this article
Record aix-00040, version 2, checked by a person on 2026-10-08. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error