~/aixsci
200 records · all checked

materials-chemistry/ai produced the result/arXiv 2024 · v2

Machine learning ranks stable materials by predicted superconducting temperature

Researchers trained a simple statistical model on measured superconductors, then used it to estimate the critical temperature of about 153,000 known compounds from chemical composition alone. Sixty-four stable candidates were predicted above 250 K.

1. Clean SuperCon data set2. Separate ambient-pressure subset3. Query Materials Project candidates4. Generate composition-based features5. Train query-aware ridge models and predict SuperCon Tc6. Evaluate prediction error against measured Tc7. Predict Tc across Materials Project8. Filter and rank high-Tc candidates

spectrum · one line per step, placed by what the step does · bright lines used AI

High-Tc superconductor candidates proposed by machine learning
arXiv, 2024

doi:10.48550/arxiv.2406.14524 · record aix-00146 v2 · checked 2026-10-09

ai-resultrole of AI
AI was for
Property prediction
Model family
Linear model
Checked by
Held-out13661 tested
Code
available

The finding the paper is about came from the AI.

read as

The science is explained before the AI appears. Switch to field specialist to go straight to the method.

Assumes the discipline and goes straight to the method.

The diagram, the record and what the paper did not report are identical in both modes. Only the framing changes — never the evidence.

Introduction by AIxSci · plain language

What this research was about

A superconductor carries electricity with no resistance at all, but only below a certain temperature, known as the critical temperature, or Tc. For most materials that temperature is very low, so the search is on for ones that work closer to room temperature. The difficulty is that Tc cannot be read off a chemical formula. It emerges from how electrons and the vibrations of the atomic lattice interact, and calculating it from first principles is costly. Measuring it means making the material and cooling it, one compound at a time, which limits how much of the space of possible compounds anyone can check.

The researchers set out to estimate Tc from chemical composition alone, with no information about how the atoms are arranged, and then to apply that estimate across a large public database of known and computed materials. Compounds predicted to have a high Tc, and calculated to be thermodynamically stable, were kept and ranked.

Where AI came in

The learning part was a ridge regression, a standard method for fitting a straight-line relationship between numbers while discouraging extreme fitted values. Rather than one model for everything, a fresh model was fitted for each material being asked about, trained only on its ten closest matches in a cleaned set of experimentally measured superconductors drawn from the SuperCon database. Closeness was judged using 147 numerical descriptors generated from each formula, built from statistics of the properties of the elements present and from the proportions in which they appear.

Accuracy was checked by predicting Tc for materials held back from training and comparing with the measured values, including against simpler baselines. The models then supplied the predicted Tc for every candidate in the Materials Project database. Those predictions stand in for measurements that were not made: no superconductivity experiments were carried out on the proposed candidates, so their reported temperatures, including the highest at 316 K for LiCuF4, exist only as model outputs.

Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.

The work

Technical · from the record

The authors predicted superconducting critical temperature from chemical composition alone by training a ridge regression model separately for each query material on its ten nearest neighbours in the curated SuperCon data set. On out-of-sample SuperCon test materials the similarity-based models gave mean absolute errors of about 5 K across the 0–250 K range, and about 3 K for leave-one-out predictions after removing samples with large spreads in their feature-weight products. Applying the approach to about 153k Materials Project materials and keeping those within 0.030 eV/atom of the convex hull, the ambient-pressure model placed sixty-four materials above 250 K, thirty-four of which have DFT-computed band gaps below 1 eV; the highest prediction was 316 K for LiCuF4. No superconductivity measurements were performed on the proposed candidates.

How AI was used

SuperCon entries were cleaned by averaging repeated measurements and removing contentious, single-element, ten-element and arbitrarily doped stoichiometries, and a separate ambient-pressure set was formed by removing samples that SuperCon2 indicated were measured under applied pressure. Each composition was encoded as 147 Matminer features built from statistics of elemental properties and stoichiometric norms, with no structural information. For a given query material, the Euclidean nearest neighbours in the training set were retrieved and a ridge regression model was fitted on those neighbours alone, with the regularisation strength chosen by cross-validation on that training subset and the matrix inversion done by Cholesky decomposition in scikit-learn; absolute values of predictions were taken. Learning curves over similarity-selected versus random training samples, and against k-nearest-neighbour regression baselines, were used to set n=10, and leave-one-out predictions were made for every SuperCon sample. The same procedure was then run over the Materials Project set, after which predictions with large feature-weight product spreads, materials above the convex hull threshold, and in a second pass materials with larger band gaps, were discarded before ranking by predicted Tc.

The shape of the work

Structural · the record, drawn

PREPARATIONPREPARATIONACQUISITIONREPRESENTATIONTRAININGVALIDATIONINFERENCESCREENING12345678AIAIClean SuperCondata setSeparateambient-pressuresubsetQuery MaterialsProjectcandidatesGeneratecomposition-basedfeaturesTrain query-awareridge models andpredict SuperCon…Evaluateprediction erroragainst measured…Predict Tc acrossMaterials ProjectFilter and rankhigh-Tccandidates↤ physical experiment↤ physical experiment
AI stepNo AI↤ what the AI stood in for
1Preparation
no AI

Clean SuperCon data set

Cleaning, filtering, normalising or labelling data already obtained.

We cleaned the data set by assigning to stoichiometries with multiple Tc measurements their mean values.where the paper describes this · verbatim
in the paper
2Preparation
no AI

Separate ambient-pressure subset

Cleaning, filtering, normalising or labelling data already obtained.

These were removed from SuperCon to create a separate data setwhere the paper describes this · verbatim
in the paper
3Acquisition
no AI

Query Materials Project candidates

Obtaining raw data, whether by measurement, download or retrieval.

We apply our similarity-based ML method to ∼ 153k samples listed in the Materials Project databasewhere the paper describes this · verbatim
in the paper
4Representation
no AI

Generate composition-based features

Encoding data into features, descriptors, embeddings or graphs.

147 features were generated for each sample from its composition using the materials informatics Python library Matminerwhere the paper describes this · verbatim
in the paper
5Training
AI

Train query-aware ridge models and predict SuperCon Tc

Fitting model parameters, including fine-tuning an existing model. The AI stood in for physical experiment.

are then used to train a ridge regression model, from which the test sample’s Tc is predictedwhere the paper describes this · verbatim
in the paper
6Validation
no AI

Evaluate prediction error against measured Tc

Testing outputs against ground truth.

the learning curves show the prediction error (mean absolute error, MAE) on the test set after training on n training sampleswhere the paper describes this · verbatim
in the paper
7Inference
AI

Predict Tc across Materials Project

Running a trained model over new data to predict, classify or score. The AI stood in for physical experiment.

Each sample’s predictions under implicit pressure (n=10) and ambient pressure (n=10) were made by training on its nearest neighborswhere the paper describes this · verbatim
in the paper
8Screening
no AI

Filter and rank high-Tc candidates

Reducing a candidate set by filtering or ranking, in a single pass.

Those with computed energies above their convex hulls of greater than 0.030 eV/atom are also disregarded as being thermodynamically unstable.where the paper describes this · verbatim
in the paper

What the record says

Technical · every part carries its own basis

+ in the paper~ our reading− not reported

How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.

~Role of AI
AI produced the resultour reading

The reported candidates and their Tc rankings exist only as outputs of the trained ridge regression models; no measurement supports them.

+What the AI was for
Training of query-aware similarity-based ridge regression models on experimental SuperCon datawhere the paper describes this · verbatim
+Model families
Linear modelin the paper
+How it was taught
Supervisedin the paper
+Models named
Similarity-based ridge regression, ambient-pressure model · Trained from scratchSimilarity-based ridge regression, implicit-pressure model · Trained from scratchk-nearest neighbors regression (baseline) · Trained from scratchin the paper
+How results were checked
Held-out13661 testedin the paper
predictions under unknown/ambient pressure are made for each of the 13,661/13,624 materialswhere the paper describes this · verbatim
+Code · weights · data
code availableweights not reporteddata availablein the paper
Refer to https://zenodo.org/records/14052692 for: Python code to generate ML features and to implement our similarity-based ML modelswhere the paper describes this · verbatim
+Compute
Several milliseconds per material on a laptop (12th gen. Intel i7-1260P, 12 cores, 2.10 GHz); the full Materials Project scan consumed ~2 node hours.in the paper

What this paper did not report

Technical · absence is published deliberately

Reported as not stated — 4 items
  • Trained model weightsWhether the trained model is available is not stated.
  • Version of Similarity-based ridge regression, ambient-pressure modelWhich version of the model was used is not stated.
  • Version of Similarity-based ridge regression, implicit-pressure modelWhich version of the model was used is not stated.
  • Version of k-nearest neighbors regression (baseline)Which version of the model was used is not stated.

About this article

Record aix-00146, version 2, checked by a person on 2026-10-09. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error