materials-chemistry/ai produced the result/arXiv 2023 · v2
GPT-4 reads band gap values from paper sentences to train better predictors
Researchers prompted GPT-4 to pull measured band gap values out of sentences from the chemistry literature, then trained neural networks on the resulting dataset and compared them with models trained on existing collections.
spectrum · one line per step, placed by what the step does · bright lines used AI
Accurate Prediction of Experimental Band Gaps from Large Language Model-Based Data Extraction
arXiv, 2023
doi:10.48550/arxiv.2311.13778 · record aix-00035 v2 · checked 2026-10-08
- AI was for
- Literature synthesis, Classification, Property prediction
- Model family
- Large language model, Graph neural network, Gradient-boosted trees, Random forest, Linear model
- Checked by
- Benchmark
- Code
- available
The finding the paper is about came from the AI.
What this research was about
A material's band gap is the energy a electron needs to jump from being bound in place to being free to carry current. It is measured in electronvolts, and it largely decides whether something behaves as a metal, a semiconductor or an insulator. Knowing the gap matters for choosing materials for solar cells, lights and transistors. Decades of measurements are scattered across tens of thousands of papers, written in prose rather than tables. Gathering them into one usable list is slow, and automated text-mining tools make mistakes: they can mix up a calculated value with a measured one, or attach a number to the wrong compound.
The researchers set out to build such a list using a large language model, and then to see whether models trained on it predicted band gaps more accurately than those trained on existing datasets.
Where AI came in
GPT-4 did the reading. It was given single sentences that an earlier text-mining tool had already flagged as mentioning band gaps, and asked, with no worked examples, to name the material, the property, the value, the unit and any descriptors. Four follow-up prompts then judged each entry: was the property really a band gap, was the material a pure bulk single crystal, what was the formula, and had the number been computed rather than measured. That second step stood in for a human curator checking each entry by hand. On 100 sentences annotated by people, the prompts were right 99% of the time they made a claim, against 81% for the earlier tool.
Two further uses of machine learning followed. Graph neural networks, which treat a compound as a network of connected atoms, were trained on the extracted data and on rival datasets, then compared on a shared set of 210 materials; the authors report a 19% drop in average error. Finally, GPT-4 was asked in plain language to write and run the code for simpler predictors, taking over work a researcher would normally script themselves.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record
GPT-4 was prompted zero-shot to pull materials, property values and descriptors out of single-sentence excerpts that had previously been collected with ChemDataExtractor, with four follow-up prompts used to keep only experimentally measured band gaps of pure single-crystalline bulk materials. The filtered result was 39391 extractions covering 2733 distinct compositions, and on 100 manually annotated sentences the prompt-based extraction reached 99% precision against 81% for ChemDataExtractor. Graph neural networks were then trained on each candidate dataset and compared on a shared hold-out set of 210 materials, using matbench_expt_gap labels as ground truth; the paper reports a 19% reduction in mean absolute error on that test set, attributed in the Results section to the dataset combining the new extractions with matbench_expt_gap. The authors also prompted the LLM to generate and execute code that trained logistic regression, gradient boosted trees and random forest models on the extracted data.
How AI was used
An LLM (GPT-4) was run over sentences released with an existing ChemDataExtractor-generated band gap dataset. A first zero-shot prompt extracted, per sentence, each material together with the property name, value, unit and any material or property descriptors. Four further prompts were then run on each extracted entry to ask whether the property is genuinely a band gap, whether the material is a pure non-doped bulk single crystal, what the chemical formula is, and whether the value was numerically calculated. Rule-based inclusion criteria then kept entries with correct prompt formatting, no evidence of numerical computation, no evidence against pure bulk single-crystalline form, units and property names consistent with band gap, and values between 0 and 20 eV; compositions with several extractions were reduced to their median. Extraction and classification prompts were scored for precision and recall against human annotation of randomly selected sentences. Message-passing graph neural networks with matscholar_el node features, an embedding size of 64 and two hidden layers of size 164 and 64 were trained for 1000 epochs with Huber loss, ensembled as the mean of 10 random initialisations, and fitted separately on each candidate training dataset for comparison by 5-fold cross validation, on a shared hold-out set of materials common to all datasets, and on the matbench leaderboard. Finally, the LLM was prompted in natural language to load and manipulate the extracted data and to generate and execute code training logistic regression, gradient boosted trees and random forest band gap models.
The shape of the work
Structural · the record, drawn
no AI
Take source sentences from an existing auto-generated database
Obtaining raw data, whether by measurement, download or retrieval.
We use the sentences in the dataset from Dong&Cole as the source of text.where the paper describes this · verbatim
AI
Extract materials, properties, values and descriptors by prompting
Running a trained model over new data to predict, classify or score. The AI stood in for conventional algorithm.
our approach uses a series of zero-shot prompting of LLMs specifically tailored to identify and extract materials and their propertieswhere the paper describes this · verbatim
AI
Classify each extraction with follow-up prompts
Running a trained model over new data to predict, classify or score. The AI stood in for manual curation.
For each extracted entry, we run four follow-up prompts to askwhere the paper describes this · verbatim
no AI
Filter to experimental band gaps of pure bulk single crystals and deduplicate
Cleaning, filtering, normalising or labelling data already obtained.
we filter the data for experimentally measured electronic band gaps of pure single-crystalline bulk materialswhere the paper describes this · verbatim
no AI
Check extraction and classification against human annotation
Testing outputs against ground truth.
First we manually annotated 100 randomly selected sentences and evaluated the precision and recall across our various prompts.where the paper describes this · verbatim
AI
Train graph neural network band gap predictors on each dataset
Fitting model parameters, including fine-tuning an existing model. The AI stood in for physical experiment.
We train a band gap property prediction model on each of the datasets in Sec. 2where the paper describes this · verbatim
AI
Compare models by cross validation, shared hold-out set and matbench leaderboard
Testing outputs against ground truth.
we perform a second evaluation on a shared hold-out test set, consisting of 210 materialswhere the paper describes this · verbatim
AI
Prompt the LLM to write and run code that trains predictors
Fitting model parameters, including fine-tuning an existing model. The AI stood in for expert judgement.
also trains and compares logistic regression, gradient boosted trees and random forest models for band gap predictionwhere the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
The LLM prompts produced the dataset that the paper's conclusions are about, and the reported band gap values on new compositions come from the trained graph neural networks
We use GPT-4 as the LLM.where the paper describes this · verbatim
we evaluated our model on the matbench leaderboard for predicting experimental band gapswhere the paper describes this · verbatim
We share a colab notebook demonstrating this at https://github.com/google-research/google-research/tree/master/matsciwhere the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- Trained model weightsWhether the trained model is available is not stated.
- DataWhether the data are available is not stated.
- ComputeThe hardware or time used is not stated.
- How many were testedThe paper gives no count of what was tested.
- Version of GPT-4Which version of the model was used is not stated.
- Version of Message-passing graph neural network (implemented in Jraph)Which version of the model was used is not stated.
- Version of Logistic regressionWhich version of the model was used is not stated.
- Version of Gradient boosted treesWhich version of the model was used is not stated.
- Version of Random forestWhich version of the model was used is not stated.
- What step 7 replacedThe paper gives no basis for what the AI stood in for.
About this article
Record aix-00035, version 2, checked by a person on 2026-10-08. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error