~/aixsci
200 records · all checked

materials-chemistry/ai produced the result/arXiv 2023 · v2

GPT-4 reads band gap values from paper sentences to train better predictors

Researchers prompted GPT-4 to pull measured band gap values out of sentences from the chemistry literature, then trained neural networks on the resulting dataset and compared them with models trained on existing collections.

1. Take source sentences from an existing auto-generated database2. Extract materials, properties, values and descriptors by prompting3. Classify each extraction with follow-up prompts4. Filter to experimental band gaps of pure bulk single crystals and deduplicate5. Check extraction and classification against human annotation6. Train graph neural network band gap predictors on each dataset7. Compare models by cross validation, shared hold-out set and matbench leaderboard8. Prompt the LLM to write and run code that trains predictors

spectrum · one line per step, placed by what the step does · bright lines used AI

Accurate Prediction of Experimental Band Gaps from Large Language Model-Based Data Extraction
arXiv, 2023

doi:10.48550/arxiv.2311.13778 · record aix-00035 v2 · checked 2026-10-08

ai-resultrole of AI
AI was for
Literature synthesis, Classification, Property prediction
Model family
Large language model, Graph neural network, Gradient-boosted trees, Random forest, Linear model
Checked by
Benchmark
Code
available

The finding the paper is about came from the AI.

read as

The science is explained before the AI appears. Switch to field specialist to go straight to the method.

Assumes the discipline and goes straight to the method.

The diagram, the record and what the paper did not report are identical in both modes. Only the framing changes — never the evidence.

Introduction by AIxSci · plain language

What this research was about

A material's band gap is the energy a electron needs to jump from being bound in place to being free to carry current. It is measured in electronvolts, and it largely decides whether something behaves as a metal, a semiconductor or an insulator. Knowing the gap matters for choosing materials for solar cells, lights and transistors. Decades of measurements are scattered across tens of thousands of papers, written in prose rather than tables. Gathering them into one usable list is slow, and automated text-mining tools make mistakes: they can mix up a calculated value with a measured one, or attach a number to the wrong compound.

The researchers set out to build such a list using a large language model, and then to see whether models trained on it predicted band gaps more accurately than those trained on existing datasets.

Where AI came in

GPT-4 did the reading. It was given single sentences that an earlier text-mining tool had already flagged as mentioning band gaps, and asked, with no worked examples, to name the material, the property, the value, the unit and any descriptors. Four follow-up prompts then judged each entry: was the property really a band gap, was the material a pure bulk single crystal, what was the formula, and had the number been computed rather than measured. That second step stood in for a human curator checking each entry by hand. On 100 sentences annotated by people, the prompts were right 99% of the time they made a claim, against 81% for the earlier tool.

Two further uses of machine learning followed. Graph neural networks, which treat a compound as a network of connected atoms, were trained on the extracted data and on rival datasets, then compared on a shared set of 210 materials; the authors report a 19% drop in average error. Finally, GPT-4 was asked in plain language to write and run the code for simpler predictors, taking over work a researcher would normally script themselves.

Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.

The work

Technical · from the record

GPT-4 was prompted zero-shot to pull materials, property values and descriptors out of single-sentence excerpts that had previously been collected with ChemDataExtractor, with four follow-up prompts used to keep only experimentally measured band gaps of pure single-crystalline bulk materials. The filtered result was 39391 extractions covering 2733 distinct compositions, and on 100 manually annotated sentences the prompt-based extraction reached 99% precision against 81% for ChemDataExtractor. Graph neural networks were then trained on each candidate dataset and compared on a shared hold-out set of 210 materials, using matbench_expt_gap labels as ground truth; the paper reports a 19% reduction in mean absolute error on that test set, attributed in the Results section to the dataset combining the new extractions with matbench_expt_gap. The authors also prompted the LLM to generate and execute code that trained logistic regression, gradient boosted trees and random forest models on the extracted data.

How AI was used

An LLM (GPT-4) was run over sentences released with an existing ChemDataExtractor-generated band gap dataset. A first zero-shot prompt extracted, per sentence, each material together with the property name, value, unit and any material or property descriptors. Four further prompts were then run on each extracted entry to ask whether the property is genuinely a band gap, whether the material is a pure non-doped bulk single crystal, what the chemical formula is, and whether the value was numerically calculated. Rule-based inclusion criteria then kept entries with correct prompt formatting, no evidence of numerical computation, no evidence against pure bulk single-crystalline form, units and property names consistent with band gap, and values between 0 and 20 eV; compositions with several extractions were reduced to their median. Extraction and classification prompts were scored for precision and recall against human annotation of randomly selected sentences. Message-passing graph neural networks with matscholar_el node features, an embedding size of 64 and two hidden layers of size 164 and 64 were trained for 1000 epochs with Huber loss, ensembled as the mean of 10 random initialisations, and fitted separately on each candidate training dataset for comparison by 5-fold cross validation, on a shared hold-out set of materials common to all datasets, and on the matbench leaderboard. Finally, the LLM was prompted in natural language to load and manipulate the extracted data and to generate and execute code training logistic regression, gradient boosted trees and random forest band gap models.

The shape of the work

Structural · the record, drawn

ACQUISITIONINFERENCEINFERENCEPREPARATIONVALIDATIONTRAININGVALIDATIONTRAINING12345678AIAIAIAIAITake sourcesentences from anexisting auto-ge…Extractmaterials,properties, valu…Classify eachextraction withfollow-up promptsFilter toexperimental bandgaps of pure bul…Check extractionandclassification a…Train graphneural networkband gap predict…Compare models bycross validation,shared hold-out …Prompt the LLM towrite and runcode that trains…↤ conventional algorithm↤ manual curation↤ physical experiment↤ expert judgement
AI stepNo AI↤ what the AI stood in for
1Acquisition
no AI

Take source sentences from an existing auto-generated database

Obtaining raw data, whether by measurement, download or retrieval.

We use the sentences in the dataset from Dong&Cole as the source of text.where the paper describes this · verbatim
in the paper
2Inference
AI

Extract materials, properties, values and descriptors by prompting

Running a trained model over new data to predict, classify or score. The AI stood in for conventional algorithm.

our approach uses a series of zero-shot prompting of LLMs specifically tailored to identify and extract materials and their propertieswhere the paper describes this · verbatim
in the paper
3Inference
AI

Classify each extraction with follow-up prompts

Running a trained model over new data to predict, classify or score. The AI stood in for manual curation.

For each extracted entry, we run four follow-up prompts to askwhere the paper describes this · verbatim
in the paper
4Preparation
no AI

Filter to experimental band gaps of pure bulk single crystals and deduplicate

Cleaning, filtering, normalising or labelling data already obtained.

we filter the data for experimentally measured electronic band gaps of pure single-crystalline bulk materialswhere the paper describes this · verbatim
in the paper
5Validation
no AI

Check extraction and classification against human annotation

Testing outputs against ground truth.

First we manually annotated 100 randomly selected sentences and evaluated the precision and recall across our various prompts.where the paper describes this · verbatim
in the paper
6Training
AI

Train graph neural network band gap predictors on each dataset

Fitting model parameters, including fine-tuning an existing model. The AI stood in for physical experiment.

We train a band gap property prediction model on each of the datasets in Sec. 2where the paper describes this · verbatim
in the paper
7Validation
AI

Compare models by cross validation, shared hold-out set and matbench leaderboard

Testing outputs against ground truth.

we perform a second evaluation on a shared hold-out test set, consisting of 210 materialswhere the paper describes this · verbatim
in the paper
8Training
AI

Prompt the LLM to write and run code that trains predictors

Fitting model parameters, including fine-tuning an existing model. The AI stood in for expert judgement.

also trains and compares logistic regression, gradient boosted trees and random forest models for band gap predictionwhere the paper describes this · verbatim
in the paper

What the record says

Technical · every part carries its own basis

+ in the paper~ our reading− not reported

How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.

~Role of AI
AI produced the resultour reading

The LLM prompts produced the dataset that the paper's conclusions are about, and the reported band gap values on new compositions come from the trained graph neural networks

+What the AI was for
We use GPT-4 as the LLM.where the paper describes this · verbatim
+How it was taught
Zero-shotSupervisedin the paper
+Models named
GPT-4 · Off the shelfMessage-passing graph neural network (implemented in Jraph) · Trained from scratchLogistic regression · Trained from scratchGradient boosted trees · Trained from scratchRandom forest · Trained from scratchin the paper
+How results were checked
Benchmarkin the paper
we evaluated our model on the matbench leaderboard for predicting experimental band gapswhere the paper describes this · verbatim
+Code · weights · data
code availableweights not reporteddata not reportedin the paper
We share a colab notebook demonstrating this at https://github.com/google-research/google-research/tree/master/matsciwhere the paper describes this · verbatim
+Compute
not reportedin the paper

What this paper did not report

Technical · absence is published deliberately

Reported as not stated — 10 items
  • Trained model weightsWhether the trained model is available is not stated.
  • DataWhether the data are available is not stated.
  • ComputeThe hardware or time used is not stated.
  • How many were testedThe paper gives no count of what was tested.
  • Version of GPT-4Which version of the model was used is not stated.
  • Version of Message-passing graph neural network (implemented in Jraph)Which version of the model was used is not stated.
  • Version of Logistic regressionWhich version of the model was used is not stated.
  • Version of Gradient boosted treesWhich version of the model was used is not stated.
  • Version of Random forestWhich version of the model was used is not stated.
  • What step 7 replacedThe paper gives no basis for what the AI stood in for.

About this article

Record aix-00035, version 2, checked by a person on 2026-10-08. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error