~/aixsci
200 records · all checked

materials-chemistry/ai in a supporting role/Journal of Chemical Information and Modeling 2025 · v2

Oxidation potentials for 15,238 molecules assembled from quantum calculations and lab measurements

Researchers built OxPot, an open dataset of oxidation potentials for 15,238 organic molecules, using a fitted linear relationship to convert quantum-chemical calculations into potentials, then trained several machine learning models on the result.

1. Curate and filter molecules from PubChem2. Measure oxidation potentials by cyclic voltammetry3. Compute HOMO energies with DFT4. Fit E_HOMO–E_ox correlation and predict E_ox5. Derive molecular descriptors and predict solubility6. Train and evaluate ML models on OxPot7. Attribute feature importance in the trained GIN

spectrum · one line per step, placed by what the step does · bright lines used AI

Predicting Oxidation Potentials with DFT-Driven Machine Learning
Journal of Chemical Information and Modeling, 2025

doi:10.1021/acs.jcim.5c00159 · record aix-00108 v2 · checked 2026-10-08

ai-supportingrole of AI
AI was for
Property prediction
Model family
Random forest, Support vector machine, Multilayer perceptron, Graph neural network, Linear model
Checked by
Held-out
Code
not reported

AI processed or interpreted data, but the main finding does not rest on it.

read as

The science is explained before the AI appears. Switch to field specialist to go straight to the method.

Assumes the discipline and goes straight to the method.

The diagram, the record and what the paper did not report are identical in both modes. Only the framing changes — never the evidence.

Introduction by AIxSci · plain language

What this research was about

An oxidation potential is the voltage at which a molecule gives up an electron. It tells chemists how easily a compound will be oxidised, which matters for batteries, dyes, drugs and anything that must survive contact with air. Measuring it means putting the substance in solution and sweeping a voltage across an electrode while watching the current — a technique called cyclic voltammetry. That works well, but one molecule at a time. There are millions of known organic compounds, and only a small fraction have ever had their oxidation potential measured, so researchers who want to screen many candidates have little data to draw on.

The researchers set out to fill that gap by computation rather than by experiment. They drew molecules from PubChem, a public chemical database, filtered them by a set of rules, and calculated a quantum-mechanical property of each one: the energy of its highest occupied molecular orbital, or HOMO — roughly, how tightly the outermost electrons are held. That calculation used density functional theory, a standard method for approximating the behaviour of electrons. They also measured oxidation potentials themselves in the laboratory for a smaller set of molecules, to anchor the calculations to reality.

Where AI came in

The link between the calculated HOMO energies and real oxidation potentials was a straight line fitted to the authors' own laboratory measurements — a simple statistical model whose slope and intercept came from the data. That fitted line stood in for the experiment itself, turning a calculation into a predicted potential for every curated molecule. Against the measurements it was fitted to, it gave a correlation of 0.977 and a root-mean-square error of 0.064; tested on molecules from the literature that were left out of the fit, the figures were 0.963 and a mean absolute error of 0.069. An off-the-shelf open-source model, AqSolPred, supplied predicted water solubility as another column of the dataset.

The team then trained machine learning models on the finished dataset to see whether the predicted potentials could be learned from molecular structure alone, without the quantum calculation. These included a random forest, support vector machines, a multilayer perceptron, and two graph neural networks, which treat a molecule as a network of atoms joined by bonds. The Graph Isomorphism Network scored best, with a mean absolute error of 0.074. A tool called GNNExplainer was then used to ask which atomic details that model leaned on; it pointed to oxidation state, the number of attached hydrogens and the number of directly bonded neighbours.

Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.

The work

Technical · from the record

The authors built OxPot, an open-access data set of oxidation potentials for 15,238 organic molecules curated from PubChem. Oxidation potentials were assigned by computing HOMO energies with DFT (PBE0/cc-pVDZ with COSMO solvation) and applying a linear correlation fitted to the authors' own cyclic voltammetry measurements, which gave an R value of 0.977 and an RMSE of 0.064; comparison with literature-reported values for molecules outside the fit gave an R of 0.963 and an MAE of 0.069. Predicted values span −0.6 to 2.7 V, with most molecules between 0.4 and 1.4 V. Several machine learning models were then trained on the data set; among those tested, the Graph Isomorphism Network had the lowest MAE (0.074) and RMSE (0.116), and GNNExplainer applied to it identified oxidation state, number of bonded hydrogens and number of directly bonded neighbours as the most influential atom-level features.

How AI was used

Machine learning entered the work at three points. First, an empirically fitted linear regression between DFT-computed E_HOMO and cyclic-voltammetry E_ox was used to assign oxidation potentials to the PubChem-curated molecules, and an off-the-shelf open-source model, AqSolPred, supplied the solubility column alongside RDKit-derived descriptors such as topological polar surface area, hydrogen-bond donor and acceptor counts, aromatic rings and fractional Csp3. Second, classical models (Random Forest, support vector machines with linear and radial basis function kernels, and a multilayer perceptron) were trained on Morgan fingerprints, and message-passing graph neural networks — a Graph Convolutional Network and a Graph Isomorphism Network — were trained on molecular graphs with atom-level descriptors including atomic number, chirality, bonded-neighbour and bonded-hydrogen counts, formal charge, radical electrons, hybridisation, aromaticity, ring membership, TPSA, fractional Csp3, LogP, pH, oxidation state and molecular mass; models were scored with MAE and RMSE. Third, GNNExplainer was applied to the trained GIN model to rank the atom-level input features by influence on its predictions.

The shape of the work

Structural · the record, drawn

PREPARATIONEXPERIMENTSIMULATIONINFERENCEREPRESENTATIONTRAININGINTERPRETATION1234567AIAIAIAICurate and filtermolecules fromPubChemMeasure oxidationpotentials bycyclic voltammet…Compute HOMOenergies with DFTFit E_HOMO–E_oxcorrelation andpredict E_oxDerive moleculardescriptors andpredict solubili…Train andevaluate MLmodels on OxPotAttribute featureimportance in thetrained GIN↤ physical experiment↤ simulation↤ expert judgement
AI stepNo AI↤ what the AI stood in for
1Preparation
no AI

Curate and filter molecules from PubChem

Cleaning, filtering, normalising or labelling data already obtained.

A set of rules was applied to filter the molecules from PubChem, developing an enhanced data set with high prediction precision.where the paper describes this · verbatim
in the paper
2Experiment
no AI

Measure oxidation potentials by cyclic voltammetry

Physical execution, by hand or by robot.

In CV, most of the molecules underwent single-electron transfer, and we used the maximum anodic peak potential (Epa) to report the potential.where the paper describes this · verbatim
in the paper
3Simulation
no AI

Compute HOMO energies with DFT

Numerical or physics simulation, including where a learned surrogate replaces it.

the hybrid functional PBE0 paired with the cc-pVDZ basis set and the Restricted Kohn–Sham approach showed the best performancewhere the paper describes this · verbatim
in the paper
4Inference
AI

Fit E_HOMO–E_ox correlation and predict E_ox

Running a trained model over new data to predict, classify or score. The AI stood in for physical experiment.

by calculating the E HOMO values for each molecule using DFT and employing an empirically derived linear correlation with experimental E oxwhere the paper describes this · verbatim
in the paper
5Representation
AI

Derive molecular descriptors and predict solubility

Encoding data into features, descriptors, embeddings or graphs.

An open-source ML model, AqSolPred, trained on a refined version of AqSolDB, was utilized for solubility predictions.where the paper describes this · verbatim
in the paper
6Training
AI

Train and evaluate ML models on OxPot

Fitting model parameters, including fine-tuning an existing model. The AI stood in for simulation.

We trained several classical ML algorithms, including Random Forest, Support Vector Machine (SVM) with both linear and radial function kernelswhere the paper describes this · verbatim
in the paper
7Interpretation
AI

Attribute feature importance in the trained GIN

Extracting understanding from model behaviour. The AI stood in for expert judgement.

we applied GNNExplainer to the trained GIN model, which demonstrated the best performancewhere the paper describes this · verbatim
in the paper

What the record says

Technical · every part carries its own basis

+ in the paper~ our reading− not reported

How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.

~Role of AI
AI in a supporting roleour reading

The reported E_ox values in OxPot are produced by DFT HOMO energies combined with an empirically fitted linear calibration; learned models (AqSolPred, and the classical/graph models trained on OxPot) add property columns and demonstrate the data set's trainability rather than producing the headline DFT-experiment correlation

+What the AI was for
we also evaluated Message Passing Neural Networks, a class of Graph Neural Networks (GNNs) specifically designed to process graph-structured datawhere the paper describes this · verbatim
+How it was taught
Supervisedin the paper
+Models named
Empirical linear E_HOMO–E_ox regression · Trained from scratchAqSolPred · Off the shelfRandom Forest · Trained from scratchSupport Vector Machine (linear kernel) · Trained from scratchSupport Vector Machine (radial basis function kernel) · Trained from scratchMultilayer Perceptron · Trained from scratchGraph Convolutional Network (GCN) · Trained from scratchGraph Isomorphism Network (GIN) · Trained from scratchin the paper
+How results were checked
Held-outin the paper
the molecules tested from the literature were not used to determine the intercept and slope of the linear relationshipwhere the paper describes this · verbatim
+Code · weights · data
code not reportedweights not reporteddata availablein the paper
we introduce OxPot, a comprehensive, open-access data set comprising oxidation potential data for 15,238 organic moleculeswhere the paper describes this · verbatim
+Compute
not reportedin the paper

What this paper did not report

Technical · absence is published deliberately

Reported as not stated — 13 items
  • CodeWhether the code is available is not stated.
  • Trained model weightsWhether the trained model is available is not stated.
  • ComputeThe hardware or time used is not stated.
  • How many were testedThe paper gives no count of what was tested.
  • Version of Empirical linear E_HOMO–E_ox regressionWhich version of the model was used is not stated.
  • Version of AqSolPredWhich version of the model was used is not stated.
  • Version of Random ForestWhich version of the model was used is not stated.
  • Version of Support Vector Machine (linear kernel)Which version of the model was used is not stated.
  • Version of Support Vector Machine (radial basis function kernel)Which version of the model was used is not stated.
  • Version of Multilayer PerceptronWhich version of the model was used is not stated.
  • Version of Graph Convolutional Network (GCN)Which version of the model was used is not stated.
  • Version of Graph Isomorphism Network (GIN)Which version of the model was used is not stated.
  • What step 5 replacedThe paper gives no basis for what the AI stood in for.

About this article

Record aix-00108, version 2, checked by a person on 2026-10-08. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error