materials-chemistry/ai in a supporting role/Journal of Chemical Information and Modeling 2025 · v2
Oxidation potentials for 15,238 molecules assembled from quantum calculations and lab measurements
Researchers built OxPot, an open dataset of oxidation potentials for 15,238 organic molecules, using a fitted linear relationship to convert quantum-chemical calculations into potentials, then trained several machine learning models on the result.
spectrum · one line per step, placed by what the step does · bright lines used AI
Predicting Oxidation Potentials with DFT-Driven Machine Learning
Journal of Chemical Information and Modeling, 2025
doi:10.1021/acs.jcim.5c00159 · record aix-00108 v2 · checked 2026-10-08
- AI was for
- Property prediction
- Model family
- Random forest, Support vector machine, Multilayer perceptron, Graph neural network, Linear model
- Checked by
- Held-out
- Code
- not reported
AI processed or interpreted data, but the main finding does not rest on it.
What this research was about
An oxidation potential is the voltage at which a molecule gives up an electron. It tells chemists how easily a compound will be oxidised, which matters for batteries, dyes, drugs and anything that must survive contact with air. Measuring it means putting the substance in solution and sweeping a voltage across an electrode while watching the current — a technique called cyclic voltammetry. That works well, but one molecule at a time. There are millions of known organic compounds, and only a small fraction have ever had their oxidation potential measured, so researchers who want to screen many candidates have little data to draw on.
The researchers set out to fill that gap by computation rather than by experiment. They drew molecules from PubChem, a public chemical database, filtered them by a set of rules, and calculated a quantum-mechanical property of each one: the energy of its highest occupied molecular orbital, or HOMO — roughly, how tightly the outermost electrons are held. That calculation used density functional theory, a standard method for approximating the behaviour of electrons. They also measured oxidation potentials themselves in the laboratory for a smaller set of molecules, to anchor the calculations to reality.
Where AI came in
The link between the calculated HOMO energies and real oxidation potentials was a straight line fitted to the authors' own laboratory measurements — a simple statistical model whose slope and intercept came from the data. That fitted line stood in for the experiment itself, turning a calculation into a predicted potential for every curated molecule. Against the measurements it was fitted to, it gave a correlation of 0.977 and a root-mean-square error of 0.064; tested on molecules from the literature that were left out of the fit, the figures were 0.963 and a mean absolute error of 0.069. An off-the-shelf open-source model, AqSolPred, supplied predicted water solubility as another column of the dataset.
The team then trained machine learning models on the finished dataset to see whether the predicted potentials could be learned from molecular structure alone, without the quantum calculation. These included a random forest, support vector machines, a multilayer perceptron, and two graph neural networks, which treat a molecule as a network of atoms joined by bonds. The Graph Isomorphism Network scored best, with a mean absolute error of 0.074. A tool called GNNExplainer was then used to ask which atomic details that model leaned on; it pointed to oxidation state, the number of attached hydrogens and the number of directly bonded neighbours.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record
The authors built OxPot, an open-access data set of oxidation potentials for 15,238 organic molecules curated from PubChem. Oxidation potentials were assigned by computing HOMO energies with DFT (PBE0/cc-pVDZ with COSMO solvation) and applying a linear correlation fitted to the authors' own cyclic voltammetry measurements, which gave an R value of 0.977 and an RMSE of 0.064; comparison with literature-reported values for molecules outside the fit gave an R of 0.963 and an MAE of 0.069. Predicted values span −0.6 to 2.7 V, with most molecules between 0.4 and 1.4 V. Several machine learning models were then trained on the data set; among those tested, the Graph Isomorphism Network had the lowest MAE (0.074) and RMSE (0.116), and GNNExplainer applied to it identified oxidation state, number of bonded hydrogens and number of directly bonded neighbours as the most influential atom-level features.
How AI was used
Machine learning entered the work at three points. First, an empirically fitted linear regression between DFT-computed E_HOMO and cyclic-voltammetry E_ox was used to assign oxidation potentials to the PubChem-curated molecules, and an off-the-shelf open-source model, AqSolPred, supplied the solubility column alongside RDKit-derived descriptors such as topological polar surface area, hydrogen-bond donor and acceptor counts, aromatic rings and fractional Csp3. Second, classical models (Random Forest, support vector machines with linear and radial basis function kernels, and a multilayer perceptron) were trained on Morgan fingerprints, and message-passing graph neural networks — a Graph Convolutional Network and a Graph Isomorphism Network — were trained on molecular graphs with atom-level descriptors including atomic number, chirality, bonded-neighbour and bonded-hydrogen counts, formal charge, radical electrons, hybridisation, aromaticity, ring membership, TPSA, fractional Csp3, LogP, pH, oxidation state and molecular mass; models were scored with MAE and RMSE. Third, GNNExplainer was applied to the trained GIN model to rank the atom-level input features by influence on its predictions.
The shape of the work
Structural · the record, drawn
no AI
Curate and filter molecules from PubChem
Cleaning, filtering, normalising or labelling data already obtained.
A set of rules was applied to filter the molecules from PubChem, developing an enhanced data set with high prediction precision.where the paper describes this · verbatim
no AI
Measure oxidation potentials by cyclic voltammetry
Physical execution, by hand or by robot.
In CV, most of the molecules underwent single-electron transfer, and we used the maximum anodic peak potential (Epa) to report the potential.where the paper describes this · verbatim
no AI
Compute HOMO energies with DFT
Numerical or physics simulation, including where a learned surrogate replaces it.
the hybrid functional PBE0 paired with the cc-pVDZ basis set and the Restricted Kohn–Sham approach showed the best performancewhere the paper describes this · verbatim
AI
Fit E_HOMO–E_ox correlation and predict E_ox
Running a trained model over new data to predict, classify or score. The AI stood in for physical experiment.
by calculating the E HOMO values for each molecule using DFT and employing an empirically derived linear correlation with experimental E oxwhere the paper describes this · verbatim
AI
Derive molecular descriptors and predict solubility
Encoding data into features, descriptors, embeddings or graphs.
An open-source ML model, AqSolPred, trained on a refined version of AqSolDB, was utilized for solubility predictions.where the paper describes this · verbatim
AI
Train and evaluate ML models on OxPot
Fitting model parameters, including fine-tuning an existing model. The AI stood in for simulation.
We trained several classical ML algorithms, including Random Forest, Support Vector Machine (SVM) with both linear and radial function kernelswhere the paper describes this · verbatim
AI
Attribute feature importance in the trained GIN
Extracting understanding from model behaviour. The AI stood in for expert judgement.
we applied GNNExplainer to the trained GIN model, which demonstrated the best performancewhere the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
The reported E_ox values in OxPot are produced by DFT HOMO energies combined with an empirically fitted linear calibration; learned models (AqSolPred, and the classical/graph models trained on OxPot) add property columns and demonstrate the data set's trainability rather than producing the headline DFT-experiment correlation
we also evaluated Message Passing Neural Networks, a class of Graph Neural Networks (GNNs) specifically designed to process graph-structured datawhere the paper describes this · verbatim
the molecules tested from the literature were not used to determine the intercept and slope of the linear relationshipwhere the paper describes this · verbatim
we introduce OxPot, a comprehensive, open-access data set comprising oxidation potential data for 15,238 organic moleculeswhere the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- CodeWhether the code is available is not stated.
- Trained model weightsWhether the trained model is available is not stated.
- ComputeThe hardware or time used is not stated.
- How many were testedThe paper gives no count of what was tested.
- Version of Empirical linear E_HOMO–E_ox regressionWhich version of the model was used is not stated.
- Version of AqSolPredWhich version of the model was used is not stated.
- Version of Random ForestWhich version of the model was used is not stated.
- Version of Support Vector Machine (linear kernel)Which version of the model was used is not stated.
- Version of Support Vector Machine (radial basis function kernel)Which version of the model was used is not stated.
- Version of Multilayer PerceptronWhich version of the model was used is not stated.
- Version of Graph Convolutional Network (GCN)Which version of the model was used is not stated.
- Version of Graph Isomorphism Network (GIN)Which version of the model was used is not stated.
- What step 5 replacedThe paper gives no basis for what the AI stood in for.
About this article
Record aix-00108, version 2, checked by a person on 2026-10-08. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error