materials-chemistry/ai produced the result/Communications Chemistry 2022 · v2
Deep-learning model designs petrol blends from octane and soot targets
Researchers trained a neural network to predict three combustion properties of fuels and their mixtures, then searched the model's internal representation for blends matching chosen targets. Eighty-six candidate mixtures came out; one blend of 22 components was put forward.
spectrum · one line per step, placed by what the step does · bright lines used AI
Artificial intelligence-driven design of fuel mixtures
Communications Chemistry, 2022
doi:10.1038/s42004-022-00722-3 · record aix-00058 v2 · checked 2026-10-08
- AI was for
- Property prediction, Candidate generation
- Model family
- Recurrent neural network, Multilayer perceptron, Graph neural network
- Checked by
- Held-out
- Code
- not reported
The finding the paper is about came from the AI.
What this research was about
Petrol is not one substance but a mixture of many hydrocarbons, and how it behaves in an engine depends on the whole blend. Two standard measures, the research octane number and the motor octane number, describe how well a fuel resists knocking, the premature ignition that damages engines. A third, the yield sooting index, describes how much soot a fuel tends to produce when it burns. All three are measured in the laboratory, which is slow and needs sizeable samples, so only a small fraction of possible blends has ever been tested.
Mixtures are harder still than single compounds. A simple assumption, that a blend's octane number is just the average of its components weighted by how much of each is present, does not always hold, because molecules interact as they burn. The researchers set out to build a computer model that predicts all three properties for pure compounds and for mixtures at once, and then to run that model backwards: instead of asking what a given blend would do, asking which blends would hit a chosen combination of octane numbers and sooting behaviour.
Where AI came in
A neural network was trained from scratch on a database of published laboratory measurements for pure hydrocarbons, laboratory blends and real fuels. Each molecule entered the model twice over, once as a text string spelling out its structure and once as a list of calculated numerical descriptors; the network compressed both into a single string of numbers, a kind of coded fingerprint. A mixture's fingerprint was formed by combining its components' fingerprints in proportion to how much of each was present, and a final layer of the network mapped fingerprints to the three properties. Tested on data held back from training, it tracked the measurements with a correlation above 0.92 for all three.
The design step depended on the model entirely. Because the space of fingerprints is continuous, the researchers could treat the trained network as a mathematical function and use gradient-based optimisation to hunt for blends whose predicted properties sat near targets of 95 research octane, 85 motor octane and 60 sooting index, while obeying European petrol specifications. That search stood in for trying candidates one by one, which the number of possible blends makes impractical. It produced 86 mixtures. Conventional thermodynamic calculations, not machine learning, then screened these for vapour pressure, leaving five. The candidates and their properties exist as model output; experimental confirmation was left to future work.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record
The authors curated a literature database of research octane number, motor octane number and yield sooting index measurements for pure hydrocarbons, surrogate blends and real fuels, and trained a deep-learning model that encodes each molecule from its SMILES string and Mordred descriptors and represents a mixture as a composition-weighted linear combination of its components' latent vectors. On an independent test set the model reached R above 0.92 for all three properties, and its mixture octane-number mean absolute errors were lower than a linear-by-mole mixing rule across the blend sizes reported. Searching the model's latent space under gasoline specification constraints for targets of RON 95, MON 85 and YSI 60 produced 86 candidate mixtures, of which five had estimated Reid vapour pressure in the stated acceptable range; one blend of 22 components was put forward as the most promising candidate, with experimental confirmation left to future work.
How AI was used
A deep-learning model was trained from scratch to predict three combustion properties jointly. SMILES strings were one-hot encoded and passed through three stacked LSTM layers, while Mordred molecular descriptors were min-max normalised and passed through fully connected layers; the two resulting fingerprints were concatenated into a per-component latent vector. A mixing operator inside the training loop formed a mixture's latent vector as a matrix-vector product of component latent vectors with their compositions, and a fully connected predictor mapped latent vectors to the target properties, with the sooting index predicted via the measured soot-volume-fraction quantities and converted using the scale endpoints. The data were split by hierarchical stratified sampling, training used a weighted mean-squared-error loss, and hyperparameters were tuned with a Bayesian optimisation platform on a validation split. Because the latent space is continuous and differentiable, the trained predictor was used as the objective of a constrained optimisation solved with SciPy using PyTorch automatic differentiation gradients, started from database points closest to the target vector; a greedy depth-first search then reduced those solutions to smaller blends, with constraints enforced by Dykstra's method. A published graph-neural-network model was also run on test-set components for comparison. Non-learned thermodynamic correlations were applied afterwards to screen the resulting candidates.
The shape of the work
Structural · the record, drawn
no AI
Curate RON, MON and YSI database
Obtaining raw data, whether by measurement, download or retrieval.
The database of experimentally obtained measurements for the three combustion-related properties (RON, MON, and YSI) for single hydrocarbons and mixtures was curatedwhere the paper describes this · verbatim
no AI
Stratified train/validation/test split
Cleaning, filtering, normalising or labelling data already obtained.
each subset was randomly split into 85% train/validation and 15% test set using stratified sampling in the scikit-learn librarywhere the paper describes this · verbatim
no AI
Encode molecules as SMILES and Mordred descriptors
Encoding data into features, descriptors, embeddings or graphs.
Generated SMILES strings were converted to a binary matrix using one-hot encoding.where the paper describes this · verbatim
AI
Train joint-property model with linear mixing operator
Fitting model parameters, including fine-tuning an existing model. The AI stood in for conventional algorithm.
the weighted loss function is used to train the modelwhere the paper describes this · verbatim
AI
Predict properties for held-out species and blends
Running a trained model over new data to predict, classify or score. The AI stood in for physical experiment.
Figure 2 shows the parity plots for the model’s independent test setwhere the paper describes this · verbatim
no AI
Compare predictions with measurements and baselines
Testing outputs against ground truth.
We compared the predictive model’s performance with (1) three data-driven models developed for predicting RON, MON, and YSI of pure componentswhere the paper describes this · verbatim
AI
Search latent space for mixtures matching targets
Iterative search over a space. The AI stood in for exhaustive search.
From the results, 20 mixtures with 5-26 components were reported using a full-scope search, whereas the greedy search generated 66 mixtureswhere the paper describes this · verbatim
no AI
Post-screen candidates on physical properties
Reducing a candidate set by filtering or ranking, in a single pass.
Five of 86 mixtures exhibited RVP in an acceptable range (50 kPa ≤ RVP ≤ 100 kPa).where the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
The designed fuel mixtures that the paper reports are produced by searching the deep-learning model's latent space; the candidate blends and their predicted properties exist only as model output
the AI fuel design tool was built on top of an end-to-end DL model based on recurrent and fully connected (FC) layerswhere the paper describes this · verbatim
MAEs were calculated for the ON predictions of 69 mixtures of varying sizes in the independent test set.where the paper describes this · verbatim
Training and test datasets for pure components and mixtures are provided in Supplementary Data 2 and Supplementary Data 3.where the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- CodeWhether the code is available is not stated.
- Trained model weightsWhether the trained model is available is not stated.
- ComputeThe hardware or time used is not stated.
- How many were testedThe paper gives no count of what was tested.
- Version of Joint-properties predictive DL model (Extractor 1 LSTM encoder, Extractor 2 fully connected encoder, mixing operator, predictor)Which version of the model was used is not stated.
- Version of GNN model by Schweidtmann et al. (RON/MON/DCN baseline)Which version of the model was used is not stated.
About this article
Record aix-00058, version 2, checked by a person on 2026-10-08. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error