~/aixsci
200 records · all checked

materials-chemistry/ai produced the result/Materials 2025 · v2

Machine learning predicts marine steel corrosion from seawater conditions and alloy make-up

Researchers trained models to predict how fast six marine engineering steels corrode in seawater. A genetic algorithm tuned the models, and a mixture model plus a generative network produced synthetic extra training samples.

1. Collect corrosion dataset for six marine steels2. Create element-property features3. Reduce features and normalise4. GA hyperparameter search and base-model selection5. Generate virtual sample inputs with GMM6. Generate virtual sample outputs with RegGAN7. Retrain base model on augmented training set8. Evaluate on test set and compare VSG methods

spectrum · one line per step, placed by what the step does · bright lines used AI

An Integrated Approach Using GA-XGBoost and GMM-RegGAN for Marine Corrosion Prediction Under Small Sample Size
Materials, 2025

doi:10.3390/ma18163760 · record aix-00068 v2 · checked 2026-10-08

ai-resultrole of AI
AI was for
Property prediction, Candidate generation
Model family
Gradient-boosted trees, Random forest, Support vector machine, Multilayer perceptron, Generative adversarial network, Clustering
Checked by
Held-out
Code
not reported

The finding the paper is about came from the AI.

read as

The science is explained before the AI appears. Switch to field specialist to go straight to the method.

Assumes the discipline and goes straight to the method.

The diagram, the record and what the paper did not report are identical in both modes. Only the framing changes — never the evidence.

Introduction by AIxSci · plain language

What this research was about

Steel used in harbours, ships and offshore structures sits in seawater, where it slowly rusts away. How fast this happens depends on the water around it — its temperature, how salty it is, how much dissolved oxygen it holds, its acidity, and a measure called ORP that describes how chemically oxidising the water is. It also depends on the steel itself, since small additions of other elements change how the metal behaves. Measuring corrosion rates means running electrochemical tests, which take time and equipment, so published measurements are relatively few. That scarcity is the difficulty: a predictive model fitted to a small table of numbers has little to learn from.

The researchers gathered corrosion records for six commonly used marine engineering steels from the published literature, together with the seawater conditions under which each was measured. They converted each steel's elemental recipe into descriptors of physical, thermal, atomic, electronegativity and orbital properties, trimmed these down, and set out to build a model that predicts corrosion rate from the remaining features — while also finding a way to work around the small size of the dataset.

Where AI came in

Machine learning carried the whole prediction task. Five learned regression methods — support vector regression, a random forest, two gradient-boosted tree methods and a neural network — were fitted to the data, with a genetic algorithm searching for each method's settings instead of an exhaustive sweep. A genetic algorithm imitates breeding and mutation to hunt for good combinations. XGBoost, a gradient-boosted tree method, came out with the lowest cross-validated error, 2.785, and the paper reports tuning cutting that error by 12.58%.

A second group of models stood in for laboratory measurements. A Gaussian mixture model, which describes data as a blend of simple statistical clumps, invented new sets of plausible input conditions, kept within the range of the real data. A regression GAN — two networks trained against each other, one producing values and one judging them — then supplied a corrosion rate for each invented input. Adding these synthetic samples to the real ones and retraining lowered test errors by 14.94%, 15.55% and 14.04% on three measures, with the best results at 300 added samples.

Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.

The work

Technical · from the record

The study builds a machine learning model that predicts the corrosion rate of marine steel from five seawater environment variables and six composition-derived property features. A genetic algorithm tuned hyperparameters for five candidate algorithms, and XGBoost gave the lowest cross-validated RMSE of 2.785 with a standard deviation of 0.054; the paper reports the cross-validation RMSE falling by 12.58% with GA tuning. A Gaussian mixture model then sampled synthetic input vectors and a regression GAN supplied their corrosion-rate outputs, and retraining on the augmented set reduced test errors by 14.94% in RMSE, 15.55% in MAE and 14.04% in MAPE relative to training on the original samples only, with the best result at 300 virtual samples.

How AI was used

Element compositions were converted into property descriptors by fixed formulas, then reduced using Pearson correlation grouping, variance selection and GBDT feature importance, and Min-Max normalised. On an 80/20 split, five learned regressors (SVR, random forest, LightGBM, XGBoost, ANN) were fitted with hyperparameters searched by a genetic algorithm under 5-fold cross-validation with RMSE as the objective, stopping when the best error changed by less than 0.01 between generations, and the best tuned model was carried forward. For data augmentation, a Gaussian mixture model was fitted to the training inputs by expectation-maximisation with the component count chosen from AIC and BIC over one to eleven components; virtual inputs were drawn from it, constrained to training-set feature bounds and checked against the original distributions with a Kolmogorov-Smirnov test. Those inputs plus noise were passed to the generator of a RegGAN whose generator and discriminator were three-layer networks trained on the training set, with the generator updated twice per discriminator update, and each virtual output taken as the mean of 20 noise draws. Virtual samples were merged with the real training set to retrain the base model, with virtual-sample counts swept from 0 to 500 and each generation method repeated 50 times, and compared against MD-MTD, t-SNE, GMM, NITAE and CGAN augmentation on test-set RMSE, MAE and MAPE.

The shape of the work

Structural · the record, drawn

ACQUISITIONREPRESENTATIONPREPARATIONTRAININGGENERATIONGENERATIONTRAININGVALIDATION12345678AIAIAIAIAIAICollect corrosiondataset for sixmarine steelsCreateelement-propertyfeaturesReduce featuresand normaliseGA hyperparametersearch andbase-model selec…Generate virtualsample inputswith GMMGenerate virtualsample outputswith RegGANRetrain basemodel onaugmented traini…Evaluate on testset and compareVSG methods↤ exhaustive search↤ physical experiment↤ physical experimentloops back
AI stepNo AI↤ what the AI stood in for
1Acquisition
no AI

Collect corrosion dataset for six marine steels

Obtaining raw data, whether by measurement, download or retrieval.

marine corrosion data of six commonly used marine engineering structural steels were collected from the literaturewhere the paper describes this · verbatim
in the paper
2Representation
no AI

Create element-property features

Encoding data into features, descriptors, embeddings or graphs.

we transformed the original metal element information into 17 types of physical, heat, atomic, electronegativity, and orbital propertieswhere the paper describes this · verbatim
in the paper
3Preparation
AI

Reduce features and normalise

Cleaning, filtering, normalising or labelling data already obtained.

Feature reduction based on GBDT feature importance analysiswhere the paper describes this · verbatim
in the paper
4Training
AI

GA hyperparameter search and base-model selection

Fitting model parameters, including fine-tuning an existing model. The AI stood in for exhaustive search.

the genetic algorithm (GA), a global search evolutionary algorithm, was selected for hyperparameter tuningwhere the paper describes this · verbatim
in the paper
5Generation
AI

Generate virtual sample inputs with GMM

Producing candidate objects that did not previously exist. The AI stood in for physical experiment.

The generation of virtual sample inputs is primarily achieved through sampling from a Gaussian Mixture Model (GMM).where the paper describes this · verbatim
in the paper
6Generation
AI

Generate virtual sample outputs with RegGAN

Producing candidate objects that did not previously exist. The AI stood in for physical experiment.

The output of the virtual samples is primarily obtained by feeding the virtual sample inputs into the RegGAN surrogate modelwhere the paper describes this · verbatim
in the paper
7Training
AI

Retrain base model on augmented training set

Fitting model parameters, including fine-tuning an existing model.

the training set samples and the generated virtual samples are merged to form a new training setwhere the paper describes this · verbatim
in the paper
8Validation
AI

Evaluate on test set and compare VSG methods

Testing outputs against ground truth. Its result feeds back into an earlier step.

The performance of the proposed model is ultimately evaluated on the testing set.where the paper describes this · verbatim
in the paper

What the record says

Technical · every part carries its own basis

+ in the paper~ our reading− not reported

How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.

~Role of AI
AI produced the resultour reading

The paper's result is the predictive model itself; all reported outcomes are model errors on the corrosion dataset

+What the AI was for
a genetic algorithm (GA)-optimized machine learning framework is employed to derive the optimal GA-XGBoost modelwhere the paper describes this · verbatim
+How it was taught
SupervisedUnsupervisedin the paper
+Models named
XGBoost (GA-optimised) · Trained from scratchLightGBM · Trained from scratchRandom forest · Trained from scratchSupport vector regression · Trained from scratchArtificial neural network · Trained from scratchGBDT (feature importance) · Trained from scratchGaussian mixture model · Trained from scratchRegGAN · Trained from scratchCGAN (comparison VSG method) · Trained from scratchNITAE (comparison VSG method) · Trained from scratcht-SNE (comparison VSG method) · Trained from scratchin the paper
+How results were checked
Held-outin the paper
The marine steel corrosion dataset was divided into a training set and a test set, with 80% allocated for trainingwhere the paper describes this · verbatim
+Code · weights · data
code not reportedweights not reporteddata availablein the paper
The detailed records of environmental factors and corrosion rates are listed in Supplementary Table S1.where the paper describes this · verbatim
+Compute
not reportedin the paper

What this paper did not report

Technical · absence is published deliberately

Reported as not stated — 18 items
  • CodeWhether the code is available is not stated.
  • Trained model weightsWhether the trained model is available is not stated.
  • ComputeThe hardware or time used is not stated.
  • How many were testedThe paper gives no count of what was tested.
  • Version of XGBoost (GA-optimised)Which version of the model was used is not stated.
  • Version of LightGBMWhich version of the model was used is not stated.
  • Version of Random forestWhich version of the model was used is not stated.
  • Version of Support vector regressionWhich version of the model was used is not stated.
  • Version of Artificial neural networkWhich version of the model was used is not stated.
  • Version of GBDT (feature importance)Which version of the model was used is not stated.
  • Version of Gaussian mixture modelWhich version of the model was used is not stated.
  • Version of RegGANWhich version of the model was used is not stated.
  • Version of CGAN (comparison VSG method)Which version of the model was used is not stated.
  • Version of NITAE (comparison VSG method)Which version of the model was used is not stated.
  • Version of t-SNE (comparison VSG method)Which version of the model was used is not stated.
  • What step 3 replacedThe paper gives no basis for what the AI stood in for.
  • What step 7 replacedThe paper gives no basis for what the AI stood in for.
  • What step 8 replacedThe paper gives no basis for what the AI stood in for.

About this article

Record aix-00068, version 2, checked by a person on 2026-10-08. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error