astronomy/ai produced the result/arXiv 2026 · v2
Neural networks sort eleven years of Mars plasma data into three regions
Researchers trained two neural networks to read ion measurements from NASA's MAVEN spacecraft and label which plasma region it was flying through. The networks were then run over all available observations from 2014 to 2025.
spectrum · one line per step, placed by what the step does · bright lines used AI
Automated Classification of Plasma Regions at Mars Using Machine Learning
arXiv, 2026
doi:10.48550/arxiv.2604.17131 · record aix-00073 v2 · checked 2026-10-08
- AI was for
- Classification
- Model family
- Convolutional neural network, Multilayer perceptron
- Checked by
- Held-out200 tested
- Code
- not reported
The finding the paper is about came from the AI.
What this research was about
Mars has no global magnetic field of the sort that shields Earth, so the stream of charged particles blowing off the Sun — the solar wind — runs almost directly into the planet's upper atmosphere. The meeting is not a simple collision. The flow slows abruptly at a shock front, piles up in a turbulent shell behind it, and closer to the planet gives way to a region governed by Mars itself. A spacecraft in orbit crosses between these zones repeatedly, and the particle instruments on board record a different pattern of particle energies in each one.
Knowing which zone a given measurement came from matters for almost any further question about how Mars loses its atmosphere. But the zones are defined by the shape of the data rather than by fixed positions in space, and the boundaries move as solar conditions change. Sorting the records has traditionally meant an expert looking at the measurements and deciding. The researchers set out to have software do that sorting instead, using only the ion energy readings from one instrument.
Where AI came in
Two networks were trained from scratch on examples that people had labelled by hand as solar wind, magnetosheath or magnetosphere, with ambiguous measurements near the boundaries marked unknown and left out. One was a multilayer perceptron, a plain stack of layers that saw a single snapshot of ion energies at one moment. The other was a convolutional network, a design that looks for patterns across a grid, and it was given a window of fifty consecutive snapshots so it could see how the energies changed over time.
On a test set of roughly 200 orbits the convolutional network reached a macro-averaged F1 score of about 0.95, a combined measure of how often it labelled correctly across all three classes. The simpler network reached about 0.87, calling 34.3% of magnetosheath measurements solar wind instead. The networks then labelled all available MAVEN observations from 2014 to 2025, with only predictions above a confidence of 0.8 kept, standing in for the hand-sorting an expert would otherwise have done.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record
The authors trained two neural networks to label the plasma region around Mars — solar wind, magnetosheath, or magnetosphere — from MAVEN SWIA ion omnidirectional energy spectra alone. A convolutional network reading 50-minute time–energy windows reached a macro-averaged F1 score of about 0.95 on a test set of roughly 200 manually labelled orbits, while a multilayer perceptron using single-time spectra reached about 0.87 and misclassified 34.3% of magnetosheath points as solar wind. Both models were then run over all available MAVEN observations from 2014 to 2025, and the resulting boundary locations were compared with published empirical bow shock and magnetic pileup boundary models.
How AI was used
Ion omnidirectional energy spectra from the MAVEN SWIA instrument were assembled for selected orbits and manually labelled as solar wind, magnetosheath, or magnetosphere, with ambiguous near-boundary samples marked unknown and excluded from both training and test sets. A labelled training pool of up to 40 orbits from January 2015 supplied the training data, and about 200 orbits spanning 2014 to 2025 formed the test set. Two supervised classifiers were trained from scratch with cross-entropy loss and the AdamW optimiser: a multilayer perceptron with two fully connected layers of 256 and 128 units taking a single 48-channel spectrum, and a convolutional network taking a 50×48 time–energy matrix of 50 consecutive one-minute spectra through convolutional stages of 16, 32 and 64 channels with batch normalisation, ReLU, dropout and residual skip connections, followed by global average pooling and a fully connected layer. Each configuration was repeated 10 times with different random seeds while the number of training orbits and the upper bound of the input energy range were varied to select the final setup, after which the trained models were run over the test orbits and over all available MAVEN observations, with class scores converted to probabilities by softmax and a 0.8 probability threshold applied to the full-mission maps.
The shape of the work
Structural · the record, drawn
no AI
Assemble MAVEN SWIA ion spectra
Obtaining raw data, whether by measurement, download or retrieval.
MAVEN ion energy spectrograms from 2014–2025 are used to build the dataset.where the paper describes this · verbatim
no AI
Manually label regions and exclude transitions
Cleaning, filtering, normalising or labelling data already obtained.
these cases are labeled as “unknown” and excluded from both the training and test datasetswhere the paper describes this · verbatim
no AI
Build single-spectrum and time-energy inputs
Encoding data into features, descriptors, embeddings or graphs.
The input to the CNN consists of 50 consecutive spectra with 48 energy channels, forming a 50×48 time–energy matrixwhere the paper describes this · verbatim
AI
Train MLP and CNN classifiers
Fitting model parameters, including fine-tuning an existing model. The AI stood in for manual curation.
All models are trained using the cross-entropy loss function, and the network parameters are optimized using the AdamW optimizer.where the paper describes this · verbatim
no AI
Sweep training-set size and input energy range
Iterative search over a space. Its result feeds back into an earlier step.
the final model configuration adopts 20 training orbits and uses the full SWIA energy channels as the neural network inputwhere the paper describes this · verbatim
AI
Classify held-out test orbits
Running a trained model over new data to predict, classify or score. The AI stood in for manual curation.
We first apply both models to the test dataset and evaluate their predictions and performancewhere the paper describes this · verbatim
no AI
Compare predictions with manual labels and empirical boundary models
Testing outputs against ground truth.
the CNN achieves a macro-averaged F1 score of approximately 95% on the test datasetwhere the paper describes this · verbatim
AI
Apply classifier to full MAVEN record
Running a trained model over new data to predict, classify or score. The AI stood in for manual curation.
we applied the trained model to all available MAVEN orbits from 2014 to 2025where the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
The paper's result is the classifier itself and the plasma-region labels it produces; the region identifications come from the trained networks
we employ two neural network architectureswhere the paper describes this · verbatim
approximately 200 MAVEN orbits from observations between 2014 and 2025 are selected as the test datasetwhere the paper describes this · verbatim
MAVEN SWIA data are publicly available through the MAVEN Science Data Centerwhere the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- CodeWhether the code is available is not stated.
- Trained model weightsWhether the trained model is available is not stated.
- ComputeThe hardware or time used is not stated.
- Version of CNN plasma-region classifierWhich version of the model was used is not stated.
- Version of MLP baseline classifierWhich version of the model was used is not stated.
About this article
Record aix-00073, version 2, checked by a person on 2026-10-08. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error