astronomy/ai produced the result/Astronomy and Astrophysics 2023 · v2
Neural networks stand in for a slow model of interstellar gas clouds
Astronomers trained neural networks to reproduce the output of the Meudon PDR code, a simulation of starlit interstellar gas that takes hours to run. The networks predicted the same spectral line brightnesses in a fraction of the time.
spectrum · one line per step, placed by what the step does · bright lines used AI
Neural network-based emulation of interstellar medium models
Astronomy and Astrophysics, 2023
doi:10.1051/0004-6361/202347074 · record aix-00204 v2 · checked 2026-10-09
- AI was for
- Simulation surrogate, Anomaly detection
- Model family
- Multilayer perceptron, Clustering
- Checked by
- Held-out3192 tested
- Code
- available
The finding the paper is about came from the AI.
What this research was about
The space between stars is not empty. It holds thin clouds of gas and dust, and where ultraviolet light from nearby stars strikes them, the light breaks molecules apart and heats the gas. Astronomers call such a region a photodissociation region. To work out the conditions inside one — its pressure, how strongly it is lit, how much dust lies along the line of sight — they compare the brightnesses of spectral lines, the narrow glows emitted by particular atoms and molecules, against what a detailed physical model predicts. The difficulty is cost. A single run of the Meudon PDR code, the model used here, is stated to last a few hours.
Fitting observations means running such a model again and again, for thousands of trial sets of conditions, which is impractical at those speeds. One common workaround is to compute a grid of models in advance and interpolate between the points, reading off values in the gaps. The researchers set out to replace that interpolation with trained networks, and to measure how the two compare on accuracy, speed and the memory they take up.
Where AI came in
A grid of precomputed model runs, spanning four input parameters, supplied the training data. Feedforward neural networks were then fitted to map those parameters straight to spectral line brightnesses, taking the place of the simulation itself at prediction time. Several design choices were themselves automated. One network, fitted with a loss function tolerant of extreme values, flagged suspect entries in the training grid for human review. A clustering algorithm sorted the lines into correlated groups, each handed its own network, in place of expert judgement. Principal component analysis, a statistical method for finding the few directions in which data mostly varies, set the width of a hidden layer rather than a search over options.
The networks were then run on a held-out test set of 3,192 points the models had not seen, alongside nearest-neighbour, linear, spline and radial basis function interpolation. They reached a mean error factor of 4.5% and a 99th-percentile error factor of 33.1%, against 10.2% and 97% for radial basis function interpolation, and were reported as 1,000 times faster than that method. The lightest network was described by 2.7 million parameters, or 43.2 MB, where storing the full training grid needs 1.65 GB.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record
The authors built neural network emulators of the Meudon PDR code, a photodissociation-region model that computes line intensities from a few physical parameters, using a grid of precomputed models as training data. Networks were trained after an outlier-removal step and with line clustering, PCA-based layer sizing, a polynomial input transform and a dense architecture, then compared on a held-out test set against nearest-neighbor, linear, spline and radial basis function interpolation. The networks reached a mean error factor of 4.5% and a 99th-percentile error factor of 33.1%, against 10.2% and 97% for radial basis function interpolation, and were reported as 1 000 times faster than that method. The lightest network was described by 2.7 million parameters, or 43.2 MB, compared with the 1.65 GB needed to store the full training grid.
How AI was used
A grid of Meudon PDR code evaluations over four input parameters (thermal pressure, UV field scaling factor, total visual extinction and viewing angle) was computed in advance and split into a training grid and a randomly drawn test set, with intensities and three of the parameters taken in decimal log scale and the parameters standardised. An artificial neural network was first fitted with a Cauchy loss function to flag training values whose errors were largest; those were manually reviewed and the confirmed outliers recorded in a binary mask used to exclude them from a masked mean-squared-error loss, while for the interpolation baselines the masked values were imputed by line-by-line radial basis function interpolation. Spectral clustering of the matrix of absolute Pearson correlations between lines produced homogeneous line subsets, each given its own network; principal component analysis on the training set set the size of the last hidden layer, and a degree-three polynomial transform of the standardised parameters was implemented as a fixed first layer. Feedforward networks with the exponential linear unit activation, and dense-architecture variants with skip concatenation of layer inputs and outputs, were trained in PyTorch with gradient-based optimisation, then run on the test points alongside SciPy interpolators so that error factors, evaluation times and parameter counts could be compared.
The shape of the work
Structural · the record, drawn
no AI
Compute grid of Meudon PDR models
Numerical or physics simulation, including where a learned surrogate replaces it.
We generated two datasets of Meudon PDR code evaluationswhere the paper describes this · verbatim
no AI
Preprocess parameters and intensities
Cleaning, filtering, normalising or labelling data already obtained.
The D parameters are thus standardized to have a zero mean and a unit standard deviation.where the paper describes this · verbatim
AI
Identify and mask outlier intensities
Cleaning, filtering, normalising or labelling data already obtained. The AI stood in for manual curation.
For this first fit, we resort to an ANN designed as described at the introduction of Sect. 4.where the paper describes this · verbatim
AI
Cluster lines into homogeneous subsets
Encoding data into features, descriptors, embeddings or graphs. The AI stood in for expert judgement.
We derive clusters of lines automatically from the correlation matrix using the spectral clustering algorithm.where the paper describes this · verbatim
AI
Size architecture with PCA and augment inputs
Encoding data into features, descriptors, embeddings or graphs. The AI stood in for exhaustive search.
We resort to a principal component analysis (PCA) on the training setwhere the paper describes this · verbatim
AI
Fit surrogate models to the training set
Fitting model parameters, including fine-tuning an existing model. The AI stood in for simulation.
The training set is used to fit all surrogate models.where the paper describes this · verbatim
AI
Predict line intensities at test points
Running a trained model over new data to predict, classify or score. The AI stood in for simulation.
Error factors are evaluated on the test set.where the paper describes this · verbatim
no AI
Compare accuracy, speed and memory
Testing outputs against ground truth.
The accuracies of surrogate models are evaluated on the test set, which contains points that they did not see during training.where the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
The reported result is the emulator itself: trained networks that stand in for the Meudon PDR code, compared against interpolation surrogates on speed, memory and accuracy.
These emulators are defined with artificial neural networks (ANNs) with adapted architectures and are fitted using regression strategies instead of interpolation methods.where the paper describes this · verbatim
It contains Ntest=3 192 points. These points were generated with 456 independent random draws from a uniform distributionwhere the paper describes this · verbatim
The code used to build the proposed ANNs can be found atwhere the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- Version of Outlier-detection ANN trained with Cauchy lossWhich version of the model was used is not stated.
- Version of ANN emulator R (hidden-layer size set by PCA)Which version of the model was used is not stated.
- Version of ANN emulator R+P (polynomial input transform)Which version of the model was used is not stated.
- Version of ANN emulator R+P+C (one network per line cluster)Which version of the model was used is not stated.
- Version of ANN emulator R+P+D (dense architecture)Which version of the model was used is not stated.
- Version of ANN emulator R+P+C+D (clustering plus dense architecture)Which version of the model was used is not stated.
- Version of Spectral clustering of line intensitiesWhich version of the model was used is not stated.
- Version of Principal component analysis of line log-intensitiesWhich version of the model was used is not stated.
- Version of Nearest-neighbor interpolation (SciPy)Which version of the model was used is not stated.
- Version of Piece-wise linear interpolation (SciPy)Which version of the model was used is not stated.
- Version of Spline interpolation (SciPy)Which version of the model was used is not stated.
- Version of Radial basis function interpolation (SciPy)Which version of the model was used is not stated.
About this article
Record aix-00204, version 2, checked by a person on 2026-10-09. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error