~/aixsci
200 records · all checked

astronomy/ai produced the result/arXiv 2024 · v2

Neural networks stand in for slow cosmology code in parameter fitting

Researchers built a framework for training and sharing neural-network stand-ins for a widely used cosmology code, then used them to fit simulated sky data. The networks produced the predicted sky patterns in place of the original calculation.

1. Generate training spectra with CAMB2. Normalise and compress spectra3. Train neural network emulators4. Assess emulator accuracy against CAMB5. Build simulated data vectors6. Run MCMC inference with the emulators7. Compare recovered cosmology with native CAMB chains

spectrum · one line per step, placed by what the step does · bright lines used AI

A complete framework for cosmological emulation and inference with CosmoPower
arXiv, 2024

doi:10.48550/arxiv.2405.07903 · record aix-00060 v2 · checked 2026-10-08

ai-resultrole of AI
AI was for
Simulation surrogate
Model family
Multilayer perceptron
Checked by
Held-out
Code
available

The finding the paper is about came from the AI.

read as

The science is explained before the AI appears. Switch to field specialist to go straight to the method.

Assumes the discipline and goes straight to the method.

The diagram, the record and what the paper did not report are identical in both modes. Only the framing changes — never the evidence.

Introduction by AIxSci · plain language

What this research was about

Cosmologists test ideas about the universe by predicting what the sky should look like under a given set of assumptions, then comparing those predictions with observations. The predictions come from so-called Einstein-Boltzmann codes, which track how matter and light behaved as the young universe expanded and cooled. Their output is a set of power spectra: summaries of how much structure there is on each scale, whether in the faint afterglow of the Big Bang known as the cosmic microwave background, or in the way galaxies are spread across the sky. The difficulty is speed. Fitting a model means running the calculation hundreds of thousands of times, once per trial set of parameters, and each run is slow.

The authors set out to build a complete route around that bottleneck, within a package called CosmoPower. That meant not just training stand-in models, but also a settled way to specify them, test them, package them and hand them to others, plus adaptors so that two established analysis codes, Cobaya and CosmoSIS, could use the packaged stand-ins in place of the original calculation.

Where AI came in

The AI here is a set of emulators: dense neural networks, each with four hidden layers of 512 units, trained to map a set of cosmological parameters directly onto the resulting spectra. Training data came from running the CAMB code on parameter sets drawn across the ranges of interest, giving pairs of input and answer for the networks to learn from. Some outputs were first compressed with principal component analysis, a standard way of describing many numbers by a few summary ones. Separate networks were trained for the standard model of cosmology and for four variants of it, and for the clumpier, non-linear spread of matter.

Once trained, the networks stood in for the Einstein-Boltzmann calculation itself. The authors checked them against held-back CAMB results, then loaded them into Cobaya and CosmoSIS to fit simulated microwave-background and galaxy-survey data, and compared the parameter ranges recovered with those from runs using CAMB directly. On the simulated microwave-background data the emulator runs matched the CAMB results and the input cosmology to within 0.1 sigma, that is a tenth of the uncertainty on each parameter. One spectrum evaluation took about 0.1 seconds, against about 20 seconds for CAMB.

Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.

The work

Technical · from the record

The authors built a framework for specifying, training, packaging and distributing machine-learning emulators of Einstein-Boltzmann codes within the CosmoPower package, together with wrappers that load the packaged emulators into the Cobaya and CosmoSIS inference codes. Using it they trained dense neural network emulators of CAMB v1.5.0 outputs - CMB TT, TE, EE, BB and lensing-potential angular power spectra, linear and non-linear matter power spectra, background evolution quantities and derived parameters - for LCDM and four extension models, using 100000 training samples for LCDM and 120000 for the extended models. Compared with direct CAMB calculations the emulators stayed within 10% of a cosmic-variance-limited experimental uncertainty, with the exception of the w0wa emulator, where some small-scale CMB outliers reached about 80% of that uncertainty. In MCMC runs on simulated cosmic-variance-limited CMB data the emulator chains reproduced the CAMB posteriors and the input cosmology to within 0.1 sigma, with a single power spectrum evaluation taking about 0.1 s against about 20 s for CAMB.

How AI was used

Cosmological parameter sets were drawn by Latin hypercube sampling over the ranges in Table 2 and passed to CAMB v1.5.0 with Stage-IV accuracy settings to produce training spectra, background quantities and derived parameters, stored in HDF5 files; a small fraction of unphysical samples was discarded. Inputs and outputs were normalised by their means and standard deviations, most quantities were emulated in the logarithm, and the TE and lensing-potential spectra were first compressed with a principal component analysis (512 and 64 retained components respectively), with scree plots used to choose the number of components. Dense neural networks with four hidden layers of 512 neurons were then fitted to map parameters to spectra, trained with the Adam optimiser over successive learning iterations whose learning rates, batch sizes, validation split, gradient accumulation steps, patience values and maximum epochs are set in the packaging prescription. Separate emulators were trained for LCDM and for the +Neff, +sum m_nu, +Neff+sum m_nu and +w0wa extensions, and for the non-linear matter power spectrum and non-linear boost with HMCode baryonic feedback parameters as additional inputs. Accuracy was evaluated by passing held-back samples through each emulator and plotting errors against the CAMB outputs, both as fractional differences and relative to cosmic-variance-limited or Simons Observatory noise curves. The packaged emulators were then loaded through the Cobaya and CosmoSIS wrappers to run Monte Carlo posterior sampling on a simulated cosmic-variance-limited CMB data vector and on a Stage-IV-like 3x2pt large-scale-structure data set, with an optional fall-through to the native Einstein-Boltzmann code for parameters outside the trained range.

The shape of the work

Structural · the record, drawn

SIMULATIONREPRESENTATIONTRAININGVALIDATIONSIMULATIONINFERENCEVALIDATION1234567AIAIAIGenerate trainingspectra with CAMBNormalise andcompress spectraTrain neuralnetwork emulatorsAssess emulatoraccuracy againstCAMBBuild simulateddata vectorsRun MCMCinference withthe emulatorsCompare recoveredcosmology withnative CAMB chai…↤ simulation↤ simulation↤ simulation
AI stepNo AI↤ what the AI stood in for
1Simulation
no AI

Generate training spectra with CAMB

Numerical or physics simulation, including where a learned surrogate replaces it.

we generate NS=105 sets of output spectra as training datawhere the paper describes this · verbatim
in the paper
2Representation
no AI

Normalise and compress spectra

Encoding data into features, descriptors, embeddings or graphs.

we follow in first decomposing the spectra with a Principal Component Analysis (PCA) and then subsequently emulating the sets of PCswhere the paper describes this · verbatim
in the paper
3Training
AI

Train neural network emulators

Fitting model parameters, including fine-tuning an existing model. The AI stood in for simulation.

CosmoPower uses the Adam optimiser to determine how to tweak the hyperparameterswhere the paper describes this · verbatim
in the paper
4Validation
AI

Assess emulator accuracy against CAMB

Testing outputs against ground truth. The AI stood in for simulation.

we perform a number of comparisons between the observables emulated and those calculated directly with CAMBwhere the paper describes this · verbatim
in the paper
5Simulation
no AI

Build simulated data vectors

Numerical or physics simulation, including where a learned surrogate replaces it.

we generate a smooth data vector with cosmic-variance-limited noisewhere the paper describes this · verbatim
in the paper
6Inference
AI

Run MCMC inference with the emulators

Running a trained model over new data to predict, classify or score. The AI stood in for simulation.

we can use our emulators in parameter inference analysis, generating posterior samples using Monte Carlo chainswhere the paper describes this · verbatim
in the paper
7Validation
no AI

Compare recovered cosmology with native CAMB chains

Testing outputs against ground truth.

we can reproduce the CAMB best-fit cosmology and posterior distributionwhere the paper describes this · verbatim
in the paper

What the record says

Technical · every part carries its own basis

+ in the paper~ our reading− not reported

How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.

~Role of AI
AI produced the resultour reading

The emulators are the object of the paper: the released neural networks supply the cosmological observables and the posterior samples that the paper's accuracy and timing claims are about

+What the AI was for
we implement the emulators as dense neural networks, with four hidden layers of 512 neurons eachwhere the paper describes this · verbatim
+Model families
+How it was taught
Supervisedin the paper
+Models named
CosmoPower emulators of CAMB (dense NN and PCAplusNN variants) · Trained from scratchin the paper
+How results were checked
Held-outin the paper
of which 20% will be used for validating the network accuracy, and the rest for trainingwhere the paper describes this · verbatim
+Code · weights · data
code availableweights availabledata not reportedin the paper
We release a full software suite for python that allows easy creation, testing, and usage of CosmoPower emulatorswhere the paper describes this · verbatim
+Compute
Training a Cl network on 10^5 samples takes O(1 h) on a GPU, or O(10 h) on a CPU; an LCDM cosmic-variance-limited CMB chain took ~20 minutes with CosmoPower versus ~10 hours with CAMB; Supercomputing Wales resources acknowledgedin the paper

What this paper did not report

Technical · absence is published deliberately

Reported as not stated — 3 items
  • DataWhether the data are available is not stated.
  • How many were testedThe paper gives no count of what was tested.
  • Version of CosmoPower emulators of CAMB (dense NN and PCAplusNN variants)Which version of the model was used is not stated.

About this article

Record aix-00060, version 2, checked by a person on 2026-10-08. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error