astronomy/ai produced the result/The Astrophysical Journal 2026 · v2
Machine-learnt stand-in for a planet model cuts interior inference from 42 hours to 8 minutes
Researchers worked out what lies inside exoplanets from their mass and radius. A trained statistical stand-in replaced the slow physics model inside the sampler, and the same calculation ran in minutes instead of hours.
spectrum · one line per step, placed by what the step does · bright lines used AI
Surrogate-accelerated Bayesian Inversion for Exoplanet Interior Characterization
The Astrophysical Journal, 2026
doi:10.3847/1538-4357/ae2ec4 · record aix-00059 v2 · checked 2026-10-08
- AI was for
- Simulation surrogate
- Model family
- Gaussian process, Linear model
- Checked by
- Held-out1000 tested
- Code
- not reported
The finding the paper is about came from the AI.
What this research was about
We cannot cut a planet open. For worlds orbiting other stars, often all we have is a mass and a radius, and from those two numbers astronomers try to work out the layers inside: an iron core, a rocky mantle of silicate minerals, and a gassy envelope of hydrogen, helium and water. The trouble is that many different recipes give the same mass and radius. A small, dense core with a thick atmosphere can weigh and measure the same as a larger core with a thin one. So the answer is not a single structure but a range of structures, each with a probability.
Mapping that range means running a physics model of the planet's interior many thousands of times, once for every candidate recipe the search tries out. Each run takes a second or so, and the search needs enough of them that the whole exercise can occupy a desk computer for days. The researchers set out to keep the same statistical machinery, which explores recipes at random and keeps those that match the observations, while making each step cheap enough to be practical.
Where AI came in
The researchers trained a stand-in for the physics model, a technique called polynomial chaos-Kriging that pairs a sparse polynomial fit with a Gaussian process, a method that interpolates between known points and reports its own uncertainty. They ran the real physics model 600 times for a given planet and used those paired inputs and outputs as training examples, then checked the stand-in against 400 more runs it had never seen. The error it made on those held-out cases was folded into the calculation, so the stand-in's own imprecision counted alongside the measurement uncertainty.
From then on, every step of the random search asked the stand-in rather than the physics model. For one tightly constrained sub-Neptune, the full physics version took about 42 hours on a single processor core while the stand-in version took 8 minutes, and the two answers agreed closely. The team also ran the pipeline on 1000 invented planets whose true interiors they knew, to see how often the stated ranges actually contained the right answer, and applied it to Earth and to the sub-Neptune TOI-270 d.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record
The study replaces a physics-based exoplanet interior model with a polynomial chaos-Kriging surrogate — a sparse polynomial trend combined with a Gaussian process — placed inside an MCMC sampler that infers core, mantle and atmosphere parameters from measured mass and radius. The surrogate was fitted to 600 forward-model evaluations per planet and checked on a held-out set of 400, with its residual variance added to the observational variance in the likelihood. For the tightly constrained sub-Neptune case, inference using the full forward model took about 42 hours on a single CPU core while the surrogate-based run took 8 minutes, and posterior means from the two agreed within 0.6σ. In a study of 1000 synthetic planets with known interiors, the 68% and 95% credible intervals contained the ground truth in 65-72% and 93-96% of runs, and the framework was also applied to Earth as a benchmark and to the sub-Neptune TOI-270 d.
How AI was used
Latin Hypercube samples were drawn from the parameter priors and passed through a physics-based interior model (iron core, silicate mantle, H2-He-H2O atmosphere) to build a simulation database; for shared priors, a 5,000-sample master database was localised to a target by selecting candidates within a four-standard-deviation hyperrectangle of the observed mass and radius and resampling them under a multivariate Gaussian centred on the observation. The localised set was split into 600 training and 400 test samples, with the training set's correlation structure also defining a Gaussian-copula prior. A sequential PC-Kriging surrogate was then fitted in UQLab: least-angle regression selects a sparse polynomial basis, and a sequence of universal Kriging models using increasing numbers of those polynomials as trend is compared by cross-validation. The held-out set was used to measure NRMSE and R2 and to estimate a single homoscedastic surrogate error term, which was added to the observational variance in a Gaussian likelihood. This surrogate likelihood was sampled with an Adaptive Metropolis MCMC, 10 chains of 10,000 steps with the first 2,500 discarded as burn-in, across five scenarios. The same pipeline was re-run on an ensemble of synthetic planets with known ground truth for a coverage analysis, and compared against one MCMC run that called the physics forward model directly.
The shape of the work
Structural · the record, drawn
no AI
Sample priors and run physics forward model
Numerical or physics simulation, including where a learned surrogate replaces it.
we create 1,000 simulations by drawing samples from the prior probability distributions of the input parameters using Latin Hypercube Sampling (LHS)where the paper describes this · verbatim
no AI
Localise and split training database
Cleaning, filtering, normalising or labelling data already obtained.
This localized dataset was then partitioned into a training dataset of Ntrain=600 samples and a held-out test dataset of Ntest=400 samples.where the paper describes this · verbatim
AI
Fit PCK surrogate
Fitting model parameters, including fine-tuning an existing model. The AI stood in for simulation.
Our implementation utilizes the sequential PC-Kriging (SPCK,) algorithm available in the UQLab software library.where the paper describes this · verbatim
AI
Evaluate surrogate on held-out test set
Testing outputs against ground truth.
The test set is used to evaluate the surrogate’s predictive performance and to estimate its empirical error term (σsurr).where the paper describes this · verbatim
AI
Surrogate-based MCMC posterior sampling
Running a trained model over new data to predict, classify or score. The AI stood in for simulation.
For each scenario, we initialize 10 chains at random points in the admissible parameter domain and run each chain for 10,000 steps.where the paper describes this · verbatim
no AI
Benchmark against full forward-model MCMC
Testing outputs against ground truth.
we performed an MCMC analysis using the computationally expensive forward model directly within the samplerwhere the paper describes this · verbatim
AI
Coverage study on synthetic planets
Testing outputs against ground truth. The AI stood in for simulation.
We ran our MCMC pipeline on the large ensemble of 1,000 synthetic planets for which the ground-truth interior parameters, θtrue, were known.where the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
The learned polynomial chaos-Kriging surrogate supplies every likelihood evaluation inside the MCMC loop, so the reported posteriors (including those for TOI-270 d) are produced by the surrogate rather than by the physics model, except in the single benchmark run.
we construct a surrogate model using polynomial chaos-Kriging (PCK,)where the paper describes this · verbatim
a large-scale coverage study with 1000 synthetic test cases to demonstrate the statistical reliability of our inferred credible intervalswhere the paper describes this · verbatim
Two identical MCMC inferences were run on a single CPU core (Apple M4 processor with 16 GB RAM)where the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- CodeWhether the code is available is not stated.
- Trained model weightsWhether the trained model is available is not stated.
- DataWhether the data are available is not stated.
- Version of Polynomial chaos-Kriging surrogate (sequential PC-Kriging, UQLab)Which version of the model was used is not stated.
- What step 4 replacedThe paper gives no basis for what the AI stood in for.
About this article
Record aix-00059, version 2, checked by a person on 2026-10-08. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error