astronomy/ai produced the result/arXiv 2025 · v2
Random forest sorts 130 million Magellanic Cloud sources into ten classes
Astronomers trained a probabilistic random forest on spectroscopically classified stars and galaxies, then used it to assign a class and a probability to each of the roughly 130 million sources in a near-infrared survey of the Magellanic Clouds.
spectrum · one line per step, placed by what the step does · bright lines used AI
The VMC Survey : LI. Classifying extragalactic sources using a probabilistic random forest supervised machine learning algorithm
arXiv, 2025
doi:10.48550/arxiv.2501.08196 · record aix-00235 v2 · checked 2026-10-09
- AI was for
- Classification
- Model family
- Random forest
- Checked by
- Held-out
- Code
- not reported
The finding the paper is about came from the AI.
What this research was about
The Magellanic Clouds are two small galaxies near our own, and surveys of them collect light from enormous numbers of points on the sky. Each point, or source, might be a star inside the Clouds, a star in the foreground of our own Galaxy, or something far beyond: a distant galaxy, or an active galactic nucleus, meaning a galaxy whose central black hole is drawing in gas and shining brightly. Telling these apart reliably usually needs a spectrum, which spreads a source's light out into its component colours. Spectra take a lot of telescope time, so only a small fraction of sources ever get one.
What is available for almost every source is photometry: brightness measurements in a handful of filters, from optical light through to the far infrared. Differences between those brightnesses, known as colours, carry clues about what a source is. The researchers set out to turn the small set of sources with known spectroscopic identities into a way of labelling every source in the VISTA Survey of the Magellanic Clouds, and to say how confident each label was.
Where AI came in
The AI here is a probabilistic random forest, a method that grows many decision trees, each asking a series of questions about a source's brightnesses and colours, and combines their votes into a class plus a probability for each class. The probabilistic version also takes the measurement errors into account. The team built a table of 237 features per source by cross-matching the near-infrared catalogue with optical, infrared and far-infrared data, labelled a training set using new and published spectra, and trained separate classifiers for each Cloud with 100 trees.
The trained classifiers then ran over both full catalogues, standing in for the spectroscopy and hand-sorting that would otherwise be needed to identify each source. The paper's catalogues are the model's output: the classifications and their probabilities. Those probabilities were used to divide the results into low-, mid- and high-confidence tiers. On held-out test data the classifiers reached average accuracies of about 79 per cent for the Small Magellanic Cloud and 87 per cent for the Large, rising to about 90 and 98 per cent for the 56,696,719 sources with class probabilities above 80 per cent.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record
A probabilistic random forest was trained on optical-to-far-infrared photometry of spectroscopically classified sources and applied to the roughly 130 million sources of the VISTA Survey of the Magellanic Clouds, assigning each source one of ten classes or an 'Unknown' label together with class probabilities. The classifiers reached average accuracies of about 79 per cent (SMC) and 87 per cent (LMC) on held-out data, and about 90 per cent (SMC) and 98 per cent (LMC) for the 56,696,719 sources with class probabilities above 80 per cent. After removing sources classed as Unknown, 707,939 sources in the SMC field and 397,899 in the LMC field were classified, including more than 77,600 extragalactic sources behind the Clouds. Cross-matching with independent X-ray and radio catalogues found 554 of 883 X-ray sources classed as AGN, and 1756 of 2694 radio sources classed as AGN with 659 classed as galaxies.
How AI was used
The authors cross-matched the VMC near-IR PSF catalogue with SMASH optical, Gaia DR3, Spitzer SAGE, AllWISE and unWISE photometry within a 1 arcsec radius, added a smoothed Herschel SPIRE 250 μm background flux measurement at each position, and formed colours between all photometric bands with propagated errors, giving 237 features with measurement errors and homogenised NaN values for missing data. Labels came from spectroscopically classified sources, including 26 sources from new SAAO 1.9m observations and 22 from new SALT observations, literature catalogues such as Milliquas, 6dFGS, SDSS in the GAMA09 field and SAGE-spec, plus a Simbad search for proper-motion stars; ten classes were used, together with an Unknown class built from randomly selected VMC sources. Training sets were split 75/25 into training and test sets with the split randomised across runs, and the training portion was upsampled by random duplication to the size of the largest class. Separate classifiers were trained for the SMC and LMC, sharing extragalactic and foreground training sources while keeping Magellanic stellar classes Cloud-specific, using the probabilistic random forest implementation with a probability threshold of 0.05 and 100 trees chosen from a scan over tree numbers; runs were repeated across ten random seed states and confusion matrices averaged. The trained classifiers were then run over the full SMC and LMC PSF catalogues, mean-decrease-impurity feature importances were averaged over ten repeats, and outputs were divided into low-, mid- and high-confidence catalogues by class probability, with sources whose combined AGN and galaxy probabilities crossed a threshold promoted between tiers. X-ray, radio, Quaia and published YSO catalogues were held out of the feature set and used only for cross-matched comparison.
The shape of the work
Structural · the record, drawn
no AI
Obtain new optical spectra of candidate sources
Obtaining raw data, whether by measurement, download or retrieval.
We observed 174 new optical spectrawhere the paper describes this · verbatim
no AI
Build multi-wavelength feature table
Encoding data into features, descriptors, embeddings or graphs.
New features were created by subtracting each feature from all the other features to create colourswhere the paper describes this · verbatim
no AI
Assemble, label and balance training sets
Cleaning, filtering, normalising or labelling data already obtained.
the ‘resample’ function of Python’s Scikit-learn module was used to upsample all the class samples to the same size as the majority classwhere the paper describes this · verbatim
AI
Train and test separate SMC and LMC classifiers
Fitting model parameters, including fine-tuning an existing model. The AI stood in for manual curation.
Individual classifiers are trained for the SMC and LMC due the different stellar populationswhere the paper describes this · verbatim
AI
Classify all VMC sources
Running a trained model over new data to predict, classify or score. The AI stood in for manual curation.
The classifiers were used on the entirety of the LMC and SMC PSF catalogueswhere the paper describes this · verbatim
AI
Compute and inspect feature importances
Extracting understanding from model behaviour.
The classifiers are trained on the full datasets for SMC and LMC, from which the feature importances were calculated.where the paper describes this · verbatim
no AI
Split output into confidence tiers
Reducing a candidate set by filtering or ranking, in a single pass.
The catalogues of sources are separated into high-confidence sources (Pclass > 80%), mid-confidence sources (60% < Pclass < 80%)where the paper describes this · verbatim
no AI
Check classifications against independent catalogues
Testing outputs against ground truth.
we use the radio and X-ray detected sources as an independent check to test the classificationswhere the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
The paper's product is the set of source classifications and class probabilities produced by the trained classifiers; the reported catalogues exist only as model output.
We used a supervised machine learning algorithm (probabilistic random forest) to classify ∼ 130 million sourceswhere the paper describes this · verbatim
each dataset was split into training and testing sets, where 75 % of the data were trained onwhere the paper describes this · verbatim
The training sets for the SMC and LMC are made available alongside this paper as online supplementary material.where the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- CodeWhether the code is available is not stated.
- Trained model weightsWhether the trained model is available is not stated.
- ComputeThe hardware or time used is not stated.
- How many were testedThe paper gives no count of what was tested.
- Version of Probabilistic Random Forest (PRF)Which version of the model was used is not stated.
- What step 6 replacedThe paper gives no basis for what the AI stood in for.
About this article
Record aix-00235, version 2, checked by a person on 2026-10-09. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error