~/aixsci
200 records · all checked

astronomy/ai produced the result/arXiv 2025 · v2

Astronomers test pre-trained image models for sifting telescope alerts

Researchers compared off-the-shelf image-recognition networks, pre-trained on everyday photographs or on galaxy images, against a purpose-built network for deciding which nightly survey alerts are real transients. The models did the sifting and the outlier hunting themselves.

1. Recompile BTSbot alert training set2. Build fine-tuning datasets and reduced-size splits3. Fine-tune off-the-shelf architectures and retrain baseline4. Score test-split alerts with trained models5. Evaluate against labels and baseline6. Embed latent representations with UMAP7. Flag outliers with Isolation Forest8. Inspect flagged outliers

spectrum · one line per step, placed by what the step does · bright lines used AI

Pre-training vision models for the classification of alerts from wide-field time-domain surveys
arXiv, 2025

doi:10.48550/arxiv.2512.11957 · record aix-00100 v2 · checked 2026-10-08

ai-resultrole of AI
AI was for
Classification, Anomaly detection
Model family
Convolutional neural network, Transformer, Multilayer perceptron
Checked by
Held-out
Code
available

The finding the paper is about came from the AI.

read as

The science is explained before the AI appears. Switch to field specialist to go straight to the method.

Assumes the discipline and goes straight to the method.

The diagram, the record and what the paper did not report are identical in both modes. Only the framing changes — never the evidence.

Introduction by AIxSci · plain language

What this research was about

Telescopes that scan the whole sky night after night look for things that change: stars that explode, objects that flare, anything that was not there before. These are called transients. Each night a survey such as the Zwicky Transient Facility sends out hundreds of thousands of alerts, small image cutouts flagging a spot of sky that has altered. Most are not real astronomy. They are artefacts of the camera, satellite trails, bad subtractions between a new image and an older reference one. Deciding which handful deserve a follow-up telescope has traditionally meant people looking at pictures, which does not scale with the alert rate.

The researchers set out to test whether general-purpose image-recognition networks, taken off the shelf and adapted, can do this vetting as well as a network built specially for the job. They rebuilt the training set used by the existing tool, BTSbot, from the survey's alert broker and its human inspectors' accept-and-reject labels: 769,056 alerts across 25,609 sources. They then asked how performance and running cost changed with how much training data was available.

Where AI came in

Two standard vision networks, ConvNeXt-pico and MaxViT-tiny, were each started three ways: from weights learned on ImageNet, a large collection of ordinary labelled photographs; from Zoobot, whose weights come from about 842,000 galaxy images annotated by volunteers; and from scratch, with random starting values. Pre-training means a network first learns general visual structure on one large picture collection, so that less data is needed for the real task. Each was then fine-tuned to answer a yes-or-no question about an alert, and compared with the retrained custom network. Some versions also took in numerical metadata alongside the images.

The trained models then scored held-out alerts, standing in for the nightly human inspection, and their speed and memory use were measured. Separately, the internal representations the networks had formed were squeezed down to two dimensions with UMAP, and an Isolation Forest, which looks for points sitting apart from the crowd, picked out unusual alerts. Astronomers then looked at those by eye and found rare sources plus two sources that had been labelled wrongly in the training set.

Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.

The work

Technical · from the record

The authors compared vision models for vetting bright transient candidates in the Zwicky Transient Facility alert stream, using an updated BTSbot training set of 769,056 alerts across 25,609 sources. ConvNeXt-pico and MaxViT-tiny were fine-tuned from ImageNet-1k pre-training, from Zoobot weights pre-trained on about 842,000 annotated Galaxy Zoo images, and from random initialisation, and compared against a retrained custom CNN baseline across four training-set sizes. For image-only models, pre-training gave higher ROC AUC than training from scratch, and Galaxy Zoo pre-training generally gave higher ROC AUC than ImageNet pre-training, with the Galaxy Zoo MaxViT highest; for multi-modal models the baseline scored higher on trigger F1 at training sets below roughly 100,000 alerts. In the reported inference tests ConvNeXt was 6.4 times faster on CPU throughput and used 5.5 times less memory than the custom CNN, while MaxViT was 3.2 times slower and used 3.0 times more memory than it. UMAP embeddings combined with an Isolation Forest surfaced rare sources and two mislabelled sources in the training set.

How AI was used

The BTSbot training set was recompiled from the ZTF alert broker and marshal, and a parallel fine-tuning set of DESI Legacy Survey DR10 colour images was assembled at matching field of view and pixel scale, with reduced splits created by randomly keeping 5, 10 and 50 per cent of alerts. Two off-the-shelf architectures, ConvNeXt-pico and MaxViT-tiny, were initialised either from ImageNet-1k weights obtained from timm, from Zoobot weights pre-trained on Galaxy Zoo volunteer annotations and obtained from Hugging Face, or from PyTorch default initialisation; each had its head replaced with a randomly initialised three-layer MLP, and all backbone and head parameters were updated during fine-tuning on the binary transient-candidate vetting task. Multi-modal variants concatenated a metadata-branch embedding with the image embedding, and the custom CNN used by BTSbot was retrained as a baseline. A separate Bayesian hyperparameter sweep was run for each configuration on the Weights and Biases platform, after which five seeded trials were run with the best hyperparameters and scored on the test split. Trained models were then run over the test split to produce scores and to measure CPU and GPU throughput and peak memory. Penultimate-layer representations were extracted and reduced with UMAP at default settings, and an Isolation Forest was run on each labelled subset of each embedding to select candidate outliers for visual inspection.

The shape of the work

Structural · the record, drawn

ACQUISITIONPREPARATIONTRAININGINFERENCEVALIDATIONREPRESENTATIONSCREENINGINTERPRETATION12345678AIAIAIAIRecompile BTSbotalert trainingsetBuild fine-tuningdatasets andreduced-size spl…Fine-tuneoff-the-shelfarchitectures an…Score test-splitalerts withtrained modelsEvaluate againstlabels andbaselineEmbed latentrepresentationswith UMAPFlag outlierswith IsolationForestInspect flaggedoutliers↤ manual curation↤ manual curation
AI stepNo AI↤ what the AI stood in for
1Acquisition
no AI

Recompile BTSbot alert training set

Obtaining raw data, whether by measurement, download or retrieval.

The updated training set contains 769,056 alerts across 25,609 sourceswhere the paper describes this · verbatim
in the paper
2Preparation
no AI

Build fine-tuning datasets and reduced-size splits

Cleaning, filtering, normalising or labelling data already obtained.

We create smaller versions of the train split by randomly selecting 5, 10, and 50% of alerts to preservewhere the paper describes this · verbatim
in the paper
3Training
AI

Fine-tune off-the-shelf architectures and retrain baseline

Fitting model parameters, including fine-tuning an existing model.

During training, all weights and biases in the vision backbone and the MLP head are unfrozen and allowed to be updatedwhere the paper describes this · verbatim
in the paper
4Inference
AI

Score test-split alerts with trained models

Running a trained model over new data to predict, classify or score. The AI stood in for manual curation.

the time it takes to compute predictions for each of them using a batch size of 32 and 32-bit floating point precisionwhere the paper describes this · verbatim
in the paper
5Validation
no AI

Evaluate against labels and baseline

Testing outputs against ground truth.

Model performance is evaluated using the area under the receiver operating characteristic curve (ROC AUC)where the paper describes this · verbatim
in the paper
6Representation
AI

Embed latent representations with UMAP

Encoding data into features, descriptors, embeddings or graphs.

We extract high-dimensional representations from the penultimate layer of the models and fit a UMAP transformation to reduce their dimensionalitywhere the paper describes this · verbatim
in the paper
7Screening
AI

Flag outliers with Isolation Forest

Reducing a candidate set by filtering or ranking, in a single pass. The AI stood in for manual curation.

We run an Isolation Forest independently on each subsetwhere the paper describes this · verbatim
in the paper
8Interpretation
no AI

Inspect flagged outliers

Extracting understanding from model behaviour.

We visually inspect these sources and identify a number of interesting outlierswhere the paper describes this · verbatim
in the paper

What the record says

Technical · every part carries its own basis

+ in the paper~ our reading− not reported

How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.

~Role of AI
AI produced the resultour reading

The study's results are the classification performance and computational cost of the models themselves; the alert-vetting output is produced entirely by the trained models

+What the AI was for
We select two pre-training regimens to compare against training from scratch: pre-training on ImageNet and pre-training on Galaxy Zoowhere the paper describes this · verbatim
+How it was taught
SupervisedTransfer / fine-tuningUnsupervisedin the paper
+Models named
ConvNeXt-pico (ImageNet-1k pre-trained) pico · Fine-tunedConvNeXt-pico (Zoobot, Galaxy Zoo pre-trained) pico · Fine-tunedConvNeXt-pico (random initialisation) pico · Trained from scratchMaxViT-tiny (ImageNet-1k pre-trained) tiny · Fine-tunedMaxViT-tiny (Zoobot, Galaxy Zoo pre-trained) tiny · Fine-tunedMaxViT-tiny (random initialisation) tiny · Trained from scratchBTSbot custom CNN baseline (uni-modal and multi-modal) · Trained from scratchUMAP · Trained from scratchIsolation Forest · Trained from scratchin the paper
+How results were checked
Held-outin the paper
The performance metrics we report are the median and standard deviation for these final trials computed on the test splitwhere the paper describes this · verbatim
+Code · weights · data
code availableweights availabledata not reportedin the paper
we release BTSbot models presented here publicly on the Hugging Face platform, along with open-source training code on GitHubwhere the paper describes this · verbatim
+Compute
Inference benchmarks on an NVIDIA A30 GPU and 12 cores of an Intel Xeon Gold 6338 CPU; computational resources provided by the Quest high performance computing facility at Northwestern Universityin the paper

What this paper did not report

Technical · absence is published deliberately

Reported as not stated — 7 items
  • DataWhether the data are available is not stated.
  • How many were testedThe paper gives no count of what was tested.
  • Version of BTSbot custom CNN baseline (uni-modal and multi-modal)Which version of the model was used is not stated.
  • Version of UMAPWhich version of the model was used is not stated.
  • Version of Isolation ForestWhich version of the model was used is not stated.
  • What step 3 replacedThe paper gives no basis for what the AI stood in for.
  • What step 6 replacedThe paper gives no basis for what the AI stood in for.

About this article

Record aix-00100, version 2, checked by a person on 2026-10-08. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error