astronomy/ai produced the result/Astronomy and Astrophysics 2022 · v2
Volunteer labels train a model to find asteroid trails in Hubble archive images
Volunteers marked asteroid streaks in archival Hubble images, and those marks trained an object-detection model that was then run across the archive. The combined search yielded 1 701 trails, 670 of them linked to known Solar System objects.
spectrum · one line per step, placed by what the step does · bright lines used AI
HubbleAsteroid Hunter
Astronomy and Astrophysics, 2022
doi:10.1051/0004-6361/202142998 · record aix-00106 v2 · checked 2026-10-08
- AI was for
- Detection, Classification
- Model family
- Convolutional neural network, Clustering
- Checked by
- Held-out
- Code
- not reported
The finding the paper is about came from the AI.
What this research was about
Asteroids are rocks orbiting the Sun, and most of them are small and faint. When a telescope stares at a distant galaxy for a long time, an asteroid drifting through the field of view does not appear as a dot. Because it moves while the shutter is open, it leaves a curved streak, or trail, across the picture. Archives of old telescope images therefore hold incidental records of asteroids that nobody was looking for. The difficulty is volume. The Hubble Space Telescope archive contains tens of thousands of images, and a trail can look much like a cosmic ray hit, a satellite streak or a stretched arc of gravitationally bent light.
The Hubble Asteroid Hunter project set out to comb that archive for such trails. Composite images from two Hubble cameras were filtered by exposure time and field of view, then cut into quadrants. Those cutouts were shown to volunteers on the Zooniverse platform, who said whether a trail was present and marked where each one started and ended. The team then measured the positions and brightnesses of the trails that survived checking, and tried to match each one to an asteroid whose orbit is already catalogued.
Where AI came in
Two pieces of software did work that people would otherwise have done by hand. First, a clustering algorithm called HDBSCAN tidied up the volunteers' clicks: many people marked the same trail, and the algorithm grouped nearby marks into a single agreed start and end point, producing clean labels. Second, those labels, together with forum tags for satellites, cosmic rays and lens arcs, were used to train an object-detection model with Google Cloud AutoML Vision. The model learns from examples to draw a box around anything it thinks is a trail and to say which of the four classes it is, with a confidence score.
The trained detector was then run over 149 292 cutouts, including images added to the archive after the volunteers had finished. It returned 2 041 asteroid trails, of which 997 had not been found by volunteers, so the final sample depends on the model as well as the people. Its labelled examples were split into training, validation and test portions so performance could be checked on images it had not learned from. Everything after detection was done without learned models: three of the authors inspected the candidates by eye, and the measurement and ephemeris matching used conventional software.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record
Volunteers on the Hubble Asteroid Hunter citizen science project provided 1 783 873 classifications of 144 559 cutouts from archival Hubble Space Telescope images, which yielded 1 488 asteroid trail labels. Those labels, together with forum tags for satellites, cosmic rays and gravitational lens arcs, were used to train a Google Cloud AutoML Vision object detection model, which was then run over 149 292 cutouts and returned 2 041 asteroid trails, 997 of which had not been found by volunteers. After visual inspection by three of the authors and removal of cosmic rays, false identifications and targeted observations, the study reports a final sample of 1 701 trails in 1 316 composite images, with trails detected to a magnitude of 24.5. Matching trail positions against ephemerides linked 670 trails to 454 known Solar System objects, while 1 031 trails were not matched and are on average 1.6 magnitudes fainter than the matched ones.
How AI was used
Composite ACS/WFC and WFC3/UVIS images from the ESA Hubble archive were filtered by exposure time and field of view and split into four quadrant cutouts, which volunteers on the Zooniverse platform classified for the presence of asteroid trails, marking trail start and end points. Per-cutout positive classification probabilities above 0.5 were taken as detections, and the volunteers' click positions were aggregated with HDBSCAN point clustering using a minimum cluster size of 5 and minimum samples of 5. The aggregated asteroid labels, plus cosmic ray, gravitational lens arc and satellite labels taken from project forum tags, were used to train a Google Cloud AutoML Vision object detection model with four classes; AutoML performs its own neural architecture search, described in the paper as based on reinforcement learning, and the labelled sample was split 70% training, 15% validation and 15% test by random selection, with the validation split used by AutoML for preprocessing, architecture and hyperparameter tuning. The trained detector, which returns bounding boxes with class scores, was applied at a 50% confidence threshold to cutouts covering both the volunteer-inspected set and images added to the archive up to 14 March 2021, and its detections were positionally cross-matched with the volunteer detections using tolerances of 4 arcsec for WFC3/UVIS and 5 arcsec for ACS/WFC. All subsequent steps used no learned models: visual inspection of candidates by three authors, trail extraction and photometry implemented in GNU Octave on re-drizzled .fits cutouts, and matching of trail coordinates to SkyBoT candidates and JPL Horizons ephemerides computed at one-minute steps.
The shape of the work
Structural · the record, drawn
no AI
Select archival HST images and split into cutouts
Obtaining raw data, whether by measurement, download or retrieval.
We used a total of 37 323 HST composite images in PNG format available from the eHST archivewhere the paper describes this · verbatim
no AI
Volunteer classification of cutouts
Cleaning, filtering, normalising or labelling data already obtained.
11 482 volunteers provided 1 783 873 classifications for 144 559 cutoutswhere the paper describes this · verbatim
AI
Aggregate volunteer votes and markings into labels
Cleaning, filtering, normalising or labelling data already obtained. The AI stood in for manual curation.
using a point clustering algorithm, Hierarchical Density-Based Spatial Clustering of Applications with Noisewhere the paper describes this · verbatim
AI
Train AutoML object detection model
Fitting model parameters, including fine-tuning an existing model. The AI stood in for expert judgement.
we trained the model with four labels: satellite, asteroid, gravitational lens arc, and cosmic raywhere the paper describes this · verbatim
AI
Run detector over the archive
Running a trained model over new data to predict, classify or score. The AI stood in for manual curation.
Our AutoML classification of 149 292 cutoutswhere the paper describes this · verbatim
no AI
Visual inspection of candidate trails
Testing outputs against ground truth.
All of them were visually inspected by three of the authorswhere the paper describes this · verbatim
no AI
Extract trail astrometry and photometry
Cleaning, filtering, normalising or labelling data already obtained.
The Trail Extractor pipeline is used to retrieve the trail from the.fits files and to obtain the calibrated astrometric and photometric datawhere the paper describes this · verbatim
no AI
Match trails against known Solar System objects
Testing outputs against ground truth.
we found 670 matches (trails associated with known SSOs)where the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
The reported sample of trails comes from the combination of volunteer labels and the AutoML detector, which contributed candidates not found by volunteers; the detection result is therefore partly produced by the model.
We used the labels provided by the volunteers to train an automated deep learning model built with Google Cloud AutoML Visionwhere the paper describes this · verbatim
We split the sample into 70% training set, 15% validation, and 15% test set, using a random selectionwhere the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- CodeWhether the code is available is not stated.
- Trained model weightsWhether the trained model is available is not stated.
- DataWhether the data are available is not stated.
- ComputeThe hardware or time used is not stated.
- How many were testedThe paper gives no count of what was tested.
- Version of Google Cloud AutoML Vision (object detection)Which version of the model was used is not stated.
- Version of HDBSCANWhich version of the model was used is not stated.
About this article
Record aix-00106, version 2, checked by a person on 2026-10-08. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error