~/aixsci
200 records · all checked

astronomy/ai produced the result/Astronomy and Astrophysics 2022 · v2

Volunteer labels train a model to find asteroid trails in Hubble archive images

Volunteers marked asteroid streaks in archival Hubble images, and those marks trained an object-detection model that was then run across the archive. The combined search yielded 1 701 trails, 670 of them linked to known Solar System objects.

1. Select archival HST images and split into cutouts2. Volunteer classification of cutouts3. Aggregate volunteer votes and markings into labels4. Train AutoML object detection model5. Run detector over the archive6. Visual inspection of candidate trails7. Extract trail astrometry and photometry8. Match trails against known Solar System objects

spectrum · one line per step, placed by what the step does · bright lines used AI

HubbleAsteroid Hunter
Astronomy and Astrophysics, 2022

doi:10.1051/0004-6361/202142998 · record aix-00106 v2 · checked 2026-10-08

ai-resultrole of AI
AI was for
Detection, Classification
Model family
Convolutional neural network, Clustering
Checked by
Held-out
Code
not reported

The finding the paper is about came from the AI.

read as

The science is explained before the AI appears. Switch to field specialist to go straight to the method.

Assumes the discipline and goes straight to the method.

The diagram, the record and what the paper did not report are identical in both modes. Only the framing changes — never the evidence.

Introduction by AIxSci · plain language

What this research was about

Asteroids are rocks orbiting the Sun, and most of them are small and faint. When a telescope stares at a distant galaxy for a long time, an asteroid drifting through the field of view does not appear as a dot. Because it moves while the shutter is open, it leaves a curved streak, or trail, across the picture. Archives of old telescope images therefore hold incidental records of asteroids that nobody was looking for. The difficulty is volume. The Hubble Space Telescope archive contains tens of thousands of images, and a trail can look much like a cosmic ray hit, a satellite streak or a stretched arc of gravitationally bent light.

The Hubble Asteroid Hunter project set out to comb that archive for such trails. Composite images from two Hubble cameras were filtered by exposure time and field of view, then cut into quadrants. Those cutouts were shown to volunteers on the Zooniverse platform, who said whether a trail was present and marked where each one started and ended. The team then measured the positions and brightnesses of the trails that survived checking, and tried to match each one to an asteroid whose orbit is already catalogued.

Where AI came in

Two pieces of software did work that people would otherwise have done by hand. First, a clustering algorithm called HDBSCAN tidied up the volunteers' clicks: many people marked the same trail, and the algorithm grouped nearby marks into a single agreed start and end point, producing clean labels. Second, those labels, together with forum tags for satellites, cosmic rays and lens arcs, were used to train an object-detection model with Google Cloud AutoML Vision. The model learns from examples to draw a box around anything it thinks is a trail and to say which of the four classes it is, with a confidence score.

The trained detector was then run over 149 292 cutouts, including images added to the archive after the volunteers had finished. It returned 2 041 asteroid trails, of which 997 had not been found by volunteers, so the final sample depends on the model as well as the people. Its labelled examples were split into training, validation and test portions so performance could be checked on images it had not learned from. Everything after detection was done without learned models: three of the authors inspected the candidates by eye, and the measurement and ephemeris matching used conventional software.

Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.

The work

Technical · from the record

Volunteers on the Hubble Asteroid Hunter citizen science project provided 1 783 873 classifications of 144 559 cutouts from archival Hubble Space Telescope images, which yielded 1 488 asteroid trail labels. Those labels, together with forum tags for satellites, cosmic rays and gravitational lens arcs, were used to train a Google Cloud AutoML Vision object detection model, which was then run over 149 292 cutouts and returned 2 041 asteroid trails, 997 of which had not been found by volunteers. After visual inspection by three of the authors and removal of cosmic rays, false identifications and targeted observations, the study reports a final sample of 1 701 trails in 1 316 composite images, with trails detected to a magnitude of 24.5. Matching trail positions against ephemerides linked 670 trails to 454 known Solar System objects, while 1 031 trails were not matched and are on average 1.6 magnitudes fainter than the matched ones.

How AI was used

Composite ACS/WFC and WFC3/UVIS images from the ESA Hubble archive were filtered by exposure time and field of view and split into four quadrant cutouts, which volunteers on the Zooniverse platform classified for the presence of asteroid trails, marking trail start and end points. Per-cutout positive classification probabilities above 0.5 were taken as detections, and the volunteers' click positions were aggregated with HDBSCAN point clustering using a minimum cluster size of 5 and minimum samples of 5. The aggregated asteroid labels, plus cosmic ray, gravitational lens arc and satellite labels taken from project forum tags, were used to train a Google Cloud AutoML Vision object detection model with four classes; AutoML performs its own neural architecture search, described in the paper as based on reinforcement learning, and the labelled sample was split 70% training, 15% validation and 15% test by random selection, with the validation split used by AutoML for preprocessing, architecture and hyperparameter tuning. The trained detector, which returns bounding boxes with class scores, was applied at a 50% confidence threshold to cutouts covering both the volunteer-inspected set and images added to the archive up to 14 March 2021, and its detections were positionally cross-matched with the volunteer detections using tolerances of 4 arcsec for WFC3/UVIS and 5 arcsec for ACS/WFC. All subsequent steps used no learned models: visual inspection of candidates by three authors, trail extraction and photometry implemented in GNU Octave on re-drizzled .fits cutouts, and matching of trail coordinates to SkyBoT candidates and JPL Horizons ephemerides computed at one-minute steps.

The shape of the work

Structural · the record, drawn

ACQUISITIONPREPARATIONPREPARATIONTRAININGINFERENCEVALIDATIONPREPARATIONVALIDATION12345678AIAIAISelect archivalHST images andsplit into cutou…Volunteerclassification ofcutoutsAggregatevolunteer votesand markings int…Train AutoMLobject detectionmodelRun detector overthe archiveVisual inspectionof candidatetrailsExtract trailastrometry andphotometryMatch trailsagainst knownSolar System obj…↤ manual curation↤ expert judgement↤ manual curation
AI stepNo AI↤ what the AI stood in for
1Acquisition
no AI

Select archival HST images and split into cutouts

Obtaining raw data, whether by measurement, download or retrieval.

We used a total of 37 323 HST composite images in PNG format available from the eHST archivewhere the paper describes this · verbatim
in the paper
2Preparation
no AI

Volunteer classification of cutouts

Cleaning, filtering, normalising or labelling data already obtained.

11 482 volunteers provided 1 783 873 classifications for 144 559 cutoutswhere the paper describes this · verbatim
in the paper
3Preparation
AI

Aggregate volunteer votes and markings into labels

Cleaning, filtering, normalising or labelling data already obtained. The AI stood in for manual curation.

using a point clustering algorithm, Hierarchical Density-Based Spatial Clustering of Applications with Noisewhere the paper describes this · verbatim
in the paper
4Training
AI

Train AutoML object detection model

Fitting model parameters, including fine-tuning an existing model. The AI stood in for expert judgement.

we trained the model with four labels: satellite, asteroid, gravitational lens arc, and cosmic raywhere the paper describes this · verbatim
in the paper
5Inference
AI

Run detector over the archive

Running a trained model over new data to predict, classify or score. The AI stood in for manual curation.

Our AutoML classification of 149 292 cutoutswhere the paper describes this · verbatim
in the paper
6Validation
no AI

Visual inspection of candidate trails

Testing outputs against ground truth.

All of them were visually inspected by three of the authorswhere the paper describes this · verbatim
in the paper
7Preparation
no AI

Extract trail astrometry and photometry

Cleaning, filtering, normalising or labelling data already obtained.

The Trail Extractor pipeline is used to retrieve the trail from the.fits files and to obtain the calibrated astrometric and photometric datawhere the paper describes this · verbatim
in the paper
8Validation
no AI

Match trails against known Solar System objects

Testing outputs against ground truth.

we found 670 matches (trails associated with known SSOs)where the paper describes this · verbatim
in the paper

What the record says

Technical · every part carries its own basis

+ in the paper~ our reading− not reported

How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.

~Role of AI
AI produced the resultour reading

The reported sample of trails comes from the combination of volunteer labels and the AutoML detector, which contributed candidates not found by volunteers; the detection result is therefore partly produced by the model.

+What the AI was for
We used the labels provided by the volunteers to train an automated deep learning model built with Google Cloud AutoML Visionwhere the paper describes this · verbatim
+Model families
+How it was taught
SupervisedUnsupervisedReinforcementin the paper
+Models named
Google Cloud AutoML Vision (object detection) · Trained from scratchHDBSCAN · Off the shelfin the paper
+How results were checked
Held-outin the paper
We split the sample into 70% training set, 15% validation, and 15% test set, using a random selectionwhere the paper describes this · verbatim
−Code · weights · data
code not reportedweights not reporteddata not reportednot reported
−Compute
not reportednot reported

What this paper did not report

Technical · absence is published deliberately

Reported as not stated — 7 items
  • CodeWhether the code is available is not stated.
  • Trained model weightsWhether the trained model is available is not stated.
  • DataWhether the data are available is not stated.
  • ComputeThe hardware or time used is not stated.
  • How many were testedThe paper gives no count of what was tested.
  • Version of Google Cloud AutoML Vision (object detection)Which version of the model was used is not stated.
  • Version of HDBSCANWhich version of the model was used is not stated.

About this article

Record aix-00106, version 2, checked by a person on 2026-10-08. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error