~/aixsci
200 records · all checked

astronomy/ai produced the result/arXiv 2024 · v2

Clustering algorithm sorts cluster stars from background stars in thirteen open clusters

Astronomers used an unsupervised Gaussian Mixture Model on Gaia DR3 motions and distances to separate the members of thirteen open clusters from unrelated field stars, then measured how well the model performed.

1. Cone-search query of Gaia DR3 around each cluster2. Quality filtering, proper-motion cut and feature normalisation3. Characterise GMM behaviour on simulated member and field populations4. Fit two-component GMM and assign membership probabilities5. Search distance cutoff and probability threshold by MSS6. Compare member samples with a published catalogue and with spectroscopic abundances7. Derive cluster parameters from revised membership with ASteCA8. Study dependence of MSS on cluster properties

spectrum · one line per step, placed by what the step does · bright lines used AI

Using GMM in Open Cluster Membership: An Insight
arXiv, 2024

doi:10.48550/arxiv.2401.10802 · record aix-00135 v2 · checked 2026-10-09

ai-resultrole of AI
AI was for
Classification
Model family
Clustering
Checked by
Benchmark
Code
available

The finding the paper is about came from the AI.

read as

The science is explained before the AI appears. Switch to field specialist to go straight to the method.

Assumes the discipline and goes straight to the method.

The diagram, the record and what the paper did not report are identical in both modes. Only the framing changes — never the evidence.

Introduction by AIxSci · plain language

What this research was about

An open cluster is a loose family of stars born together from the same cloud of gas. Because they share a birthplace, they also share a journey: they sit at roughly the same distance and drift across the sky at roughly the same speed. That shared motion is the main clue to who belongs. The difficulty is that a cluster is seen through a crowd. Many unrelated stars, called field stars, lie along the same line of sight and simply look like part of the group. Telling members from interlopers has traditionally meant looking for crowded patches of sky, inspecting brightness-and-colour plots, and checking stellar motions by hand.

The researchers set out to do that sorting statistically instead, using the Gaia satellite's catalogue of stellar positions, motions and distances. They applied the same method to thirteen clusters spanning a wide range of ages and distances, from 441 to 5183 parsecs, and asked how the model's performance depended on a cluster's properties. They also defined their own score, the Modified Silhouette Score, to judge how cleanly the two groups had been pulled apart.

Where AI came in

The AI here is a Gaussian Mixture Model, an unsupervised clustering method. Unsupervised means nobody told it which stars were members; it was given no examples to copy. It assumes the data are a blend of two overlapping bell-shaped populations and works out, by repeated adjustment, which blend best fits what it sees. It was fitted from scratch on just three measurements per star: the two components of proper motion, meaning the star's drift across the sky, and parallax, the small wobble that encodes distance. Sky position was deliberately left out, because members and field stars pile up in the same place there. The model then gave every star a probability of belonging.

The group with the tighter spread in those three measurements was labelled as members. Before working on real data, the model was run on simulated mixtures of members and scattered field stars to see how its separation depended on the range of values it was shown. For each real cluster, the researchers scanned distance limits and probability thresholds, keeping the setting that scored best. The resulting member lists were compared with a previously published catalogue and with spectroscopic measurements of stellar chemistry, and were then fed to the ASteCA software to estimate each cluster's properties. In short, the model stood in for the hand-worked judgements that membership lists had relied on.

Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.

The work

Technical · from the record

The authors applied an unsupervised two-component Gaussian Mixture Model to Gaia DR3 proper motions and parallaxes for thirteen open clusters spanning log t of about 6.38 to 9.64 and distances of 441 to 5183 pc, in order to separate cluster members from field stars. They defined a Modified Silhouette Score (MSS), based on the ratio of field-star to member standard deviations across features, and used it both to pick the parallax and proper-motion cutoffs and the membership probability threshold, and to assess performance. Members were found down to G of about 20 mag, compared with about 18 mag in the published catalogue used as a benchmark, and cluster parameters were then derived with ASteCA. Comparing MSS across clusters, the authors report no significant difference between younger and older clusters but a moderate correlation with distance, with lower MSS for clusters beyond about 3 kpc.

How AI was used

Stars were drawn from Gaia DR3 by cone search around each cluster centre, filtered on parallax sign, parallax over error greater than 3 and proper motion errors, restricted to proper motions between -20 and 20 mas/yr, and normalised. A two-component Gaussian Mixture Model, fitted by expectation-maximisation with full covariance and five random initialisations, was run on three features only — pmra, pmdec and parallax — with sky coordinates deliberately excluded because member and field distributions peak together there. The group with the lower per-feature standard deviation was labelled as members, and the model supplied a membership probability for each star. Before this, the same model was run on simulated normally distributed members plus randomly distributed field stars, over 40 feature half-widths from 2σ to 10σ with 20 trials each and varying field-star counts, to characterise how the feature range affects separation. For each real cluster the distance filter was then scanned from d±50 pc up to d±1400 pc, and the membership probability threshold likewise varied, with the Modified Silhouette Score recorded at each setting and the best-scoring setting retained; members were taken above the chosen probability threshold and non-members below 0.2. The resulting member samples were compared with a published Gaia DR2 membership catalogue and with APOGEE and GALAH abundances, and passed to ASteCA for cluster parameter determination.

The shape of the work

Structural · the record, drawn

ACQUISITIONPREPARATIONSIMULATIONTRAININGOPTIMISATIONVALIDATIONINTERPRETATIONINTERPRETATION12345678AIAICone-search queryof Gaia DR3around each clus…Qualityfiltering,proper-motion cu…Characterise GMMbehaviour onsimulated member…Fit two-componentGMM and assignmembership proba…Search distancecutoff andprobability thre…Compare membersamples with apublished catalo…Derive clusterparameters fromrevised membersh…Study dependenceof MSS on clusterproperties↤ conventional algorithmloops back
AI stepNo AI↤ what the AI stood in for
1Acquisition
no AI

Cone-search query of Gaia DR3 around each cluster

Obtaining raw data, whether by measurement, download or retrieval.

we query the data using a cone search for a specific position in the skywhere the paper describes this · verbatim
in the paper
2Preparation
no AI

Quality filtering, proper-motion cut and feature normalisation

Cleaning, filtering, normalising or labelling data already obtained.

Then we normalized each of the features before feeding it to our model.where the paper describes this · verbatim
in the paper
3Simulation
AI

Characterise GMM behaviour on simulated member and field populations

Numerical or physics simulation, including where a learned surrogate replaces it.

We used a simulated dataset of normally distributed members and randomly distributed field starswhere the paper describes this · verbatim
in the paper
4Training
AI

Fit two-component GMM and assign membership probabilities

Fitting model parameters, including fine-tuning an existing model. The AI stood in for conventional algorithm.

We ran the 2-component GMM algorithm with 5 different initial conditions and chose the best onewhere the paper describes this · verbatim
in the paper
5Optimisation
no AI

Search distance cutoff and probability threshold by MSS

Iterative search over a space. Its result feeds back into an earlier step.

We used an empirical approach to choose the optimal distance cutoff.where the paper describes this · verbatim
in the paper
6Validation
no AI

Compare member samples with a published catalogue and with spectroscopic abundances

Testing outputs against ground truth.

We compared the chemical abundances of our members and the members found by using APOGEE and GALAH datawhere the paper describes this · verbatim
in the paper
7Interpretation
no AI

Derive cluster parameters from revised membership with ASteCA

Extracting understanding from model behaviour.

We use ASteca to determine parameters for the clusters from our revised membership data.where the paper describes this · verbatim
in the paper
8Interpretation
no AI

Study dependence of MSS on cluster properties

Extracting understanding from model behaviour.

We study the dependence of MSS on age, distance, extinction, galactic latitude and longitude, and other parameterswhere the paper describes this · verbatim
in the paper

What the record says

Technical · every part carries its own basis

+ in the paper~ our reading− not reported

How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.

~Role of AI
AI produced the resultour reading

The membership samples the paper reports, and the cluster parameters derived from them, come from the GMM clustering itself; the paper's subject is the performance of that model.

+What the AI was for
Classificationin the paper
we apply the unsupervised Gaussian Mixture Model (GMM) to a sample of thirteen clusterswhere the paper describes this · verbatim
+Model families
Clusteringin the paper
+How it was taught
Unsupervisedin the paper
+Models named
Gaussian Mixture Model (2-component, full covariance) · Trained from scratchin the paper
+How results were checked
Benchmarkin the paper
We use it as a benchmark and find fainter members at the low mass end G≈20where the paper describes this · verbatim
+Code · weights · data
code availableweights not reporteddata not reportedin the paper
All the code for GMM is available at https://github.com/mahmud-nobe/Cluster-Membership/tree/master/GMMwhere the paper describes this · verbatim
+Compute
not reportedin the paper

What this paper did not report

Technical · absence is published deliberately

Reported as not stated — 6 items
  • Trained model weightsWhether the trained model is available is not stated.
  • DataWhether the data are available is not stated.
  • ComputeThe hardware or time used is not stated.
  • How many were testedThe paper gives no count of what was tested.
  • Version of Gaussian Mixture Model (2-component, full covariance)Which version of the model was used is not stated.
  • What step 3 replacedThe paper gives no basis for what the AI stood in for.

About this article

Record aix-00135, version 2, checked by a person on 2026-10-09. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error