2020A&A...633A.154L


Query : 2020A&A...633A.154L

2020A&A...633A.154L - Astronomy and Astrophysics, volume 633A, 154-154 (2020/1-1)

Unsupervised star, galaxy, QSO classification. Application of HDBSCAN.

LOGAN C.H.A. and FOTOPOULOU S.

Abstract (from CDS):


Context. Classification will be an important first step for upcoming surveys aimed at detecting billions of new sources, such as LSST and Euclid, as well as DESI, 4MOST, and MOONS. The application of traditional methods of model fitting and colour-colour selections will face significant computational constraints, while machine-learning methods offer a viable approach to tackle datasets of that volume.
Aims. While supervised learning methods can prove very useful for classification tasks, the creation of representative and accurate training sets is a task that consumes a great deal of resources and time. We present a viable alternative using an unsupervised machine learning method to separate stars, galaxies and QSOs using photometric data.
Methods. The heart of our work uses Hierarchical Density-Based Spatial Clustering of Applications with Noise (HDBSCAN) to find the star, galaxy, and QSO clusters in a multidimensional colour space. We optimized the hyperparameters and input attributes of three separateHDBSCAN runs, each to select a particular object class and, thus, treat the output of each separate run as a binary classifier. We subsequently consolidated the output to give our final classifications, optimized on the basis of their F1 scores. We explored the use of Random Forest and PCA as part of the pre-processing stage for feature selection and dimensionality reduction.
Results. Using our dataset of ∼50000 spectroscopically labelled objects we obtain F1 scores of 98.9, 98.9, and 93.13 respectively for star, galaxy, and QSO selection using our unsupervised learning method. We find that careful attribute selection is a vital part of accurate classification withHDBSCAN. We applied our classification to a subset of the SDSS spectroscopic catalogue and demonstrated the potential of our approach in correcting misclassified spectra useful for DESI and 4MOST. Finally, we created a multiwavelength catalogue of 2.7 million sources using the KiDS, VIKING, and ALLWISE surveys and published corresponding classifications and photometric redshifts.

Abstract Copyright: © ESO 2020

Journal keyword(s): stars: general - galaxies: general - galaxies: active - methods: data analysis - surveys

VizieR on-line data: <Available at CDS (J/A+A/633/A154): cpz.dat klabels.dat>

Status at CDS : All or part of tables of objects will not be ingested in SIMBAD.

Simbad objects: 0

goto Full paper

goto View the references in ADS

Number of rows : 0

To bookmark this query, right click on this link: simbad:objects in 2020A&A...633A.154L and select 'bookmark this link' or equivalent in the popup menu