SearcharxivSearch

arXiv subjects

Sara Jamal

Publications and source records attributed to Sara Jamal.

5 recordsLinked to original sources

Performance analysis of extragalactic classifications in Gaia Data Release 4

The Discrete Source Classifier (DSC) provides probabilistic classifications of sources in Gaia Data Release 4 (GDR4) based on empirically-trained Bayesian classifiers. Using Gaia astrometry, photometry, and low-resolution spectra (XP), DSC classifies all sources as quasars, galaxies, or stars. DSC comprises three trained neural networks and three combinations of their probabilities. When evaluated as a function of brightness and sky position on a test set excluding the Magellanic Clouds, the DSC purity in GDR4 has improved for a small loss in completeness. The average performance of the best classifiers at magnitudes brighter than G=20 is at least 88% completeness and 96% purity for the extragalactic classes, namely the quasar and galaxy classes. At fainter magnitudes, performance is lower due to increased noise. The average performance at magnitudes of 20$\leq$G<20.5 is a minimum of 55% completeness and 71% purity for the extragalactic classes. At G>20.5 mag, completeness is considerably reduced, primarily for the models that depend on the XP spectra. Furthermore, we train additional models on Gaia optical data together with mid-infrared photometry from the CatWISE2020 catalogue. Inclusion of infrared photometry increases the completeness of extragalactic samples at G>20 mag between 9 and 29 percentage points, at the cost of reducing purity between 1 and 9 percentage points. In GDR4, the best DSC-combined classifier prioritising completeness identifies three million quasars and two million galaxies, but with expected high contamination among fainter sources. In contrast, the combined classifiers prioritising purity identify approximately two million quasars and 1.3 million galaxies with an expected lower level of contamination. Finally, we provide recommendations for enhancing the purity of the DSC extragalactic selection by applying quality cuts to the Gaia photometry and astrometry.

astro-ph.GA

Improved source classification and performance analysis using Gaia DR3

The Discrete Source Classifier (DSC) provides probabilistic classification of sources in Gaia Data Release 3 using a Bayesian framework and a global prior. The DSC Combmod classifier in GDR3 achieved for the extragalactic classes (quasars and galaxies) a high completeness of 92%, but a low purity of 22% due to contamination from the far larger star class. However, these single metrics mask significant variation in performance with magnitude and sky position. Furthermore, a better combination of the individual classifiers is possible. Here we compute two-dimensional representations of the completeness and the purity as function of Galactic latitude and source brightness, and also exclude the Magellanic Clouds where stellar contamination significantly reduces the purity. Reevaluated on a cleaner validation set and without introducing changes to the published GDR3 DSC probabilities themselves, we achieve for Combmod average 2D completenesses of 92% and 95% and average 2D purities of 55% and 89% for the quasar and galaxy classes, respectively. Since the relative proportions of extragalactic objects to stars in Gaia is expected to vary significantly with brightness and latitude, we introduce a new prior as a continuous function of brightness and latitude, and compute new class probabilities. This variable prior only improves the performance by a few percentage points, mostly at the faint end. Significant improvement, however, is obtained by a new additive combination of Specmod and Allosmod. This classifier, Combmod-$\alpha$, achieves average 2D completenesses of 82% and 93% and average 2D purities of 79% and 93% for the quasar and galaxy classes, respectively, when using the global prior. Thus, we achieve a significant improvement in purity for a small loss of completeness. The improvement is most significant for faint quasars where the purity rises from 20% to 62%.

astro-ph.GA

Quasar and galaxy classification using Gaia EDR3 and CatWise2020

In this work, we assess the combined use of Gaia photometry and astrometry with infrared data from CatWISE in improving the identification of extragalactic sources compared to the classification obtained using Gaia data. We evaluate different input feature configurations and prior functions, with the aim of presenting a classification methodology integrating prior knowledge stemming from realistic class distributions in the universe. In our work, we compare different classifiers, namely Gaussian Mixture Models (GMMs), XGBoost and CatBoost, and classify sources into three classes - star, quasar, and galaxy, with the target quasar and galaxy class labels obtained from SDSS16 and the star label from Gaia EDR3. In our approach, we adjust the posterior probabilities to reflect the intrinsic distribution of extragalactic sources in the universe via a prior function. We introduce two priors, a global prior reflecting the overall rarity of quasars and galaxies, and a mixed prior that incorporates in addition the distribution of the these sources as a function of Galactic latitude and magnitude. Our best classification performances, in terms of completeness and purity of the galaxy and quasar classes, are achieved using the mixed prior for sources at high latitudes and in the magnitude range G = 18.5 to 19.5. We apply our identified best-performing classifier to three application datasets from Gaia DR3, and find that the global prior is more conservative in what it considers to be a quasar or a galaxy compared to the mixed prior. In particular, when applied to the pure quasar and galaxy candidates samples, we attain a purity of 97% for quasars and 99.9% for galaxies using the global prior, and purities of 96% and 99% respectively using the mixed prior. We conclude our work by discussing the importance of applying adjusted priors portraying realistic class distributions in the universe.

astro-ph.GA

Identification of high order closure terms from fully kinetic simulations using machine learning

Simulations of large-scale plasma systems are typically based on a fluid approximation approach. These models construct a moment-based system of equations that approximate the particle-based physics as a fluid, but as a result lack the small-scale physical processes available to fully kinetic models. Traditionally, empirical closure relations are used to close the moment-based system of equations, which typically approximate the pressure tensor or heat flux. The more accurate the closure relation, the stronger the simulation approaches kinetic-based results. In this paper, new closure terms are constructed using machine learning techniques. Two different machine learning models, a multi-layer perceptron and a gradient boosting regressor, synthesize a local closure relation for the pressure tensor and heat flux vector from fully kinetic simulations of a 2D magnetic reconnection problem. The models are compared to an existing closure relation for the pressure tensor, and the applicability of the models is discussed. The initial results show that the models can capture the diagonal components of the pressure tensor accurately, and show promising results for the heat flux, opening the way for new experiments in multi-scale modeling. We find that the sampling of the points used to train both models play a capital role in their accuracy.

physics.plasm-ph

On Neural Architectures for Astronomical Time-series Classification with Application to Variable Stars

Despite the utility of neural networks (NNs) for astronomical time-series classification, the proliferation of learning architectures applied to diverse datasets has thus far hampered a direct intercomparison of different approaches. Here we perform the first comprehensive study of variants of NN-based learning and inference for astronomical time-series, aiming to provide the community with an overview on relative performance and, hopefully, a set of best-in-class choices for practical implementations. In both supervised and self-supervised contexts, we study the effects of different time-series-compatible layer choices, namely the dilated temporal convolutional neural network (dTCNs), Long-Short Term Memory (LSTM) NNs, Gated Recurrent Units (GRUs) and temporal convolutional NNs (tCNNs). We also study the efficacy and performance of encoder-decoder (i.e., autoencoder) networks compared to direct classification networks, different pathways to include auxiliary (non-time-series) metadata, and different approaches to incorporate multi-passband data (i.e., multiple time-series per source). Performance---applied to a sample of 17,604 variable stars from the MACHO survey across 10 imbalanced classes---is measured in training convergence time, classification accuracy, reconstruction error, and generated latent variables. We find that networks with Recurrent NN (RNNs) generally outperform dTCNs and, in many scenarios, yield to similar accuracy as tCNNs. In learning time and memory requirements, convolution-based layers are more performant. We conclude by discussing the advantages and limitations of deep architectures for variable star classification, with a particular eye towards next-generation surveys such as LSST, WFIRST and ZTF2.

astro-ph.IM