SearcharxivSearch

arXiv subjects

Gursharanjit Kaur

Publications and source records attributed to Gursharanjit Kaur.

5 recordsLinked to original sources

Exploring Machine Learning Regression Models for Advancing Foreground Mitigation and Global 21cm Signal Parameter Extraction

Extracting parameters from the global 21cm signal is crucial for understanding the early Universe. However, detecting the 21cm signal is challenging due to the brighter foreground and associated observational difficulties. In this study, we evaluate the performance of various machine-learning regression models to improve parameter extraction and foreground removal. This evaluation is essential for selecting the most suitable machine learning regression model based on computational efficiency and predictive accuracy. We compare four models: Random Forest Regressor (RFR), Gaussian Process Regressor (GPR), Support Vector Regressor (SVR), and Artificial Neural Networks (ANN). The comparison is based on metrics such as the root mean square error (RMSE) and $R^2$ scores. We examine their effectiveness across different dataset sizes and conditions, including scenarios with foreground contamination. Our results indicate that ANN consistently outperforms the other models, achieving the lowest RMSE and the highest $R^2$ scores across multiple cases. While GPR also performs well, it is computationally intensive, requiring significant RAM and longer execution times. SVR struggles with large datasets due to its high computational costs, and RFR demonstrates the weakest accuracy among the models tested. We also found that employing Principal Component Analysis (PCA) as a preprocessing step significantly enhances model performance, especially in the presence of foregrounds.

astro-ph.CO

The fifth data release of the Kilo Degree Survey: Multi-epoch optical/NIR imaging covering wide and legacy-calibration fields

We present the final data release of the Kilo-Degree Survey (KiDS-DR5), a public European Southern Observatory (ESO) wide-field imaging survey optimised for weak gravitational lensing studies. We combined matched-depth multi-wavelength observations from the VLT Survey Telescope and the VISTA Kilo-degree INfrared Galaxy (VIKING) survey to create a nine-band optical-to-near-infrared survey spanning $1347$ deg$^2$. The median $r$-band $5σ$ limiting magnitude is 24.8 with median seeing $0.7^{\prime\prime}$. The main survey footprint includes $4$ deg$^2$ of overlap with existing deep spectroscopic surveys. We complemented these data in DR5 with a targeted campaign to secure an additional $23$ deg$^2$ of KiDS- and VIKING-like imaging over a range of additional deep spectroscopic survey fields. From these fields, we extracted a catalogue of $126\,085$ sources with both spectroscopic and photometric redshift information, which enables the robust calibration of photometric redshifts across the full survey footprint. In comparison to previous releases, DR5 represents a $34\%$ areal extension and includes an $i$-band re-observation of the full footprint, thereby increasing the effective $i$-band depth by $0.4$ magnitudes and enabling multi-epoch science. Our processed nine-band imaging, single- and multi-band catalogues with masks, and homogenised photometry and photometric redshifts can be accessed through the ESO Archive Science Portal.

astro-ph.GA

Target Selection for the Redshift-Limited WAVES-Wide with Machine Learning

The forthcoming Wide Area Vista Extragalactic Survey (WAVES) on the 4-metre Multi-Object Spectroscopic Telescope (4MOST) has a key science goal of probing the halo mass function to lower limits than possible with previous surveys. For that purpose, in its Wide component, galaxies targetted by WAVES will be flux-limited to $Z<21.1$ mag and will cover the redshift range of $z<0.2$, at a spectroscopic success rate of $\sim95\%$. Meeting this completeness requirement, when the redshift is unknown a priori, is a challenge. We solve this problem with supervised machine learning to predict the probability of a galaxy falling within the WAVES-Wide redshift limit, rather than estimate each object's redshift. This is done by training an XGBoost tree-based classifier to decide if a galaxy should be a target or not. Our photometric data come from 9-band VST+VISTA observations, including KiDS+VIKING surveys. The redshift labels for calibration are derived from an extensive spectroscopic sample overlapping with KiDS and ancillary fields. Our current results indicate that with our approach, we should be able to achieve the completeness of $\sim95\%$, which is the WAVES success criterion.

astro-ph.CO

Wide Area VISTA Extra-galactic Survey (WAVES): Unsupervised star-galaxy separation on the WAVES-Wide photometric input catalogue using UMAP and ${\rm{\scriptsize HDBSCAN}}$

Star-galaxy separation is a crucial step in creating target catalogues for extragalactic spectroscopic surveys. A classifier biased towards inclusivity risks including spurious stars, wasting fibre hours, while a more conservative classifier might overlook galaxies, compromising completeness and hence survey objectives. To avoid bias introduced by a training set in supervised methods, we employ an unsupervised machine learning approach. Using photometry from the Wide Area VISTA Extragalactic Survey (WAVES)-Wide catalogue comprising 9-band $u-K_s$ data, we create a feature space with colours, fluxes, and apparent size information extracted by ${\rm P{\scriptsize RO} F{\scriptsize OUND}}$. We apply the non-linear dimensionality reduction method UMAP (Uniform Manifold Approximation and Projection) combined with the classifier ${\rm{\scriptsize HDBSCAN}}$ to classify stars and galaxies. Our method is verified against a baseline colour and morphological method using a truth catalogue from Gaia, SDSS, GAMA, and DESI. We correctly identify 99.72% of galaxies within the AB magnitude limit of $Z = 21.2$, with an F1 score of $0.9971 \pm 0.0018$ across the entire ground truth sample, compared to $0.9879 \pm 0.0088$ from the baseline method. Our method's higher purity ($0.9967 \pm 0.0021$) compared to the baseline ($0.9795 \pm 0.0172$) increases efficiency, identifying 11% fewer galaxy or ambiguous sources, saving approximately 70,000 fibre hours on the 4MOST instrument. We achieve reliable classification statistics for challenging sources including quasars, compact galaxies, and low surface brightness galaxies, retrieving 95.1%, 84.6%, and 99.5% of them respectively. Angular clustering analysis validates our classifications, showing consistency with expected galaxy clustering, regardless of the baseline classification.

astro-ph.GA

Comparing sampling techniques to chart parameter space of 21 cm Global signal with Artificial Neural Networks

Understanding the first billion years of the universe requires studying two critical epochs: the Epoch of Reionization (EoR) and Cosmic Dawn (CD). However, due to limited data, the properties of the Intergalactic Medium (IGM) during these periods remain poorly understood, leading to a vast parameter space for the global 21cm signal. Training an Artificial Neural Network (ANN) with a narrowly defined parameter space can result in biased inferences. To mitigate this, the training dataset must be uniformly drawn from the entire parameter space to cover all possible signal realizations. However, drawing all possible realizations is computationally challenging, necessitating the sampling of a representative subset of this space. This study aims to identify optimal sampling techniques for the extensive dimensionality and volume of the 21cm signal parameter space. The optimally sampled training set will be used to train the ANN to infer from the global signal experiment. We investigate three sampling techniques: random, Latin hypercube (stratified), and Hammersley sequence (quasi-Monte Carlo) sampling, and compare their outcomes. Our findings reveal that sufficient samples must be drawn for robust and accurate ANN model training, regardless of the sampling technique employed. The required sample size depends primarily on two factors: the complexity of the data and the number of free parameters. More free parameters necessitate drawing more realizations. Among the sampling techniques utilized, we find that ANN models trained with Hammersley sequence sampling demonstrate greater robustness compared to those trained with Latin hypercube and Random sampling.

astro-ph.IM