SearcharxivSearch

arXiv subjects

Vahid Asadi

Publications and source records attributed to Vahid Asadi.

8 recordsLinked to original sources

Linear gate bounds against natural functions for position-verification

A quantum position-verification scheme attempts to verify the spatial location of a prover. The prover is issued a challenge with quantum and classical inputs and must respond with appropriate timings. We consider two well-studied position-verification schemes known as $f$-routing and $f$-BB84. Both schemes require an honest prover to locally compute a classical function $f$ of inputs of length $n$, and manipulate $O(1)$ size quantum systems. We prove the number of quantum gates plus single qubit measurements needed to implement a function $f$ is lower bounded linearly by the communication complexity of $f$ in the simultaneous message passing model with shared entanglement. Taking $f(x,y)=\sum_i x_i y_i$ to be the inner product function, we obtain a $Ω(n)$ lower bound on quantum gates plus single qubit measurements. The scheme is feasible for a prover with linear classical resources and $O(1)$ quantum resources, and secure against sub-linear quantum resources.

quant-ph

CatBoost versus Spectral Energy Distribution-Fitting: Estimating Galaxy Properties under Controlled Photometric Incompleteness

Estimating galaxy physical parameters from photometric data is fundamentally challenged by missing measurements that are endemic to astronomical surveys. Using a mock catalog from the Horizon-AGN hydrodynamical simulation that provides the true physical parameters, we evaluate CatBoost, a gradient-boosting algorithm that natively handles missing data, under a deliberately adversarial scenario: we train it on 12 photometric bands with increasing levels of injected missingness (10%, 20%, 30% missing per band) and compare its performance against an idealised parametric spectral energy distribution (SED)-fitting reference that uses complete 26-band photometry (no missing data, all bands available). This asymmetric design represents an upper-bound, best-case baseline for traditional methods. Despite this intentional disadvantage, CatBoost's performance degradation is limited: moving from complete data to 30% missing, mass RMSE increases from 0.08 to 0.18 dex, SFR RMSE from 0.41 to 0.53 dex, and redshift RMSE from 0.20 to 0.28, while bias remains near zero. Against the ideal SED-fitting reference, CatBoost trained on only 12 bands with 30% missing values achieves lower errors for mass (0.18 vs. 0.28 dex) and SFR (0.53 vs. 0.57 dex) and removes the systematic biases present in the SED-fitting results. For redshift, the SED-fitting reference has lower NMAD (0.030 vs. 0.093 at extreme missingness) while CatBoost maintains smaller bias. These results suggest that CatBoost's native handling of missing values can offer practical advantages for extracting galaxy properties from imperfect photometric surveys, at least under the conditions explored here.

astro-ph.GA

COSMOS2025: A Machine Learning Census of Massive Quiescent Galaxies at $2.5 < \text{z} < 5$

The existence of massive quiescent galaxies at high redshifts ($ \text{z} \gtrsim 2$) strongly constrains the rapid quenching mechanisms in galaxy evolution models. We present a machine learning framework to identify massive ($\log(\text{M}_*/\text{M}_\odot) > 9.5$) quiescent galaxies at $2.5 < \text{z} < 5$ in the COSMOS2025 catalog. We train a \texttt{CatBoostClassifier} on mock photometry from the Santa Cruz semi-analytic models (SAMs), incorporating key JWST NIRCam bands and realistic noise to transfer the SAM-derived quiescent label (based on specific star-formation rate) to the observational space. When validated against the SAM ground truth, our classifier achieves a significantly higher recall (completeness) of 78\% (compared to 53\% for spectral energy distribution (SED)-fitting), while maintaining a high purity of 82\%. Applied to the COSMOS2025 sample, and assuming the SAM definition of quiescence transfers to the real Universe, the model identifies 1111 quiescent candidates, a population 2.6 times larger than the 427 candidates identified via the catalog's simple SED-fitting configuration. Under the SAM definition of quiescence, this consistent pattern of high purity but poor completeness suggests that the SED-fitting methods, constrained by simplified parametric star-formation histories, may miss a significant fraction of the quiescent population, likely galaxies in crucial transitional evolutionary stages. The trained classifier and classified COSMOS2025 sample are publicly available.

astro-ph.GA

COSMOS2025: Machine Learning Classification of Early- and Late-type Galaxies at 0 < z < 3

We present a fast, interpretable machine learning framework to classify early- and late-type galaxies in the COSMOS2025 catalog at $0 < z < 3$, without relying on image-based training labels or computationally expensive structural fitting. Using the Santa Cruz Semi-Analytic Model, we generate a training set with secure morphological labels defined by bulge-to-total mass ratio and specific star formation rate. We bridge the simulation-to-observation domain gap by injecting realistic photometric noise derived from COSMOS2025. A CatBoostClassifier trained on 66 broadband colors achieves excellent performance in the simulated domain, recovering late-types with 98\% precision/recall and early-types with 91\% precision and 88\% recall. Applied to 44,132 COSMOS2025 galaxies, the model reveals a striking bimodality: only about 6\% of galaxies receive intermediate probabilities ($0.3 < P(\text{Early type}) < 0.7$) -- nearly identical to the fraction observed in the simulation. This demonstrates that broadband colors are a decisive morphological discriminant, with the remaining 94\% classified at high confidence. Validation against independent bulge+disk decompositions yields 70\% overall accuracy, with late-types identified at 78\% purity and 74\% completeness. The most important color feature, F277W-F444W, reflects the expected optical/NIR contrast between old and young stellar populations. The full pipeline completes in under 30 minutes on standard hardware, demonstrating that simulation-trained color-based classifiers offer a scalable, physically interpretable route to approximate morphology for large next-generation surveys.

astro-ph.GA

Machine Learning vs. Spectral Energy Distribution Fitting: A Comparative Analysis of Accuracy in Stellar Mass Estimation

Traditional spectral energy distribution (SED)-fitting methods for stellar mass estimation face persistent challenges including systematic biases and computational constraints. We present a controlled comparison of machine learning (ML) and SED-fitting methods, assessing their accuracy, robustness, and computational efficiency. Using a sample of COSMOS-like galaxies from the Horizon-AGN simulation as a benchmark with known true masses, we evaluate the Parametric t-SNE (Pt-SNE) algorithm -- trained on noise-injected BC03 models -- against the established SED-fitting code LePhare. Our results demonstrate that Pt-SNE achieves superior accuracy, with a root-mean-square error (sigma_F) of 0.169 dex compared to LePhare's 0.306 dex. Crucially, Pt-SNE exhibits significantly lower bias (0.029 dex) compared to LePhare (0.286 dex). Pt-SNE also shows greater robustness across all stellar mass ranges, particularly for low-mass galaxies (10^9 to 10^10 solar masses), where it reduces errors by 47-53 %. Even when restricted to only six optical bands, Pt-SNE outperforms LePhare using all 26 available photometric bands, underscoring its superior informational efficiency. Computationally, Pt-SNE processes large datasets approximately 3.2 x 10^3 times faster than LePhare. These findings highlight the fundamental advantages of ML methods for stellar mass estimation, demonstrating their potential to deliver more accurate, stable, and scalable measurements for large-scale galaxy surveys.

astro-ph.GA

Machine Learning Classification of COSMOS2020 Galaxies: Quiescent vs. Star-Forming

Accurately distinguishing between quiescent and star-forming galaxies is essential for understanding galaxy evolution. Traditional methods, such as spectral energy distribution (SED) fitting, can be computationally expensive and may struggle to capture complex galaxy properties. This study aims to develop a robust and efficient machine learning (ML) classification method to identify quiescent and star-forming galaxies within the Farmer COSMOS2020 catalog. We utilized JWST wide-field light cones from the Santa Cruz semi-analytical modeling framework to train a supervised ML model, the CatBoostClassifier, using 28 color features derived from 8 mutual photometric bands within the COSMOS catalog. The model was validated against a testing set and compared to the SED-fitting method in terms of precision, recall, F1-score, and execution time. Preprocessing steps included addressing missing data, injecting observational noise, and applying a magnitude cut (ch1 < 26 AB) along with a redshift range of 0.2 < z < 3.5 to align the simulated and observational datasets. The ML method achieved an F1-score of 89\% for quiescent galaxies, significantly outperforming the SED-fitting method, which achieved 54%. The ML model demonstrated superior recall (88% vs. 38%) while maintaining comparable precision. When applied to the COSMOS2020 catalog, the ML model predicted a systematically higher fraction of quiescent galaxies across all redshift bins within 0.2 < z < 3.5 compared to traditional methods like NUVrJ and SED-fitting. This study shows that ML, combined with multi-wavelength data, can effectively identify quiescent and star-forming galaxies, providing valuable insights into galaxy evolution. The trained classifier and full classification catalog are publicly available.

astro-ph.GA

Semi-supervised classification of stars, galaxies and quasars using K-means and random-forest approaches

Classifying stars, galaxies, and quasars is essential for understanding cosmic structure and evolution; however, the vast data from modern surveys make manual classification impractical, while supervised learning methods remain constrained by the scarcity of labeled spectroscopic data. We aim to develop a scalable, label-efficient method for astronomical classification by leveraging semi-supervised learning (SSL) to overcome the limitations of fully supervised approaches. We propose a novel SSL framework combining K-means clustering with random forest classification. Our method partitions unlabeled data into 50 clusters, propagates labels from spectroscopically confirmed centroids to 95% of cluster members, and trains a random forest on the expanded pseudo-labeled dataset. We applied this to the CPz catalog, containing multi-survey photometric and spectroscopic data, and compared performance with a fully supervised random forest. Our SSL approach achieves F1 scores of 98.8%, 98.9%, and 92.0% for stars, galaxies, and quasars, respectively, closely matching the supervised method with F1 scores of 99.1%, 99.1%, and 93.1%, while outperforming traditional color-cut techniques. The method demonstrates robustness in high-dimensional feature spaces and superior label efficiency compared to prior work. This work highlights SSL as a scalable solution for astronomical classification when labeled data is limited, though performance may be degraded in lower dimensional settings.

astro-ph.GA

Leveraging Machine Learning for Accurate and Fast Stellar Mass Estimation of Galaxies

Unveiling the evolutionary history of galaxies necessitates a precise understanding of their physical properties. Traditionally, astronomers achieve this through spectral energy distribution (SED) fitting. However, this approach can be computationally intensive and time-consuming, particularly for large datasets. This study investigates the viability of machine learning (ML) algorithms as an alternative to traditional SED-fitting for estimating stellar masses in galaxies. We compare a diverse range of unsupervised and supervised learning approaches including prominent algorithms such as K-means, HDBSCAN, Parametric t-Distributed Stochastic Neighbor Embedding (Pt-SNE), Principal Component Analysis (PCA), Random Forest, and Self-Organizing Maps (SOM) against the well-established LePhare code, which performs SED-fitting as a benchmark. We train various ML algorithms using simple model SEDs in photometric space, generated with the BC03 code. These trained algorithms are then employed to estimate the stellar masses of galaxies within a subset of the COSMOS survey dataset. The performance of these ML methods is subsequently evaluated and compared with the results obtained from LePhare, focusing on both accuracy and execution time. Our evaluation reveals that ML algorithms can achieve comparable accuracy to LePhare while offering significant speed advantages (1,000 to 100,000 times faster). K-means and HDBSCAN emerge as top performers among our selected ML algorithms. Supervised learning algorithms like Random Forest and manifold learning techniques such as Pt-SNE and SOM also show promising results. These findings suggest that ML algorithms hold significant promise as a viable alternative to traditional SED-fitting methods for estimating the stellar masses of galaxies.

astro-ph.GA