Searcharxiv⌕ Search

arXiv subjects

A. I. Malz

Publications and source records attributed to A. I. Malz.

13 recordsLinked to original sources

VAR-PZnn: A machine-learning framework for AGN photometric redshifts using color and variability-based features

Photometric redshift estimation for active galactic nuclei (AGNs) remains a fundamental challenge for current and upcoming large-scale photometric surveys. Traditional spectral energy distribution (SED) fitting suffers from color-redshift degeneracies, particularly for AGNs whose power-law continua hide the strong spectral features required to anchor redshift estimates. While AGN variability provides additional constraining power, existing frameworks require multi-band light curves that are not always available. This work presents VAR-PZnn, a fully connected mixture density network that integrates 26 variability features extracted from ZTF g-band light curves with optical photometry from Pan-STARRS1, mid-infrared (MIR) photometry from CatWISE, and, for a subsample, NIR photometry from UKIDSS. The model is trained and tested on 72,728 spectroscopically confirmed AGNs/QSOs spanning 0.01 < z < 4.5 and g-band magnitudes from 17 to 21.5. For the main sample, we achieve σ_{NMAD} = 0.058 and an outlier fraction of η= 8.2%, which reduces to 5.4% when the 10% of sources with the highest predicted uncertainty are excluded. An ablation study demonstrates that MIR photometry provides the dominant constraint for photo-z accuracy, while variability features serve as a secondary refiner. Using UKIDSS NIR data as a proxy for future synergies between LSST and space-based missions like Euclid and Roman, we obtain η= 13.3% without MIR data and η= 4.6% when MIR is available. We benchmark against Low-Resolution Templates (LRT) SED fitting (η= 28.7%) and the VAR-PZ framework; applying single-band VAR-PZ priors worsens LRT performance to η= 39.4% due to single-band light-curve degeneracies, confirmed via simulations (η= 27.6% to 28.1%). This framework provides a scalable approach for the Legacy Survey of Space and Time (LSST).

astro-ph.GA↗

Toward decision-aware AI for LSST-scale time-domain astronomy

The Vera C. Rubin Observatory's Legacy Survey of Space and Time (LSST) will generate approximately (10^7) alerts per night, pushing time-domain astronomy beyond pipelines that treat discovery as a static labeling problem. We argue that LSST is better understood as a partially observed dynamical environment, in which scientific return depends on the quality of follow-up decisions made under uncertainty and finite observational resources. The central challenge is therefore to maintain evolving, uncertainty-aware representations of astrophysical sources and to select actions that maximize long-term scientific value. We propose that foundation models trained on heterogeneous time-domain data can learn survey-scale representations of source state, while decision-theoretic policies support principled, auditable allocation of follow-up resources. Embedded within human-supervised agentic systems, these components position AI as part of the operational inference loop rather than as a downstream predictive tool. The way such systems represent belief, optimize utility, and expose their reasoning will shape observational efficiency, the distribution of scientific agency, including who participates in discovery and the scientific questions that receive priority.

astro-ph.IM↗

photoD with Rubin's Data Preview 1: first stellar photometric distances and deficit of faint blue stars. Stellar distances with Rubin's DP1

Aims: We investigate the utility of Rubin's Data Preview 1 for estimating stellar number density profile in the Milky Way halo. Methods: Stellar broad-band near-UV to near-IR $ugrizy$ photometry released in Rubin's Data Preview 1 is used to estimate distance and metallicity for blue main sequence stars brighter than $r=24$ in three $\sim$1.1. sq.~deg. fields at southern Galactic latitudes. Results: Compared to TRILEGAL simulations of the Galaxy's stellar content by (Dal Tio, 2022), we find a significant deficit of blue main sequence turn-off stars with $22 < r < 24$. We interpret this discrepancy as a signature of a much steeper halo number density profile at galactocentric distances $10-50$ kpc than the cannonical $\sim1/r^3$ profile assumed in TRILEGAL simulations. Conclusions: This interpretation is consistent with earlier suggestions based on observations of more luminous, but much less numerous, evolved stellar populations, and a few pencil beam surveys of blue main sequence stars in the northern sky. These results bode well for the future Galactic halo exploration with Rubin's Legacy Survey of Space and Time.

astro-ph.GA↗

CLMM: a LSST-DESC Cluster weak Lensing Mass Modeling library for cosmology

We present the v1.0 release of CLMM, an open source Python library for the estimation of the weak lensing masses of clusters of galaxies. CLMM is designed as a standalone toolkit of building blocks to enable end-to-end analysis pipeline validation for upcoming cluster cosmology analyses such as the ones that will be performed by the LSST-DESC. Its purpose is to serve as a flexible, easy-to-install and easy-to-use interface for both weak lensing simulators and observers and can be applied to real and mock data to study the systematics affecting weak lensing mass reconstruction. At the core of CLMM are routines to model the weak lensing shear signal given the underlying mass distribution of galaxy clusters and a set of data operations to prepare the corresponding data vectors. The theoretical predictions rely on existing software, used as backends in the code, that have been thoroughly tested and cross-checked. Combined, theoretical predictions and data can be used to constrain the mass distribution of galaxy clusters as demonstrated in a suite of example Jupyter Notebooks shipped with the software and also available in the extensive online documentation.

astro-ph.CO↗

Evaluation of probabilistic photometric redshift estimation approaches for The Rubin Observatory Legacy Survey of Space and Time (LSST)

Many scientific investigations of photometric galaxy surveys require redshift estimates, whose uncertainty properties are best encapsulated by photometric redshift (photo-z) posterior probability density functions (PDFs). A plethora of photo-z PDF estimation methodologies abound, producing discrepant results with no consensus on a preferred approach. We present the results of a comprehensive experiment comparing twelve photo-z algorithms applied to mock data produced for The Rubin Observatory Legacy Survey of Space and Time (LSST) Dark Energy Science Collaboration (DESC). By supplying perfect prior information, in the form of the complete template library and a representative training set as inputs to each code, we demonstrate the impact of the assumptions underlying each technique on the output photo-z PDFs. In the absence of a notion of true, unbiased photo-z PDFs, we evaluate and interpret multiple metrics of the ensemble properties of the derived photo-z PDFs as well as traditional reductions to photo-z point estimates. We report systematic biases and overall over/under-breadth of the photo-z PDFs of many popular codes, which may indicate avenues for improvement in the algorithms or implementations. Furthermore, we raise attention to the limitations of established metrics for assessing photo-z PDF accuracy; though we identify the conditional density estimate (CDE) loss as a promising metric of photo-z PDF performance in the case where true redshifts are available but true photo-z PDFs are not, we emphasize the need for science-specific performancemetrics.

astro-ph.CO↗

Approximating photo-$z$ PDFs for large surveys

Modern galaxy surveys produce redshift probability density functions (PDFs) in addition to traditional photometric redshift (photo-$z$) point estimates. However, the storage of photo-$z$ PDFs may present a challenge with increasingly large catalogs, as we face a trade-off between the accuracy of subsequent science measurements and the limitation of finite storage resources. This paper presents $\texttt{qp}$, a Python package for manipulating parametrizations of 1-dimensional PDFs, as suitable for photo-$z$ PDF compression. We use $\texttt{qp}$ to investigate the performance of three simple PDF storage formats (quantiles, samples, and step functions) as a function of the number of stored parameters on two realistic mock datasets, representative of upcoming surveys with different data qualities. We propose some best practices for choosing a photo-$z$ PDF approximation scheme and demonstrate the approach on a science case using performance metrics on both ensembles of individual photo-$z$ PDFs and an estimator of the overall redshift distribution function. We show that both the properties of the set of PDFs we wish to approximate and the chosen fidelity metric(s) affect the optimal parametrization. Additionally, we find that quantiles and samples outperform step functions, and we encourage further consideration of these formats for PDF approximation.

astro-ph.IM↗

The Photometric LSST Astronomical Time-series Classification Challenge (PLAsTiCC): Selection of a performance metric for classification probabilities balancing diverse science goals

Classification of transient and variable light curves is an essential step in using astronomical observations to develop an understanding of their underlying physical processes. However, upcoming deep photometric surveys, including the Large Synoptic Survey Telescope (LSST), will produce a deluge of low signal-to-noise data for which traditional labeling procedures are inappropriate. Probabilistic classification is more appropriate for the data but are incompatible with the traditional metrics used on deterministic classifications. Furthermore, large survey collaborations intend to use these classification probabilities for diverse science objectives, indicating a need for a metric that balances a variety of goals. We describe the process used to develop an optimal performance metric for an open classification challenge that seeks probabilistic classifications and must serve many scientific interests. The Photometric LSST Astronomical Time-series Classification Challenge (PLAsTiCC) is an open competition aiming to identify promising techniques for obtaining classification probabilities of transient and variable objects by engaging a broader community both within and outside astronomy. Using mock classification probability submissions emulating archetypes of those anticipated of PLAsTiCC, we compare the sensitivity of metrics of classification probabilities under various weighting schemes, finding that they yield qualitatively consistent results. We choose as a metric for PLAsTiCC a weighted modification of the cross-entropy because it can be meaningfully interpreted. Finally, we propose extensions of our methodology to ever more complex challenge goals and suggest some guiding principles for approaching the choice of a metric of probabilistic classifications.

astro-ph.IM↗

Results of the Photometric LSST Astronomical Time-series Classification Challenge (PLAsTiCC)

Next-generation surveys like the Legacy Survey of Space and Time (LSST) on the Vera C. Rubin Observatory will generate orders of magnitude more discoveries of transients and variable stars than previous surveys. To prepare for this data deluge, we developed the Photometric LSST Astronomical Time-series Classification Challenge (PLAsTiCC), a competition which aimed to catalyze the development of robust classifiers under LSST-like conditions of a non-representative training set for a large photometric test set of imbalanced classes. Over 1,000 teams participated in PLAsTiCC, which was hosted in the Kaggle data science competition platform between Sep 28, 2018 and Dec 17, 2018, ultimately identifying three winners in February 2019. Participants produced classifiers employing a diverse set of machine learning techniques including hybrid combinations and ensemble averages of a range of approaches, among them boosted decision trees, neural networks, and multi-layer perceptrons. The strong performance of the top three classifiers on Type Ia supernovae and kilonovae represent a major improvement over the current state-of-the-art within astronomy. This paper summarizes the most promising methods and evaluates their results in detail, highlighting future directions both for classifier development and simulation needs for a next generation PLAsTiCC data set.

astro-ph.IM↗

Gaia DR2 unravels incompleteness of nearby cluster population: New open clusters in the direction of Perseus

Open clusters (OCs) are popular tracers of the structure and evolutionary history of the Galactic disk. The OC population is often considered to be complete within 1.8 kpc of the Sun. The recent Gaia Data Release 2 (DR2) allows the latter claim to be challenged. We perform a systematic search for new OCs in the direction of Perseus using precise and accurate astrometry from Gaia DR2. We implement a coarse-to-fine search method. First, we exploit spatial proximity using a fast density-aware partitioning of the sky via a k-d tree in the spatial domain of Galactic coordinates, (l, b). Secondly, we employ a Gaussian mixture model in the proper motion space to quickly tag fields around OC candidates. Thirdly, we apply an unsupervised membership assignment method, UPMASK, to scrutinise the candidates. We visually inspect colour-magnitude diagrams to validate the detected objects. Finally, we perform a diagnostic to quantify the significance of each identified overdensity in proper motion and in parallax space We report the discovery of 41 new stellar clusters. This represents an increment of at least 20% of the previously known OC population in this volume of the Milky Way. We also report on the clear identification of NGC 886, an object previously considered an asterism. This letter challenges the previous claim of a near-complete sample of open clusters up to 1.8 kpc. Our results reveal that this claim requires revision, and a complete census of nearby open clusters is yet to be found.

astro-ph.GA↗

Bayesian Redshift Classification of Emission-line Galaxies with Photometric Equivalent Widths

We present a Bayesian approach to the redshift classification of emission-line galaxies when only a single emission line is detected spectroscopically. We consider the case of surveys for high-redshift Lyman-alpha-emitting galaxies (LAEs), which have traditionally been classified via an inferred rest-frame equivalent width (EW) greater than 20 angstrom. Our Bayesian method relies on known prior probabilities in measured emission-line luminosity functions and equivalent width distributions for the galaxy populations, and returns the probability that an object in question is an LAE given the characteristics observed. This approach will be directly relevant for the Hobby-Eberly Telescope Dark Energy Experiment (HETDEX), which seeks to classify ~10^6 emission-line galaxies into LAEs and low-redshift [O II] emitters. For a simulated HETDEX catalog with realistic measurement noise, our Bayesian method recovers 86% of LAEs missed by the traditional EW > 20 angstrom cutoff over 2 < z < 3, outperforming the EW cut in both contamination and incompleteness. This is due to the method's ability to trade off between the two types of binary classification error by adjusting the stringency of the probability requirement for classifying an observed object as an LAE. In our simulations of HETDEX, this method reduces the uncertainty in cosmological distance measurements by 14% with respect to the EW cut, equivalent to recovering 29% more cosmological information. Rather than using binary object labels, this method enables the use of classification probabilities in large-scale structure analyses. It can be applied to narrowband emission-line surveys as well as upcoming large spectroscopic surveys including Euclid and WFIRST.

astro-ph.IM↗

Physical and Morphological Properties of [O II] Emitting Galaxies in the HETDEX Pilot Survey

The Hobby-Eberly Dark Energy Experiment pilot survey identified 284 [O II] 3727 emitting galaxies in a 169 square-arcminute field of sky in the redshift range 0 < z < 0.57. This line flux limited sample provides a bridge between studies in the local universe and higher-redshift [O II] surveys. We present an analysis of the star formation rates (SFRs) of these galaxies as a function of stellar mass as determined via spectral energy distribution fitting. The [O II] emitters fall on the "main sequence" of star-forming galaxies with SFR decreasing at lower masses and redshifts. However, the slope of our relation is flatter than that found for most other samples, a result of the metallicity dependence of the [O II] star formation rate indicator. The mass specific SFR is higher for lower mass objects, supporting the idea that massive galaxies formed more quickly and efficiently than their lower mass counterparts. This is confirmed by the fact that the equivalent widths of the [O II] emission lines trend smaller with larger stellar mass. Examination of the morphologies of the [O II] emitters reveals that their star formation is not a result of mergers, and the galaxies' half-light radii do not indicate evolution of physical sizes.

astro-ph.GA↗

HST Emission Line Galaxies at z ~ 2: The Ly-alpha Escape Fraction

We compare the H-beta line strengths of 1.90 < z < 2.35 star-forming galaxies observed with the near-IR grism of the Hubble Space Telescope with ground-based measurements of Ly-alpha from the HETDEX Pilot Survey and narrow-band imaging. By examining the line ratios of 73 galaxies, we show that most star-forming systems at this epoch have a Ly-alpha escape fraction below ~6%. We confirm this result by using stellar reddening to estimate the effective logarithmic extinction of the H-beta emission line (c_Hbeta = 0.5) and measuring both the H-beta and Ly-alpha luminosity functions in a ~ 100,000 cubic Mpc volume of space. We show that in our redshift window, the volumetric Ly-alpha escape fraction is at most 4.4+/-2.1(1.2)%, with an additional systematic ~25% uncertainty associated with our estimate of extinction. Finally, we demonstrate that the bulk of the epoch's star-forming galaxies have Ly-alpha emission line optical depths that are significantly greater than that for the underlying UV continuum. In our predominantly [O~III] 5007-selected sample of galaxies, resonant scattering must be important for the escape of Ly-alpha photons.

astro-ph.GA↗

Spectral Energy Distribution Fitting of HETDEX Pilot Survey Lyman-alpha Emitters in COSMOS and GOODS-N

We use broadband photometry extending from the rest-frame UV to the near-IR to fit the individual spectral energy distributions (SEDs) of 63 bright (L(Ly-alpha) > 10^43 ergs/s) Ly-alpha emitting galaxies (LAEs) in the redshift range 1.9 < z < 3.6. We find that these LAEs are quite heterogeneous, with stellar masses that span over three orders of magnitude, from 7.5 < log M < 10.5. Moreover, although most LAEs have small amounts of extinction, some high-mass objects have stellar reddenings as large as E(B-V) ~0.4. Interestingly, in dusty objects the optical depths for Ly-alpha and the UV continuum are always similar, indicating that Ly-alpha photons are not undergoing many scatters before escaping their galaxy. In contrast, the ratio of optical depths in low-reddening systems can vary widely, illustrating the diverse nature of the systems. Finally, we show that in the star formation rate (SFR)-log mass diagram, our LAEs fall above the "main-sequence" defined by z ~ 3 continuum selected star-forming galaxies. In this respect, they are similar to sub-mm-selected galaxies, although most LAEs have much lower mass.

astro-ph.GA↗