SearcharxivSearch

arXiv subjects

Maximilian Autenrieth

Publications and source records attributed to Maximilian Autenrieth.

8 recordsLinked to original sources

On the origin of the environmental dependence of SN Ia magnitudes: A BayeSN view of the ZTF SN Ia DR2

Astrophysical variabilities of Type Ia supernovae (SNe Ia), such as their link with their birth environment, are now one of the leading sources of systematic uncertainties on the measurement of the dark energy equation-of-state parameter $w$. Population studies of SNe Ia, using large samples, give precious insights into these variabilities. We analyse a volume-limited subsample of 932 SNe from the ZTF SN Ia DR2 with BayeSN, a hierarchical Bayesian model for SN Ia SEDs. We investigate the distributions of SN Ia light curve parameters and their link with SN environment. Using a new training of BayeSN released in a companion paper, we find a smaller scatter of Hubble residuals compared to SALT. We then investigate the magnitude step, which accounts for the correlation between SN Ia standardised absolute magnitude and host environments. We find a posteriori steps of $0.103\pm0.010$ mag (a $10.1σ$ difference from 0) when using global stellar mass as an environmental proxy, and $0.086\pm0.010$ mag ($8.3σ$) when using local colour, in accordance with steps computed using SALT light curve fits. This confirms that the large step seen in the ZTF SN Ia DR2 data was not due to the SALT fit or the associated standardisation process. We then investigate the origin of the step, using a BayeSN model which accounts for both an intrinsic magnitude step and differing dust properties with the SN environment. We find a $0.103\pm0.018$ mag ($5.6σ$) step in global mass and a $0.085\pm0.019$ mag ($4.5σ$) step in local colour. The means of the $R_V$ distribution are similar between different host environments, with $Δ\mathbb{E}(R_V)\leq0.2$ across all environment proxies, with significances ranging from $0.6σ$ to $1.2σ$. This is a strong signal of the existence of an intrinsic dependence of SN Ia absolute magnitude on environment.

astro-ph.CO

FlowSN: Neural Simulation-Based Inference under Realistic Selection Effects applied to Supernova Cosmology

We present FlowSN, a statistical framework using simulation-based inference (SBI) with normalising flows to account for selection effects in observational astronomy. Failure to account for selection effects can lead to biased inference on global parameters. An example is Malmquist bias, where detection limits result in a sample skewed towards brighter objects. In Type Ia supernova (SN Ia) cosmology, these selection effects can systematically shift the inferred posterior distributions of cosmological parameters, necessitating the development of robust statistical frameworks to account for the biases. SBI enables us to implicitly learn probability distributions that are analytically intractable to calculate. In this work, we introduce a novel approach that employs a normalising flow to learn the non-analytic selected SN likelihood for a given survey from forward simulations, independent of the assumed cosmological model. The resulting likelihood approximation is incorporated into a hierarchical Bayesian framework and posterior sampling is performed using Hamiltonian Monte Carlo to obtain constraints on cosmological parameters conditioned on the observed data. The modular learnt likelihood approximation can be reused without retraining to evaluate different cosmological models, providing a key advantage over other SBI approaches. We demonstrate the performance of this methodology by training and testing the SBI technique using realistic LSST-like SNANA simulations for the first time. Our FlowSN approach yields accurate posterior estimates on cosmological parameters, including the dark energy equation of state $w_0$, that are an order of magnitude less biased than those obtained with conventional techniques and also exhibit improved frequentist calibration.

astro-ph.CO

Detecting Localized Density Anomalies in Multivariate Data via Coin-Flip Statistics

Detecting localized differences between two samples is a central task in scientific data analysis, required for the identification of signal events, regime changes, or model mismatch. We introduce EagleEye, a method that pinpoints local over- and under-densities in multivariate feature spaces. EagleEye assigns each point an anomaly score by encoding its ordered k-nearest-neighbour list as a binary membership sequence and testing whether the cumulative number of successes in this sequence is consistent with a binomial (coin-flipping) null model. In the presence of a genuine local anomaly, neighbours will preferentially belong to one of the two datasts, yielding an excess of ``successes'' relative to the binomial null model. These local, pointwise detections are consolidated into interpretable anomaly sets through a deterministic refinement procedure that can also estimate the irreducible background and local density anomaly purity. We demonstrate EagleEye's efficacy in three scenarios. We first consider an artificial data example with known localized over- and under-densities. Second, we demonstrate how EagleEye may be used for new physics searches at particle collider experiments in the presence of systematic background modelling differences. Finally, we conduct a climate analysis study that reveals localized changes in spatiotemporal temperature-pattern recurrence.

stat.ML

Using fractional derivatives to derive marginal densities

This paper presents a novel method for analytical derivations of marginal densities using the fractional derivatives of moment-generating functions. Although the method requires likelihood functions to take specific forms, its assumptions are otherwise modest. It only requires that the prior moment-generating functions exist, are finite, and are continuous and differentiable at certain points. We also present the probabilistic and statistical insights behind this method.

stat.ME

StratLearn-z: Improved photo-$z$ estimation from spectroscopic data subject to selection effects

A precise measurement of photometric redshifts (photo-z) is key for the success of modern photometric galaxy surveys. Machine learning (ML) methods show great promise in this context, but suffer from covariate shift (CS) in training sets due to selection bias where interesting sources are underrepresented, and the corresponding ML models show poor generalisation properties. We present an application of the StratLearn method to the estimation of photo-z, validating against simulations where we enforce the presence of CS to different degrees. StratLearn is a statistically principled approach that relies on splitting the source and target datasets into strata based on estimated propensity scores (i.e. the probability for an object to be in the source set given its observed covariates). After stratification, two conditional density estimators are fit separately to each stratum, then combined via a weighted average. We benchmark our results against the GPz algorithm, quantifying the performance of the two codes with a set of metrics. Our results show that the StratLearn-z metrics are only marginally affected by the presence of CS, while GPz shows a significant degradation of performance in the photo-z prediction for fainter objects. For the strongest CS scenario, StratLearn-z yields a reduced fraction of catastrophic errors, a factor of 2 improvement for the RMSE and one order of magnitude improvement on the bias. We also assess the quality of the conditional redshift estimates with the probability integral transform (PIT). The PIT distribution obtained from StratLearn-z features fat fewer outliers and is symmetric, i.e. the predictions appear to be centered around the true redshift value, despite showing a conservative estimation of the spread of the conditional redshift distributions. Our julia implementation of the method is available at https://github.com/chiaramoretti/StratLearn-z.

astro-ph.CO

Improved Weak Lensing Photometric Redshift Calibration via StratLearn and Hierarchical Modeling

Discrepancies between cosmological parameter estimates from cosmic shear surveys and from recent Planck cosmic microwave background measurements challenge the ability of the highly successful $Λ$CDM model to describe the nature of the Universe. To rule out systematic biases in cosmic shear survey analyses, accurate redshift calibration within tomographic bins is key. In this paper, we improve photo-$z$ calibration via Bayesian hierarchical modeling of full galaxy photo-$z$ conditional densities, by employing $\textit{StratLearn}$, a recently developed statistical methodology, which accounts for systematic differences in the distribution of the spectroscopic training/source set and the photometric target set. Using realistic simulations that were designed to resemble the KiDS+VIKING-450 dataset, we show that $\textit{StratLearn}$-estimated conditional densities improve the galaxy tomographic bin assignment, and that our $\textit{StratLearn}$-Bayesian framework leads to nearly unbiased estimates of the target population means. This leads to a factor of $\sim 2$ improvement upon the previously best photo-$z$ calibration method. Our approach delivers a maximum bias per tomographic bin of $Δ\langle z \rangle = 0.0095 \pm 0.0089$, with an average absolute bias of $0.0052 \pm 0.0067$ across the five tomographic bins.

astro-ph.CO

Stratified Learning: A General-Purpose Statistical Method for Improved Learning under Covariate Shift

We propose a simple, statistically principled, and theoretically justified method to improve supervised learning when the training set is not representative, a situation known as covariate shift. We build upon a well-established methodology in causal inference, and show that the effects of covariate shift can be reduced or eliminated by conditioning on propensity scores. In practice, this is achieved by fitting learners within strata constructed by partitioning the data based on the estimated propensity scores, leading to approximately balanced covariates and much-improved target prediction. We demonstrate the effectiveness of our general-purpose method on two contemporary research questions in cosmology, outperforming state-of-the-art importance weighting methods. We obtain the best reported AUC (0.958) on the updated "Supernovae photometric classification challenge", and we improve upon existing conditional density estimation of galaxy redshift from Sloan Data Sky Survey (SDSS) data.

stat.ML

Uncertainty Based Detection and Relabeling of Noisy Image Labels

Deep neural networks (DNNs) are powerful tools in computer vision tasks. However, in many realistic scenarios label noise is prevalent in the training images, and overfitting to these noisy labels can significantly harm the generalization performance of DNNs. We propose a novel technique to identify data with noisy labels based on the different distributions of the predictive uncertainties from a DNN over the clean and noisy data. Additionally, the behavior of the uncertainty over the course of training helps to identify the network weights which best can be used to relabel the noisy labels. Data with noisy labels can therefore be cleaned in an iterative process. Our proposed method can be easily implemented, and shows promising performance on the task of noisy label detection on CIFAR-10 and CIFAR-100.

cs.CV