SearcharxivSearch

arXiv subjects

Alexander Henzi

Publications and source records attributed to Alexander Henzi.

16 recordsLinked to original sources

Survival Isotonic Distributional Regression

We introduce Survival-IDR (S-IDR), a nonparametric estimator of conditional survival distributions under order restrictions, extending Isotonic Distributional Regression (IDR; Henzi et al., 2021) to right-censored outcomes. S-IDR has no tuning parameters and accommodates continuous, discrete, and partially ordered covariates. We first study the direct Kaplan-Meier adaptation of IDR: it is uniformly consistent at the minimax rate, but only when the conditional outcomes are hazard-rate ordered. We trace this restriction to the Kaplan-Meier estimator's failure to satisfy the Cauchy mean value property on non-i.i.d. samples, and use the diagnosis to construct S-IDR. The S-IDR estimator is uniformly consistent under only stochastic dominance of the conditional outcomes, attains the minimax rate when the smoothness of the conditional CDFs is known, and admits a known cross-threshold PAVA acceleration. We further embed S-IDR in a distributional single-index framework on a benchmark suite, and apply it in a case study that validates the MELD score used for liver-transplant wait list management. Accompanying R, Python and Rust packages are available at https://github.com/AlexanderHenzi/isodistrreg.

math.ST

In-sample calibration yields conformal calibration guarantees

Conformal prediction produces a set of predictions that has out-of-sample calibration guarantees by construction, under the assumption of exchangeability. In this work, we study the calibration properties of conformal predictive systems, which issue sets of predictive distributions for real-valued outcomes. We demonstrate that conformal predictive systems implicitly exploit prediction methods that are in-sample calibrated to construct sets that are guaranteed to contain a calibrated predictive distribution out-of-sample. This allows us to take any prediction method that is in-sample calibrated, and conformalize it to obtain a predictive system with out-of-sample calibration guarantees. While the satisfied notion of calibration is typically that prediction intervals derived from the predictive distribution have the correct marginal coverage, we show that this line of reasoning can be extended to stronger conditional notions of calibration that are common in statistical forecasting theory. Using this, we introduce two predictive systems that satisfy stronger out-of-sample calibration guarantees than existing conformal predictive systems. The first method corresponds to a binning of the data, while the second leverages isotonic distributional regression (IDR), a non-parametric distributional regression method under order constraints. We study the theoretical properties of these new predictive systems, and compare their performance in a simulation experiment. They are then applied to two case studies on European temperature forecasts and on predictions for the length of patient stay in Swiss intensive care units. Both approaches are found to outperform existing conformal predictive systems, while conformal IDR additionally provides a natural method for quantifying epistemic uncertainty of the predictions.

stat.ME

Estimation and convergence rates in the distributional single index model

The distributional single index model is a semiparametric regression model in which the conditional distribution functions $P(Y \leq y | X = x) = F_0(\theta_0(x), y)$ of a real-valued outcome variable $Y$ depend on $d$-dimensional covariates $X$ through a univariate, parametric index function $\theta_0(x)$, and increase stochastically as $\theta_0(x)$ increases. We propose least squares approaches for the joint estimation of $\theta_0$ and $F_0$ in the important case where $\theta_0(x) = \alpha_0^{\top}x$ and obtain convergence rates of $n^{-1/3}$, thereby improving an existing result that gives a rate of $n^{-1/6}$. A simulation study indicates that the convergence rate for the estimation of $\alpha_0$ might be faster. Furthermore, we illustrate our methods in a real data application that demonstrates the advantages of shape restrictions in single index models.

math.ST

Invariant Probabilistic Prediction

In recent years, there has been a growing interest in statistical methods that exhibit robust performance under distribution changes between training and test data. While most of the related research focuses on point predictions with the squared error loss, this article turns the focus towards probabilistic predictions, which aim to comprehensively quantify the uncertainty of an outcome variable given covariates. Within a causality-inspired framework, we investigate the invariance and robustness of probabilistic predictions with respect to proper scoring rules. We show that arbitrary distribution shifts do not, in general, admit invariant and robust probabilistic predictions, in contrast to the setting of point prediction. We illustrate how to choose evaluation metrics and restrict the class of distribution shifts to allow for identifiability and invariance in the prototypical Gaussian heteroscedastic linear model. Motivated by these findings, we propose a method to yield invariant probabilistic predictions, called IPP, and study the consistency of the underlying parameters. Finally, we demonstrate the empirical performance of our proposed procedure on simulated as well as on single-cell data.

stat.ME

Various New Inequalities for Beta Distributions

This note provides some new inequalities and approximations for beta distributions, including tail inequalities, exponential inequalities of Hoeffding and Bernstein type, Gaussian inequalities and approximations.

math.ST

Easy Uncertainty Quantification (EasyUQ): Generating Predictive Distributions from Single-valued Model Output

How can we quantify uncertainty if our favorite computational tool - be it a numerical, a statistical, or a machine learning approach, or just any computer model - provides single-valued output only? In this article, we introduce the Easy Uncertainty Quantification (EasyUQ) technique, which transforms real-valued model output into calibrated statistical distributions, based solely on training data of model output-outcome pairs, without any need to access model input. In its basic form, EasyUQ is a special case of the recently introduced Isotonic Distributional Regression (IDR) technique that leverages the pool-adjacent-violators algorithm for nonparametric isotonic regression. EasyUQ yields discrete predictive distributions that are calibrated and optimal in finite samples, subject to stochastic monotonicity. The workflow is fully automated, without any need for tuning. The Smooth EasyUQ approach supplements IDR with kernel smoothing, to yield continuous predictive distributions that preserve key properties of the basic form, including both, stochastic monotonicity with respect to the original model output, and asymptotic consistency. For the selection of kernel parameters, we introduce multiple one-fit grid search, a computationally much less demanding approximation to leave-one-out cross-validation. We use simulation examples and forecast data from weather prediction to illustrate the techniques. In a study of benchmark problems from machine learning, we show how EasyUQ and Smooth EasyUQ can be integrated into the workflow of neural network learning and hyperparameter tuning, and find EasyUQ to be competitive with conformal prediction, as well as more elaborate input-based approaches.

stat.ME

A Rank-Based Sequential Test of Independence

We consider the problem of independence testing for two univariate random variables in a sequential setting. By leveraging recent developments on safe, anytime-valid inference, we propose a test with time-uniform type I error control and derive explicit bounds on the finite sample performance of the test. We demonstrate the empirical performance of the procedure in comparison to existing sequential and non-sequential independence tests. Furthermore, since the proposed test is distribution free under the null hypothesis, we empirically simulate the gap due to Ville's inequality, the supermartingale analogue of Markov's inequality, that is commonly applied to control type I error in anytime-valid inference, and apply this to construct a truncated sequential test.

stat.ME

A safe Hosmer-Lemeshow test

This article proposes an alternative to the Hosmer-Lemeshow (HL) test for evaluating the calibration of probability forecasts for binary events. The approach is based on e-values, a new tool for hypothesis testing. An e-value is a random variable with expected value less or equal to one under a null hypothesis. Large e-values give evidence against the null hypothesis, and the multiplicative inverse of an e-value is a p-value. Our test uses online isotonic regression to estimate the calibration curve as a `betting strategy' against the null hypothesis. We show that the test has power against essentially all alternatives, which makes it theoretically superior to the HL test and at the same time resolves the well-known instability problem of the latter. A simulation study shows that a feasible version of the proposed eHL test can detect slight miscalibrations in practically relevant sample sizes, but trades its universal validity and power guarantees against a reduced empirical power compared to the HL test in a classical simulation setup.We illustrate our test on recalibrated predictions for credit card defaults during the Taiwan credit card crisis, where the classical HL test delivers equivocal results.

stat.ME

Anytime Valid Tests of Conditional Independence Under Model-X

We propose a sequential, anytime-valid method to test the conditional independence of a response $Y$ and a predictor $X$ given a random vector $Z$. The proposed test is based on e-statistics and test martingales, which generalize likelihood ratios and allow valid inference at arbitrary stopping times. In accordance with the recently introduced model-X setting, our test depends on the availability of the conditional distribution of $X$ given $Z$, or at least a sufficiently sharp approximation thereof. Within this setting, we derive a general method for constructing e-statistics for testing conditional independence, show that it leads to growth-rate optimal e-statistics for simple alternatives, and prove that our method yields tests with asymptotic power one in the special case of a logistic regression model. A simulation study is done to demonstrate that the approach is competitive in terms of power when compared to established sequential and nonsequential testing methods, and robust with respect to violations of the model-X assumption.

stat.ME

Honest calibration assessment for binary outcome predictions

Probability predictions from binary regressions or machine learning methods ought to be calibrated: If an event is predicted to occur with probability $x$, it should materialize with approximately that frequency, which means that the so-called calibration curve $p(\cdot)$ should equal the identity, $p(x) = x$ for all $x$ in the unit interval. We propose honest calibration assessment based on novel confidence bands for the calibration curve, which are valid only subject to the natural assumption of isotonicity. Besides testing the classical goodness-of-fit null hypothesis of perfect calibration, our bands facilitate inverted goodness-of-fit tests whose rejection allows for the sought-after conclusion of a sufficiently well specified model. We show that our bands have a finite sample coverage guarantee, are narrower than existing approaches, and adapt to the local smoothness of the calibration curve $p$ and the local variance of the binary observations. In an application to model predictions of an infant having a low birth weight, the bounds give informative insights on model calibration.

math.ST

Distributional (Single) Index Models

A Distributional (Single) Index Model (DIM) is a semi-parametric model for distributional regression, that is, estimation of conditional distributions given covariates. The method is a combination of classical single index models for the estimation of the conditional mean of a response given covariates, and isotonic distributional regression. The model for the index is parametric, whereas the conditional distributions are estimated non-parametrically under a stochastic ordering constraint. We show consistency of our estimators and apply them to a highly challenging data set on the length of stay (LoS) of patients in intensive care units. We use the model to provide skillful and calibrated probabilistic predictions for the LoS of individual patients, that outperform the available methods in the literature.

stat.ME

Consistent estimation of distribution functions under increasing concave and convex stochastic ordering

A random variable $Y_1$ is said to be smaller than $Y_2$ in the increasing concave stochastic order if $\mathbb{E}[ϕ(Y_1)] \leq \mathbb{E}[ϕ(Y_2)]$ for all increasing concave functions $ϕ$ for which the expected values exist, and smaller than $Y_2$ in the increasing convex order if $\mathbb{E}[ψ(Y_1)] \leq \mathbb{E}[ψ(Y_2)]$ for all increasing convex $ψ$. This article develops nonparametric estimators for the conditional cumulative distribution functions $F_x(y) = \mathbb{P}(Y \leq y \mid X = x)$ of a response variable $Y$ given a covariate $X$, solely under the assumption that the conditional distributions are increasing in $x$ in the increasing concave or increasing convex order. Uniform consistency and rates of convergence are established both for the $K$-sample case $X \in \{1, \dots, K\}$ and for continuously distributed $X$.

math.ST

Valid sequential inference on probability forecast performance

Probability forecasts for binary events play a central role in many applications. Their quality is commonly assessed with proper scoring rules, which assign forecasts a numerical score such that a correct forecast achieves a minimal expected score. In this paper, we construct e-values for testing the statistical significance of score differences of competing forecasts in sequential settings. E-values have been proposed as an alternative to p-values for hypothesis testing, and they can easily be transformed into conservative p-values by taking the multiplicative inverse. The e-values proposed in this article are valid in finite samples without any assumptions on the data generating processes. They also allow optional stopping, so a forecast user may decide to interrupt evaluation taking into account the available data at any time and still draw statistically valid inference, which is generally not true for classical p-value based tests. In a case study on postprocessing of precipitation forecasts, state-of-the-art forecasts dominance tests and e-values lead to the same conclusions.

stat.ME

Sequentially valid tests for forecast calibration

Forecasting and forecast evaluation are inherently sequential tasks. Predictions are often issued on a regular basis, such as every hour, day, or month, and their quality is monitored continuously. However, the classical statistical tools for forecast evaluation are static, in the sense that statistical tests for forecast calibration are only valid if the evaluation period is fixed in advance. Recently, e-values have been introduced as a new, dynamic method for assessing statistical significance. An e-value is a non-negative random variable with expected value at most one under a null hypothesis. Large e-values give evidence against the null hypothesis, and the multiplicative inverse of an e-value is a conservative p-value. E-values are particularly suitable for sequential forecast evaluation, since they naturally lead to statistical tests which are valid under optional stopping. This article proposes e-values for testing probabilistic calibration of forecasts, which is one of the most important notions of calibration. The proposed methods are also more generally applicable for sequential goodness-of-fit testing. We demonstrate that the e-values are competitive in terms of power when compared to extant methods, which do not allow sequential testing. Furthermore, they provide important and useful insights in the evaluation of probabilistic weather forecasts.

stat.ME

Isotonic Distributional Regression

Isotonic distributional regression (IDR) is a powerful nonparametric technique for the estimation of conditional distributions under order restrictions. In a nutshell, IDR learns conditional distributions that are calibrated, and simultaneously optimal relative to comprehensive classes of relevant loss functions, subject to isotonicity constraints in terms of a partial order on the covariate space. Nonparametric isotonic quantile regression and nonparametric isotonic binary regression emerge as special cases. For prediction, we propose an interpolation method that generalizes extant specifications under the pool adjacent violators algorithm. We recommend the use of IDR as a generic benchmark technique in probabilistic forecast problems, as it does not involve any parameter tuning nor implementation choices, except for the selection of a partial order on the covariate space. The method can be combined with subsample aggregation, with the benefits of smoother regression functions and gains in computational efficiency. In a simulation study, we compare methods for distributional regression in terms of the continuous ranked probability score (CRPS) and $L_2$ estimation error, which are closely linked. In a case study on raw and postprocessed quantitative precipitation forecasts from a leading numerical weather prediction system, IDR is competitive with state of the art techniques.

stat.ME