SearcharxivSearch

arXiv subjects

Katarzyna Reluga

Publications and source records attributed to Katarzyna Reluga.

9 recordsLinked to original sources

Distribution Shift in Missing Data Imputation: A Risk-Based Perspective and Importance-Weighted Correction under MAR

Missing data imputation, where a model is trained on observed data to estimate unobserved values, is a fundamental problem in machine learning. In this paper, we rigorously formulate imputation model learning as a mean-squared error risk minimisation problem. We show that when the probability of missingness depends on the data, many state-of-the-art methods fail to account for the resulting distribution shift between the observed data used for training and the full data distribution used for evaluation. Consequently, these approaches do not minimise mean-squared error on the full data distribution. Instead, we propose a novel imputation algorithm designed to learn an imputation model from the observed data while explicitly accounting for this distribution shift. Simulation studies show consistent improvements over otherwise identical uncorrected baselines, with average reductions of 3% in RMSE and 7% in Wasserstein distance.

stat.ML

Direct Doubly Robust Estimation of Conditional Quantile Contrasts

Within heterogeneous treatment effect (HTE) analysis, various estimands have been proposed to capture the effect of a treatment conditional on covariates. Recently, the conditional quantile comparator (CQC) has emerged as a promising estimand, offering quantile-level summaries akin to the conditional quantile treatment effect (CQTE) while preserving some interpretability of the conditional average treatment effect (CATE). It achieves this by summarising the treated response conditional on both the covariates and the untreated response. Despite these desirable properties, the CQC's current estimation is limited by the need to first estimate the difference in conditional cumulative distribution functions and then invert it. This inversion obscures the CQC estimate, hampering our ability to both model and interpret it. To address this, we propose the first direct estimator of the CQC, allowing for explicit modelling and parameterisation. This explicit parameterisation enables better interpretation of our estimate while also providing a means to constrain and inform the model. We show, both theoretically and empirically, that our estimation error depends directly on the complexity of the CQC itself, improving upon the existing estimation procedure. Furthermore, it retains the desirable double robustness property with respect to nuisance parameter estimation. We further show our method to outperform existing procedures in estimation accuracy across multiple data scenarios while varying sample size and nuisance error. Finally, we apply it to real-world data from an employment scheme, uncovering a reduced range of potential earnings improvement as participant age increases.

stat.ME

The impact of job stability on monetary poverty in Italy: causal small area estimation

Job stability - encompassing secure contracts, adequate wages, social benefits, and career opportunities - is a critical determinant in reducing monetary poverty, as it provides households with reliable income and enhances economic well-being. This study leverages EU-SILC survey and census data to estimate the causal effect of job stability on monetary poverty across Italian provinces, quantifying its influence and analyzing regional disparities. We introduce a novel causal small area estimation (CSAE) framework that integrates global and local estimation strategies for heterogeneous treatment effect estimation, effectively addressing data sparsity at the provincial level. Furthermore, we develop a general bootstrap scheme to construct reliable confidence intervals, applicable regardless of the method used for estimating nuisance parameters. Extensive simulation studies demonstrate that our proposed estimators outperform classical causal inference methods in terms of stability while maintaining computational scalability for large datasets. Applying this methodology to real-world data, we uncover significant relationships between job stability and poverty across six Italian regions, offering critical insights into regional disparities and their implications for evidence-based policy design.

stat.AP

Conditional Outcome Equivalence: A Quantile Alternative to CATE

Conditional quantile treatment effect (CQTE) can provide insight into the effect of a treatment beyond the conditional average treatment effect (CATE). This ability to provide information over multiple quantiles of the response makes CQTE especially valuable in cases where the effect of a treatment is not well-modelled by a location shift, even conditionally on the covariates. Nevertheless, the estimation of CQTE is challenging and often depends upon the smoothness of the individual quantiles as a function of the covariates rather than smoothness of the CQTE itself. This is in stark contrast to CATE where it is possible to obtain high-quality estimates which have less dependency upon the smoothness of the nuisance parameters when the CATE itself is smooth. Moreover, relative smoothness of the CQTE lacks the interpretability of smoothness of the CATE making it less clear whether it is a reasonable assumption to make. We combine the desirable properties of CATE and CQTE by considering a new estimand, the conditional quantile comparator (CQC). The CQC not only retains information about the whole treatment distribution, similar to CQTE, but also having more natural examples of smoothness and is able to leverage simplicity in an auxiliary estimand. We provide finite sample bounds on the error of our estimator, demonstrating its ability to exploit simplicity. We validate our theory in numerical simulations which show that our method produces more accurate estimates than baselines. Finally, we apply our methodology to a study on the effect of employment incentives on earnings across different age groups. We see that our method is able to reveal heterogeneity of the effect across different quantiles.

stat.ME

A unified analysis of regression adjustment in randomized experiments

Regression adjustment is broadly applied in randomized trials under the premise that it usually improves the precision of a treatment effect estimator. However, previous work has shown that this is not always true. To further understand this phenomenon, we develop a unified comparison of the asymptotic variance of a class of linear regression-adjusted estimators. Our analysis is based on the classical theory for linear regression with heteroscedastic errors and thus does not assume that the postulated linear model is correct. For a completely randomized binary treatment, we provide sufficient conditions under which some regression-adjusted estimators are guaranteed to be more asymptotically efficient than others. We explore other settings such as general treatment assignment mechanisms and generalized linear models, and find that the variance dominance phenomenon no longer occurs.

stat.ME

Simple bootstrap for linear mixed effects under model misspecification

Linear mixed effects are considered excellent predictors of cluster-level parameters in various domains. However, previous work has shown that their performance can be seriously affected by departures from modelling assumptions. Since the latter are common in applied studies, there is a need for inferential methods which are to certain extent robust to misspecfications, but at the same time simple enough to be appealing for practitioners. We construct statistical tools for cluster-wise and simultaneous inference for mixed effects under model misspecification using straightforward semiparametric random effect bootstrap. In our theoretical analysis, we show that our methods are asymptotically consistent under general regularity conditions. In simulations our intervals were robust to severe departures from model assumptions and performed better than their competitors in terms of empirical coverage probability.

stat.ME

Post-selection inference for linear mixed model parameters using the conditional Akaike information criterion

We investigate the issue of post-selection inference for a fixed and a mixed parameter in a linear mixed model using a conditional Akaike information criterion as a model selection procedure. Within the framework of linear mixed models we develop complete theory to construct confidence intervals for regression and mixed parameters under three frameworks: nested and general model sets as well as misspecified models. Our theoretical analysis is accompanied by a simulation experiment and a post-selection examination on mean income across Galicia's counties. Our numerical studies confirm a good performance of our new procedure. Moreover, they reveal a startling robustness to the model misspecification of a naive method to construct the confidence intervals for a mixed parameter which is in contrast to our findings for the fixed parameters.

stat.ME

Simultaneous inference for linear mixed model parameters with an application to small area estimation

Over the past decades, linear mixed models have attracted considerable attention in various fields of applied statistics. They are popular whenever clustered, hierarchical or longitudinal data are investigated. Nonetheless, statistical tools for valid simultaneous inference for mixed parameters are rare. This is surprising because one often faces inferential problems beyond the pointwise examination of fixed or mixed parameters. For example, there is an interest in a comparative analysis of cluster-level parameters or subject-specific estimates in studies with repeated measurements. We discuss methods for simultaneous inference assuming a linear mixed model. Specifically, we develop simultaneous prediction intervals as well as multiple testing procedures for mixed parameters. They are useful for joint considerations or comparisons of cluster-level parameters. We employ a consistent bootstrap approximation of the distribution of max-type statistic to construct our tools. The numerical performance of the developed methodology is studied in simulation experiments and illustrated in a data example on household incomes in small areas.

stat.ME

Simultaneous Inference for Empirical Best Predictors with a Poverty Study in Small Areas

Today, generalized linear mixed models are broadly used in many fields. However, the development of tools for performing simultaneous inference has been largely neglected in this domain. A framework for joint inference is indispensable to carry out statistically valid multiple comparisons of parameters of interest between all or several clusters. We therefore develop simultaneous confidence intervals and multiple testing procedures for empirical best predictors under generalized linear mixed models. In addition, we implement our methodology to study widely employed examples of mixed models, that is, the unit-level binomial, the area-level Poisson-gamma and the area-level Poisson-lognormal mixed models. The asymptotic results are accompanied by extensive simulations. A case study on predicting poverty rates illustrates applicability and advantages of our simultaneous inference tools.

stat.AP