SearcharxivSearch

arXiv subjects

Marina Valdora

Publications and source records attributed to Marina Valdora.

9 recordsLinked to original sources

Robust estimation in generalized linear models based on the normal quantiles of the probability integral transformation

A new approach to robust estimation in generalized linear models is introduced. The idea of the method is to first transform the responses applying the composition of the normal quantile function and the probability integral transformation. Then, using that the transformed responses should follow a standard normal distribution, find the values of the parameters that minimize a robust measure of their size. In practice an approximation of this transformation is used. The proposed estimators are studied theoretically for distributions that depend on a single parameter and through simulations and examples for the particular cases of Poisson and logistic regression.

stat.ME

Robust Penalized Estimators for High--Dimensional Generalized Linear Models

Robust estimators for generalized linear models (GLMs) are not easy to develop due to the nature of the distributions involved. Recently, there has been growing interest in robust estimation methods, particularly in contexts involving a potentially large number of explanatory variables. Transformed M-estimators (MT-estimators) provide a natural extension of M-estimation techniques to the GLM framework, offering robust methodologies. We propose a penalized variant of MT-estimators to address high-dimensional data scenarios. Under suitable assumptions, we demonstrate the consistency and asymptotic normality of this novel class of estimators. Our theoretical development focuses on redescending rho-functions and penalization functions that satisfy specific regularity conditions. We present an Iterative Re-Weighted Least Squares algorithm, together with a deterministic initialization procedure, which is crucial since the estimating equations may have multiple solutions. We evaluate the finite sample performance of this method for Poisson distribution and well known penalization functions through Monte Carlo simulations that consider various types of contamination, as well as an empirical application using a real dataset.

stat.ME

An unbiased estimator of the case fatality rate

During an epidemic outbreak of a new disease, the probability of dying once infected is considered an important though difficult task to be computed. Since it is very hard to know the true number of infected people, the focus is placed on estimating the case fatality rate, which is defined as the probability of dying once tested and confirmed as infected. The estimation of this rate at the beginning of an epidemic remains challenging for several reasons, including the time gap between diagnosis and death, and the rapid growth in the number of confirmed cases. In this work, an unbiased estimator of the case fatality rate of a virus is presented. The consistency of the estimator is demonstrated, and its asymptotic distribution is derived, enabling the corresponding confidence intervals (C.I.) to be established. The proposed method is based on the distribution F of the time between confirmation and death of individuals who die because of the virus. The estimator's performance is analyzed in both simulation scenarios and the real-world context of Argentina in 2020 for the COVID-19 pandemic, consistently achieving excellent results when compared to an existing proposal as well as to the conventional \naive" estimator that was employed to report the case fatality rates during the last COVID-19 pandemic. In the simulated scenarios, the empirical coverage of our C.I. is studied, both using the F employed to generate the data and an estimated F, and it is observed that the desired level of confidence is reached quickly when using real F and in a reasonable period of time when estimating F.

stat.ME

Robust estimation for functional logistic regression models

This paper addresses the problem of providing robust estimators under a functional logistic regression model. Logistic regression is a popular tool in classification problems with two populations. As in functional linear regression, regularization tools are needed to compute estimators for the functional slope. The traditional methods are based on dimension reduction or penalization combined with maximum likelihood or quasi--likelihood techniques and for that reason, they may be affected by misclassified points especially if they are associated to functional covariates with atypical behaviour. The proposal given in this paper adapts some of the best practices used when the covariates are finite--dimensional to provide reliable estimations. Under regularity conditions, consistency of the resulting estimators and rates of convergence for the predictions are derived. A numerical study illustrates the finite sample performance of the proposed method and reveals its stability under different contamination scenarios. A real data example is also presented.

stat.ME

Each student with her/his own data: understanding sampling distributions

Sampling distribution, a foundational concept in statistics, is difficult to understand, since we usually have only one realization of the estimator of interest. In this work, we present an innovative method for helping university students understand the variability of an estimator. In our approach, each student uses a different data set, getting diverse estimations. Then, sharing the results, we can empirically study the sampling distribution. After some "handmade" experiences, we have built a web page to deliver a personalized data set for each student. Through this web page, we can also reformulate the role of the student in the classroom, inviting him/her to became an active player, submitting the solution to different problems and checking whether they are correct. In this work we present a recent experience in such direction.

stat.OT

Robust Doubly Protected Estimators for Quantiles with Missing Data

Doubly protected estimators are widely used for estimating the population mean of an outcome Y from a sample where the response is missing in some individuals. To compensate for the missing responses, a vector X of covariates is observed at each individual, and the missing mechanism is assumed to be independent of the response, conditioned on X (missing at random). In recent years, many authors have moved from the mean to the median, and more generally, doubly protected estimators of the quantiles have been proposed, assuming a parametric regression model for the relationship between X and Y and a parametric form for the propensity score. In this work, we present doubly protected estimators for the quantiles that are also robust, in the sense that they are resistant to the presence of outliers in the sample. We also flexibilize the model for the relationship between X and Y . Thus we present robust doubly protected estimators for the quantiles of the response in the presence of missing observations, postulating a semiparametric regression model for the relationship between the response and the covariates and a parametric model for the propensity score.

stat.ME

Robust Estimation in High Dimensional Generalized Linear Models

Generalized Linear Models are routinely used in data analysis. The classical procedures for estimation are based on Maximum Likelihood and it is well known that the presence of outliers can have a large impact on this estimator. Robust procedures are presented in the literature but they need a robust initial estimate in order to be computed. This is especially important for robust procedures with non convex loss function such as redescending M-estimators. Subsampling techniques are often used to determine a robust initial estimate; however when the number of unknown parameters is large the number of subsamples needed in order to have a high probability of having one subsample free of outliers become infeasible. Furthermore the subsampling procedure provides a non deterministic starting point. Based on ideas in Pena and Yohai (1999), we introduce a deterministic robust initial estimate for M-estimators based on transformations Valdora and Yohai (2014) for which we also develop an iteratively reweighted least squares algorithm. The new methods are studied by Monte Carlo experiments.

stat.CO

Models for the Propensity Score that Contemplate the Positivity Assumption and their Application to Missing Data and Causality

Generalized linear models are often assumed to fit propensity scores, which are used to compute inverse probability weighted (IPW) estimators. In order to derive the asymptotic properties of IPW estimators, the propensity score is supposed to be bounded away from cero. This condition is known in the literature as strict positivity (or positivity assumption) and, in practice, when it does not hold, IPW estimators are very unstable and have a large variability. Although strict positivity is often assumed, it is not upheld when some of the covariates are continuous. In this work, we attempt to conciliate between the strict positivity condition and the theory of generalized linear models by incorporating an extra parameter, which results in an explicit lower bound for the propensity scores.

stat.ME

Robust estimators for generalized linear models with a dispersion parameter

Highly robust and efficient estimators for the generalized linear model with a dispersion parameter are proposed. The estimators are based on three steps. In the first step the maximum rank correlation estimator is used to consistently estimate the slopes up to a scale factor. In the second step, the scale factor, the intercept, and the dispersion parameter are consistently estimated using a MT-estimator of a simple regression model. The combined estimator is highly robust but inefficient. Then, randomized quantile residuals based on the initial estimators are used to detect outliers to be rejected and to define a set S of observations to be retained. Finally, a conditional maximum likelihood (CML) estimator given the observations in S is computed. We show that, under the model, S tends to the complete sample for increasing sample size. Therefore, the CML tends to the unconditional maximum likelihood estimator. It is therefore highly efficient, while maintaining the high degree of robustness of the initial estimator. The case of the negative binomial regression model is studied in detail.

stat.ME