SearcharxivSearch

arXiv subjects

Stefan Sperlich

Publications and source records attributed to Stefan Sperlich.

14 recordsLinked to original sources

An Optimal Transportation Approach for Improved Confidence Intervals

Optimal transport methods have recently attracted a lot of attention in statistics. Their appeal lies in providing a geometric framework for comparing probability measures, leading to new perspectives on classical problems. A central problem in statistics is the construction of valid confidence sets as fundamental inferential tools in practice. A well-known problem is that for complex problems or relatively small samples, their asymptotic approximations often show poor performance. This suggests to apply optimal transport methods when constructing confidence sets for hard problems to improve their coverage properties. We introduce such a procedure, derive the theoretical framework studying consistency and error bounds for the coverage probability of the resulting intervals. To guarantee feasibility in practice, we propose data-driven choices for our hyper parameters. This approach extends classical quantile-based confidence intervals by leveraging optimal couplings to minimize coverage deviations. Simulations demonstrate striking performance in different estimation problems, outperforming standard methods in accuracy and robustness.

stat.ME

Inference on panel data models with a generalized factor structure

We consider identification, inference and validation of linear panel data models when both factors and factor loadings are accounted for by a nonparametric function. This general specification encompasses rather popular models such as the two-way fixed effects and the interactive fixed effects ones. By applying a conditional mean independence assumption between unobserved heterogeneity and the covariates, we obtain consistent estimators of the parameters of interest at the optimal rate of convergence, for fixed and large $T$. We also provide a specification test for the modeling assumption based on the methodology of conditional moment tests and nonparametric estimation techniques. Using degenerate and nondegenerate theories of U-statistics we show its convergence and asymptotic distribution under the null, and that it diverges under the alternative at a rate arbitrarily close to $\sqrt{NT}$. Finite sample inference is based on bootstrap. Simulations reveal an excellent performance of our methods and an empirical application is conducted.

econ.EM

Simple bootstrap for linear mixed effects under model misspecification

Linear mixed effects are considered excellent predictors of cluster-level parameters in various domains. However, previous work has shown that their performance can be seriously affected by departures from modelling assumptions. Since the latter are common in applied studies, there is a need for inferential methods which are to certain extent robust to misspecfications, but at the same time simple enough to be appealing for practitioners. We construct statistical tools for cluster-wise and simultaneous inference for mixed effects under model misspecification using straightforward semiparametric random effect bootstrap. In our theoretical analysis, we show that our methods are asymptotically consistent under general regularity conditions. In simulations our intervals were robust to severe departures from model assumptions and performed better than their competitors in terms of empirical coverage probability.

stat.ME

Marginal and Conditional Multiple Inference for Linear Mixed Model Predictors

In spite of its high practical relevance, cluster specific multiple inference for linear mixed model predictors has hardly been addressed so far. While marginal inference for population parameters is well understood, conditional inference for the cluster specific predictors is more intricate. This work introduces a general framework for multiple inference in linear mixed models for cluster specific predictors. Consistent confidence sets for multiple inference are constructed under both, the marginal and the conditional law. Furthermore, it is shown that, remarkably, corresponding multiple marginal confidence sets are also asymptotically valid for conditional inference. Those lend themselves for testing linear hypotheses using standard quantiles without the need of re-sampling techniques. All findings are validated in simulations and illustrated along a study on Covid-19 mortality in US state prisons.

math.ST

Post-selection inference for linear mixed model parameters using the conditional Akaike information criterion

We investigate the issue of post-selection inference for a fixed and a mixed parameter in a linear mixed model using a conditional Akaike information criterion as a model selection procedure. Within the framework of linear mixed models we develop complete theory to construct confidence intervals for regression and mixed parameters under three frameworks: nested and general model sets as well as misspecified models. Our theoretical analysis is accompanied by a simulation experiment and a post-selection examination on mean income across Galicia's counties. Our numerical studies confirm a good performance of our new procedure. Moreover, they reveal a startling robustness to the model misspecification of a naive method to construct the confidence intervals for a mixed parameter which is in contrast to our findings for the fixed parameters.

stat.ME

Simultaneous Inference for Empirical Best Predictors with a Poverty Study in Small Areas

Today, generalized linear mixed models are broadly used in many fields. However, the development of tools for performing simultaneous inference has been largely neglected in this domain. A framework for joint inference is indispensable to carry out statistically valid multiple comparisons of parameters of interest between all or several clusters. We therefore develop simultaneous confidence intervals and multiple testing procedures for empirical best predictors under generalized linear mixed models. In addition, we implement our methodology to study widely employed examples of mixed models, that is, the unit-level binomial, the area-level Poisson-gamma and the area-level Poisson-lognormal mixed models. The asymptotic results are accompanied by extensive simulations. A case study on predicting poverty rates illustrates applicability and advantages of our simultaneous inference tools.

stat.AP

New Bias Calibration for Robust Estimation in Small Areas

Using sample surveys as a cost effective tool to provide estimates for characteristics of interest at population and sub-populations (area/domain) level has a long tradition in "small area estimation". However, the existence of outliers in the sample data can significantly affect the estimation for areas in which they occur, especially where the domain-sample size is small. Based on existing robust estimators for small area estimation we propose two novel approaches for bias calibration. A series of simulations shows that our methods lead to more efficient estimators in comparison with other existing bias-calibration methods. As a real data example we apply our estimators to obtain \textit{Gini} coefficients in labour market areas of the Tuscany region of Italy, where our sources of information are the EU-SILC survey and the Italian census. This analysis shows that the new methods reveal a different picture than existing methods. We extend our ideas to predictions for non-sampled areas.

stat.ME

The economics of minority language use: theory and empirical evidence for a language game model

Language and cultural diversity is a fundamental aspect of the present world. We study three modern multilingual societies -- the Basque Country, Ireland and Wales -- which are endowed with two, linguistically distant, official languages: $A$, spoken by all individuals, and $B$, spoken by a bilingual minority. In the three cases it is observed a decay in the use of minoritarian $B$, a sign of diversity loss. However, for the "Council of Europe" the key factor to avoid the shift of $B$ is its use in all domains. Thus, we investigate the language choices of the bilinguals by means of an evolutionary game theoretic model. We show that the language population dynamics has reached an evolutionary stable equilibrium where a fraction of bilinguals have shifted to speak $A$. Thus, this equilibrium captures the decline in the use of $B$. To test the theory we build empirical models that predict the use of $B$ for each proportion of bilinguals. We show that model-based predictions fit very well the observed use of Basque, Irish, and Welsh.

econ.EM

The Africa-Dummy: Gone with the Millennium?

A fixed effects regression estimator is introduced that can directly identify and estimate the Africa-Dummy in one regression step so that its correct standard errors as well as correlations to other coefficients can easily be estimated. We can estimate the Nickel bias and found it to be negligibly tiny. Semiparametric extensions check whether the Africa-Dummy is simply a result of misspecification of the functional form. In particular, we show that the returns to growth factors are different for Sub-Saharan African countries compared to the rest of the world. For example, returns to population growth are positive and beta-convergence is faster. When extending the model to identify the development of the Africa-Dummy over time we see that it has been changing dramatically over time and that the punishment for Sub-Saharan African countries has been decreasing incrementally to reach insignificance around the turn of the millennium.

econ.EM

A Varying Coefficient Model for Assessing the Returns to Growth to Account for Poverty and Inequality

Various papers demonstrate the importance of inequality, poverty and the size of the middle class for economic growth. When explaining why these measures of the income distribution are added to the growth regression, it is often mentioned that poor people behave different which may translate to the economy as a whole. However, simply adding explanatory variables does not reflect this behavior. By a varying coefficient model we show that the returns to growth differ a lot depending on poverty and inequality. Furthermore, we investigate how these returns differ for the poorer and for the richer part of the societies. We argue that the differences in the coefficients impede, on the one hand, that the means coefficients are informative, and, on the other hand, challenge the credibility of the economic interpretation. In short, we show that, when estimating mean coefficients without accounting for poverty and inequality, the estimation is likely to suffer from a serious endogeneity bias.

econ.EM

Kernel-based semiparametric multinomial logit modelling of political party affiliation

Conventional, parametric multinomial logit models are in general not sufficient for detecting the complex patterns voter profiles nowadays typically exhibit. In this manuscript, we use a semiparametric multinomial logit model to give a detailed analysis of the composition of a subsample of the German electorate in 2006. Germany is a particularly strong case for more flexible nonparametric approaches in this context, since due to the reunification and the preceding different political histories the composition of the electorate is very complex and nuanced. Our analysis reveals strong interactions of the covariates age and income, and highly nonlinear shapes of the factor impacts for each party's likelihood to be voted. Notably, we develop and provide a smoothed likelihood estimator for semiparametric multinomial logit models, which can be applied also in other application fields, such as, e.g., marketing.

stat.AP

A comparative study of new cross-validated bandwidth selectors for kernel density estimation

Recent contributions to kernel smoothing show that the performance of cross-validated bandwidth selectors improve significantly from indirectness. Indirect crossvalidation first estimates the classical cross-validated bandwidth from a more rough and difficult smoothing problem than the original one and then rescales this indirect bandwidth to become a bandwidth of the original problem. The motivation for this approach comes from the observation that classical crossvalidation tends to work better when the smoothing problem is difficult. In this paper we find that the performance of indirect crossvalidation improves theoretically and practically when the polynomial order of the indirect kernel increases, with the Gaussian kernel as limiting kernel when the polynomial order goes to infinity. These theoretical and practical results support the often proposed choice of the Gaussian kernel as indirect kernel. However, for do-validation our study shows a discrepancy between asymptotic theory and practical performance. As for indirect crossvalidation, in asymptotic theory the performance of indirect do-validation improves with increasing polynomial order of the used indirect kernel. But this theoretical improvements do not carry over to practice and the original do-validation still seems to be our preferred bandwidth selector. We also consider plug-in estimation and combinations of plug-in bandwidths and crossvalidated bandwidths. These latter bandwidths do not outperform the original do-validation estimator either.

stat.ME

Fractional White Noise Perturbations of Parabolic Volterra Equations

Aim of this work is to extend the results of Clément, Da Prato & Prüss on the fractional white noise perturbation with Hurst parameter 0<H<1. We will obtain similar results and it will turn out that the regularity of the solution u(t) of the stochastic Volterra equation increases with Hurst parameter H.

math.AP

Estimation of a semiparametric transformation model

This paper proposes consistent estimators for transformation parameters in semiparametric models. The problem is to find the optimal transformation into the space of models with a predetermined regression structure like additive or multiplicative separability. We give results for the estimation of the transformation when the rest of the model is estimated non- or semi-parametrically and fulfills some consistency conditions. We propose two methods for the estimation of the transformation parameter: maximizing a profile likelihood function or minimizing the mean squared distance from independence. First the problem of identification of such models is discussed. We then state asymptotic results for a general class of nonparametric estimators. Finally, we give some particular examples of nonparametric estimators of transformed separable models. The small sample performance is studied in several simulations.

math.ST