SearcharxivSearch

arXiv subjects

Ioannis Kosmidis

Publications and source records attributed to Ioannis Kosmidis.

At least 19 recordsLinked to original sources

Focused median bias reduction

Median bias reduction of maximum likelihood estimators can substantially improve estimation and inference. Existing generally applicable methods are, however, typically implicit, requiring the solution of nonlinear systems of estimating equations, which is computationally demanding. They also require a fully specified nuisance parameterization, and their application to transformations of parameters involves tedious algebra and bespoke implementations. We develop an explicit median bias-corrected estimator for focus parameters that are smooth scalar transformations of a chosen reference parameterization. The estimator is obtained by solving, to the required order, an equation based on the Cornish-Fisher expansion of the centred and scaled maximum likelihood estimator of the focus parameter. It requires only the maximum likelihood or an asymptotically equivalent estimator at the reference parameterization, the gradient and Hessian of the transformation, and expectations of products of log-likelihood derivatives. These expectations are available for many models from the existing bias reduction literature and can also be estimated by Monte Carlo simulation. The resulting estimators are third-order median unbiased and provide one-step approximations to estimators from implicit median bias reduction when the focus parameter is included in the reference parameterization. The method improves standard asymptotic inference and integrates naturally with hull-based confidence procedures, yielding intervals with near nominal finite-sample coverage under median bias control. We demonstrate the framework through post-selection inference using the Focused Information Criterion, Mahalanobis distances, quantiles, and other scalar focus parameters in regression, circular, and stratified models.

stat.ME

Diaconis-Ylvisaker prior penalized likelihood for $p/n \to κ\in (0,1)$ logistic regression

We characterise the behavior of the maximum Diaconis--Ylvisaker prior penalized likelihood estimator in high-dimensional logistic regression, where the number of covariates is a fraction $κ\in (0,1)$ of the number of observations $n$, as $n \to \infty$. We construct a rescaled estimator with zero asymptotic aggregate bias and define adjusted $Z$-statistics and rescaled penalized likelihood ratio statistics that exhibit the typical null asymptotic distributions, when the covariates are independent multivariate normal with an arbitrary covariance matrix and the linear predictor has asymptotic variance $γ^2$. While the maximum likelihood estimate asymptotically exists only for a narrow range of $(κ, γ)$ values, the maximum Diaconis--Ylvisaker prior penalized likelihood estimate always exists and can be computed directly using standard maximum likelihood routines. Thus, our asymptotic results extend to $(κ, γ)$ values where the maximum likelihood framework breaks down, with no additional implementation or computational cost. We study the estimator's shrinkage properties, compare the proposed estimation and inference procedures with alternatives that also accommodate proportional asymptotics, and formulate a conjecture -- supported by strong empirical evidence -- that extends our results when the model includes an intercept parameter. Finally, we propose estimation methods for all unknown constants involved in our procedures and demonstrate the theoretical advances through extensive simulation studies and the analysis of digit recognition data.

math.ST

Bias-Reduced Estimation of Structural Equation Models

Finite-sample bias is a pervasive challenge in the estimation of structural equation models (SEMs), especially when sample sizes are small or measurement reliability is low. A range of methods have been proposed to improve finite-sample bias in the SEM literature, ranging from analytic bias corrections to resampling-based techniques, with each carrying trade-offs in scope, computational burden, and statistical performance. We apply the reduced-bias M-estimation framework (RBM, Kosmidis & Lunardon, 2024, J. R. Stat. Soc. Series B Stat. Methodol.) to SEMs. The RBM framework is attractive as it requires only first- and second-order derivatives of the log-likelihood, which renders it both straightforward to implement, and computationally more efficient compared to resampling-based alternatives such as bootstrap and jackknife. It is also robust to departures from modelling assumptions. Using the same simulation setup as in Dhaene and Rosseel (2022), we illustrate that RBM estimators consistently reduce mean bias in the estimation of SEMs without inflating mean squared error. They also deliver improvements in both median bias and inference relative to maximum likelihood estimators, while maintaining robustness under non-normality. Our findings suggest that RBM offers a promising, practical, and broadly applicable tool for mitigating bias in the estimation of SEMs, particularly in small-sample research contexts.

stat.ME

Extended-support beta regression for $[0, 1]$ responses

We introduce the XBX regression model, a continuous mixture of extended-support beta regressions for modelling bounded responses with boundary observations. The core building block of XBX regression is the extended-support beta distribution, a censored version of a four-parameter beta distribution with the same exceedance on the left and right of $(0, 1)$. Hence, XBX regression is a direct extension of beta regression. We prove that beta regression and heteroscedastic normal regression with censoring at both $0$ and $1$ -- also known as the heteroscedastic two-limit tobit model in the econometrics literature -- are special cases of extended-support beta regression, depending on whether a single extra parameter is zero or infinity, respectively. To overcome identifiability issues due to the similarity of the beta and normal distributions for certain parameter values, we shrink towards beta regression by letting the extra parameter have an exponential distribution with unknown mean. A Gauss-Laguerre quadrature approximation results in efficient likelihood-based estimation and inference procedures, which the betareg R package implements since version 3.2.0. We use XBX regression to analyze investment decisions in a behavioural economics experiment about the occurrence and extent of loss aversion. In contrast to standard approaches, we capture both the probability of rational behaviour and the mean of loss aversion. Extensive comparisons with alternative approaches illustrate the effectiveness of the new model.

stat.ME

Jeffreys-prior penalty for high-dimensional logistic regression: A conjecture about aggregate bias

Firth (1993, Biometrika) shows that the maximum Jeffreys' prior penalized likelihood estimator in logistic regression has asymptotic bias decreasing with the square of the number of observations when the number of parameters is fixed, which is an order faster than the typical rate from maximum likelihood. The widespread use of that estimator in applied work is supported by the results in Kosmidis and Firth (2021, Biometrika), who show that it takes finite values, even in cases where the maximum likelihood estimate does not exist. Kosmidis and Firth (2021, Biometrika) also provide empirical evidence that the estimator has good bias properties in high-dimensional settings where the number of parameters grows asymptotically linearly but slower than the number of observations. We design and carry out a large-scale computer experiment covering a wide range of such high-dimensional settings and produce strong empirical evidence for a simple rescaling of the maximum Jeffreys' prior penalized likelihood estimator that delivers high accuracy in signal recovery, in terms of aggregate bias, in the presence of an intercept parameter. The rescaled estimator is effective even in cases where estimates from maximum likelihood and other recently proposed corrective methods based on approximate message passing do not exist.

stat.ME

Bounded-memory adjusted scores estimation in generalized linear models with large data sets

The widespread use of maximum Jeffreys'-prior penalized likelihood in binomial-response generalized linear models, and in logistic regression, in particular, are supported by the results of Kosmidis and Firth (2021, Biometrika), who show that the resulting estimates are always finite-valued, even in cases where the maximum likelihood estimates are not, which is a practical issue regardless of the size of the data set. In logistic regression, the implied adjusted score equations are formally bias-reducing in asymptotic frameworks with a fixed number of parameters and appear to deliver a substantial reduction in the persistent bias of the maximum likelihood estimator in high-dimensional settings where the number of parameters grows asymptotically as a proportion of the number of observations. In this work, we develop and present two new variants of iteratively reweighted least squares for estimating generalized linear models with adjusted score equations for mean bias reduction and maximization of the likelihood penalized by a positive power of the Jeffreys-prior penalty, which eliminate the requirement of storing $O(n)$ quantities in memory, and can operate with data sets that exceed computer memory or even hard drive capacity. We achieve that through incremental QR decompositions, which enable IWLS iterations to have access only to data chunks of predetermined size. Both procedures can also be readily adapted to fit generalized linear models when distinct parts of the data is stored across different sites and, due to privacy concerns, cannot be fully transferred across sites. We assess the procedures through a real-data application with millions of observations.

stat.ME

Bayesian Tensor Factorisations for Time Series of Counts

We propose a flexible nonparametric Bayesian modelling framework for multivariate time series of count data based on tensor factorisations. Our models can be viewed as infinite state space Markov chains of known maximal order with non-linear serial dependence through the introduction of appropriate latent variables. Alternatively, our models can be viewed as Bayesian hierarchical models with conditionally independent Poisson distributed observations. Inference about the important lags and their complex interactions is achieved via MCMC. When the observed counts are large, we deal with the resulting computational complexity of Bayesian inference via a two-step inferential strategy based on an initial analysis of a training set of the data. Our methodology is illustrated using simulation experiments and analysis of real-world data.

stat.ME

Empirical bias-reducing adjustments to estimating functions

We develop a novel and general framework for reduced-bias $M$-estimation from asymptotically unbiased estimating functions. The framework relies on an empirical approximation of the bias by a function of derivatives of estimating function contributions. Reduced-bias $M$-estimation operates either implicitly, by solving empirically-adjusted estimating equations, or explicitly, by subtracting the estimated bias from the original $M$-estimates, and applies to models that are partially- or fully-specified, with either likelihoods or other surrogate objectives. Automatic differentiation can be used to abstract away the only algebra required to implement reduced-bias $M$-estimation. As a result, the bias reduction methods we introduce have markedly broader applicability with more straightforward implementation and less algebraic or computational effort than other established bias-reduction methods that require resampling or evaluation of expectations of products of log-likelihood derivatives. If $M$-estimation is by maximizing an objective, then there always exists a bias-reducing penalized objective. That penalized objective relates closely to information criteria for model selection, and can be further enhanced with plug-in penalties to deliver reduced-bias $M$-estimates with extra properties, like finiteness in models for categorical data. The reduced-bias $M$-estimators have the same asymptotic distribution as the original $M$-estimators, and, hence, standard procedures for inference and model selection apply unaltered with the improved estimates. We demonstrate and assess the properties of reduced-bias $M$-estimation in well-used, prominent modelling settings of varying complexity.

stat.ME

Scalable Marked Point Processes for Exchangeable and Non-Exchangeable Event Sequences

We adopt the interpretability offered by a parametric, Hawkes-process-inspired conditional probability mass function for the marks and apply variational inference techniques to derive a general and scalable inferential framework for marked point processes. The framework can handle both exchangeable and non-exchangeable event sequences with minimal tuning and without any pre-training. This contrasts with many parametric and non-parametric state-of-the-art methods that typically require pre-training and/or careful tuning, and can only handle exchangeable event sequences. The framework's competitive computational and predictive performance against other state-of-the-art methods are illustrated through real data experiments. Its attractiveness for large-scale applications is demonstrated through a case study involving all events occurring in an English Premier League season.

stat.ML

Maximum softly-penalized likelihood for mixed effects logistic regression

Maximum likelihood estimation in logistic regression with mixed effects is known to often result in estimates on the boundary of the parameter space. Such estimates, which include infinite values for fixed effects and singular or infinite variance components, can cause havoc to numerical estimation procedures and inference. We introduce an appropriately scaled additive penalty to the log-likelihood function, or an approximation thereof, which penalizes the fixed effects by the Jeffreys' invariant prior for the model with no random effects and the variance components by a composition of negative Huber loss functions. The resulting maximum penalized likelihood estimates are shown to lie in the interior of the parameter space. Appropriate scaling of the penalty guarantees that the penalization is soft enough to preserve the optimal asymptotic properties expected by the maximum likelihood estimator, namely consistency, asymptotic normality, and Cramér-Rao efficiency. Our choice of penalties and scaling factor preserves equivariance of the fixed effects estimates under linear transformation of the model parameters, such as contrasts. Maximum softly-penalized likelihood is compared to competing approaches on two real-data examples, and through comprehensive simulation studies that illustrate its superior finite sample performance.

stat.ME

Flexible marked spatio-temporal point processes with applications to event sequences from association football

We develop a new family of marked point processes by focusing the characteristic properties of marked Hawkes processes exclusively to the space of marks, providing the freedom to specify a different model for the occurrence times. This is possible through the decomposition of the joint distribution of marks and times that allows to separately specify the conditional distribution of marks given the filtration of the process and the current time. We develop a Bayesian framework for the inference and prediction from this family of marked point processes that can naturally accommodate process and point-specific covariate information to drive cross-excitations, offering wide flexibility and applicability in the modelling of real-world processes. The framework is used here for the modelling of in-game event sequences from association football, resulting not only in inferences about previously unquantified characteristics of game dynamics and extraction of event-specific team abilities, but also in predictions for the occurrence of events of interest, such as goals, corners or fouls in a specified interval of time.

stat.AP

Mean and median bias reduction: A concise review and application to adjacent-categories logit models

The estimation of categorical response models using bias-reducing adjusted score equations has seen extensive theoretical research and applied use. The resulting estimates have been found to have superior frequentist properties to what maximum likelihood generally delivers and to be finite, even in cases where the maximum likelihood estimates are infinite. We briefly review mean and median bias reduction of maximum likelihood estimates via adjusted score equations in an illustration-driven way, and discuss their particular equivariance properties under parameter transformations. We then apply mean and median bias reduction to adjacent-categories logit models for ordinal responses. We show how ready bias reduction procedures for Poisson log-linear models can be used for mean and median bias reduction in adjacent-categories logit models with proportional odds and mean bias-reduced estimation in models with non-proportional odds. As in binomial logistic regression, the reduced-bias estimates are found to be finite even in cases where the maximum likelihood estimates are infinite. We also use the approximation of the bias of transformations of mean bias-reduced estimators to correct for the mean bias of model-based ordinal superiority measures. All developments are motivated and illustrated using real-data case studies and simulations

stat.ME

Parametric bootstrap inference for stratified models with high-dimensional nuisance specifications

Inference about a scalar parameter of interest typically relies on the asymptotic normality of common likelihood pivots, such as the signed likelihood root, the score and Wald statistics. Nevertheless, the resulting inferential procedures are known to perform poorly when the dimension of the nuisance parameter is large relative to the sample size and when the information about the parameters is limited. In many such cases, the use of asymptotic normality of analytical modifications of the signed likelihood root is known to recover inferential performance. It is proved here that parametric bootstrap of standard likelihood pivots results in as accurate inferences as analytical modifications of the signed likelihood root do in stratified models with stratum specific nuisance parameters. We focus on the challenging case where the number of strata increases as fast or faster than the stratum samples size. It is also shown that this equivalence holds regardless of whether constrained or unconstrained bootstrap is used. This is in contrast to when the number of strata is fixed or increases slower than the stratum sample size, where we show that constrained bootstrap corrects inference to a higher order than unconstrained bootstrap. Simulation experiments support the theoretical findings and demonstrate the excellent performance of bootstrap in extreme scenarios.

math.ST

Bias Reduction as a Remedy to the Consequences of Infinite Estimates in Poisson and Tobit Regression

Data separation is a well-studied phenomenon that can cause problems in the estimation and inference from binary response models. Complete or quasi-complete separation occurs when there is a combination of regressors in the model whose value can perfectly predict one or both outcomes. In such cases, and such cases only, the maximum likelihood estimates and the corresponding standard errors are infinite. It is less widely known that the same can happen in further microeconometric models. One of the few works in the area is Santos Silva and Tenreyro (2010) who note that the finiteness of the maximum likelihood estimates in Poisson regression depends on the data configuration and propose a strategy to detect and overcome the consequences of data separation. However, their approach can lead to notable bias on the parameter estimates when the regressors are correlated. We illustrate how bias-reducing adjustments to the maximum likelihood score equations can overcome the consequences of separation in Poisson and Tobit regression models.

stat.ME

Two-way sparsity for time-varying networks, with applications in genomics

We propose a novel way of modelling time-varying networks, by inducing two-way sparsity on local models of node connectivity. This two-way sparsity separately promotes sparsity across time and sparsity across variables (within time). Separation of these two types of sparsity is achieved through a novel prior structure, which draws on ideas from the Bayesian lasso and from copula modelling. We provide an efficient implementation of the proposed model via a Gibbs sampler, and we apply the model to data from neural development. In doing so, we demonstrate that the proposed model is able to identify changes in genomic network structure that match current biological knowledge. Such changes in genomic network structure can then be used by neuro-biologists to identify potential targets for further experimental investigation.

stat.ME

A Bayesian inference approach for determining player abilities in football

We consider the task of determining a football player's ability for a given event type, for example, scoring a goal. We propose an interpretable Bayesian model which is fit using variational inference methods. We implement a Poisson model to capture occurrences of event types, from which we infer player abilities. Our approach also allows the visualisation of differences between players, for a specific ability, through the marginal posterior variational densities. We then use these inferred player abilities to extend the Bayesian hierarchical model of Baio and Blangiardo (2010) which captures a team's scoring rate (the rate at which they score goals). We apply the resulting scheme to the English Premier League, capturing player abilities over the 2013/2014 season, before using output from the hierarchical model to predict whether over or under 2.5 goals will be scored in a given game in the 2014/2015 season. This validates our model as a way of providing insights into team formation and the individual success of sports teams.

stat.AP

Jeffreys-prior penalty, finiteness and shrinkage in binomial-response generalized linear models

Penalization of the likelihood by Jeffreys' invariant prior, or by a positive power thereof, is shown to produce finite-valued maximum penalized likelihood estimates in a broad class of binomial generalized linear models. The class of models includes logistic regression, where the Jeffreys-prior penalty is known additionally to reduce the asymptotic bias of the maximum likelihood estimator; and also models with other commonly used link functions such as probit and log-log. Shrinkage towards equiprobability across observations, relative to the maximum likelihood estimator, is established theoretically and is studied through illustrative examples. Some implications of finiteness and shrinkage for inference are discussed, particularly when inference is based on Wald-type procedures. A widely applicable procedure is developed for computation of maximum penalized likelihood estimates, by using repeated maximum likelihood fits with iteratively adjusted binomial responses and totals. These theoretical results and methods underpin the increasingly widespread use of reduced-bias and similarly penalized binomial regression models in many applied fields.

math.ST

Modelling rankings in R: the PlackettLuce package

This paper presents the R package PlackettLuce, which implements a generalization of the Plackett-Luce model for rankings data. The generalization accommodates both ties (of arbitrary order) and partial rankings (complete rankings of subsets of items). By default, the implementation adds a set of pseudo-comparisons with a hypothetical item, ensuring that the underlying network of wins and losses between items is always strongly connected. In this way, the worth of each item always has a finite maximum likelihood estimate, with finite standard error. The use of pseudo-comparisons also has a regularization effect, shrinking the estimated parameters towards equal item worth. In addition to standard methods for model summary, PlackettLuce provides a method to compute quasi standard errors for the item parameters. This provides the basis for comparison intervals that do not change with the choice of identifiability constraint placed on the item parameters. Finally, the package provides a method for model-based partitioning using covariates whose values vary between rankings, enabling the identification of subgroups of judges or settings that have different item worths. The features of the package are demonstrated through application to classic and novel data sets.

stat.CO