SearcharxivSearch

arXiv subjects

Philipp Sterzinger

Publications and source records attributed to Philipp Sterzinger.

4 recordsLinked to original sources

Proportional-limit asymptotics for Diaconis-Ylvisaker-penalised logistic regression with fitted intercept

This paper develops estimator-level asymptotic theory for maximum Diaconis-Ylvisaker prior penalised likelihood for logistic regression with a jointly fitted intercept and nonzero prior slope direction in the proportional-limit regime. For $\mathrm{N}(\mathbf{0}_p, p^{-1}\mathbf{I}_p)$ Gaussian covariates and $p/n\to\kappa\in(0,1)$, a conditional convex Gaussian min--max analysis yields almost-sure convergence of the fitted intercept and a pseudo-Lipschitz empirical law for the slope estimator. This estimator-level law gives asymptotic limits for out-of-sample scores, classification error, optimal thresholding and oracle calibration. It also yields oracle-adjusted fixed-block $Z$-statistics under isotropic Gaussian covariates, and yields the main ingredient in establishing the limiting distribution of the penalised likelihood-ratio test statistic and identifies the rescaling to recover the nominal chi-square distribution. We extend these results to Gaussian designs with arbitrary deterministic mean and positive-definite covariance via affine centering and whitening and discuss extensions to subgaussian covariates. Finally, we propose a consistent response-moment estimator of the oracle parameters entering the state equations that govern the slope limiting law and are required for feasible inference.

math.ST

Maximum softly penalised likelihood in factor analysis

Estimation in exploratory factor analysis often yields estimates on the boundary of the parameter space. Such occurrences, known as Heywood cases, are characterised by non-positive variance estimates and can cause issues in numerical optimisation procedures or convergence failures, which, in turn, can lead to misleading inferences, particularly regarding factor scores and model selection. We derive sufficient conditions on the model and a penalty to the log-likelihood function that i) guarantee the existence of maximum penalised likelihood estimates in the interior of the parameter space, and ii) ensure that the corresponding estimators possess the desirable asymptotic properties expected by the maximum likelihood estimator, namely consistency and asymptotic normality. Consistency and asymptotic normality are achieved when the penalisation is soft enough, in a way that adapts to the information accumulation about the model parameters. We formally show, for the first time, that the penalties of Akaike (1987) and Hirose et al. (2011) to the log-likelihood of the normal linear factor model satisfy the conditions for existence, and, hence, deal with Heywood cases. Their vanilla versions, though, can result in questionable finite-sample properties in estimation, inference, and model selection. The maximum softly-penalised likelihood framework we introduce enables the careful scaling of those penalties to ensure that the resulting estimation and inference procedures inherit the ML estimator's optimal properties. Through comprehensive simulation studies and the analysis of real data sets, we illustrate the desirable finite-sample properties of the maximum softly penalised likelihood estimators and associated procedures.

stat.ME

Diaconis-Ylvisaker prior penalized likelihood for $p/n \to \kappa \in (0,1)$ logistic regression

We characterise the behavior of the maximum Diaconis--Ylvisaker prior penalized likelihood estimator in high-dimensional logistic regression, where the number of covariates is a fraction $\kappa \in (0,1)$ of the number of observations $n$, as $n \to \infty$. We construct a rescaled estimator with zero asymptotic aggregate bias and define adjusted $Z$-statistics and rescaled penalized likelihood ratio statistics that exhibit the typical null asymptotic distributions, when the covariates are independent multivariate normal with an arbitrary covariance matrix and the linear predictor has asymptotic variance $\gamma^2$. While the maximum likelihood estimate asymptotically exists only for a narrow range of $(\kappa, \gamma)$ values, the maximum Diaconis--Ylvisaker prior penalized likelihood estimate always exists and can be computed directly using standard maximum likelihood routines. Thus, our asymptotic results extend to $(\kappa, \gamma)$ values where the maximum likelihood framework breaks down, with no additional implementation or computational cost. We study the estimator's shrinkage properties, compare the proposed estimation and inference procedures with alternatives that also accommodate proportional asymptotics, and formulate a conjecture -- supported by strong empirical evidence -- that extends our results when the model includes an intercept parameter. Finally, we propose estimation methods for all unknown constants involved in our procedures and demonstrate the theoretical advances through extensive simulation studies and the analysis of digit recognition data.

math.ST

Maximum softly-penalized likelihood for mixed effects logistic regression

Maximum likelihood estimation in logistic regression with mixed effects is known to often result in estimates on the boundary of the parameter space. Such estimates, which include infinite values for fixed effects and singular or infinite variance components, can cause havoc to numerical estimation procedures and inference. We introduce an appropriately scaled additive penalty to the log-likelihood function, or an approximation thereof, which penalizes the fixed effects by the Jeffreys' invariant prior for the model with no random effects and the variance components by a composition of negative Huber loss functions. The resulting maximum penalized likelihood estimates are shown to lie in the interior of the parameter space. Appropriate scaling of the penalty guarantees that the penalization is soft enough to preserve the optimal asymptotic properties expected by the maximum likelihood estimator, namely consistency, asymptotic normality, and Cramér-Rao efficiency. Our choice of penalties and scaling factor preserves equivariance of the fixed effects estimates under linear transformation of the model parameters, such as contrasts. Maximum softly-penalized likelihood is compared to competing approaches on two real-data examples, and through comprehensive simulation studies that illustrate its superior finite sample performance.

stat.ME