SearcharxivSearch

arXiv subjects

Maria Kateri

Publications and source records attributed to Maria Kateri.

8 recordsLinked to original sources

Sparse Latent Class Analysis For Dichotomous Responses: Post-Estimation Refinement via Item-level Pseudo-Likelihood

Latent Class Analysis (LCA) is widely used to identify unobserved subgroups in social and behavioural sciences. A long-standing challenge for LCA is the interpretability of the latent classes, due to the high complexity of the estimated item response probability matrix. To address this, we propose a computationally efficient post-estimation refinement procedure that enhances model interpretability by a sparse model estimate. The method begins by estimating a classical, unrestricted, latent class model and determining the number of classes using the Bayesian information criterion (BIC). It is followed by a refinement step that further performs model selection on the item-specific response probabilities based on the initial estimate. This refinement penalises the number of distinct response probability levels per item, collapsing redundant levels to yield a sparse matrix that is significantly easier to interpret than those produced by classical LCA. We provide asymptotic theory showing that the proposed procedure consistently recovers the sparse pattern of the item response probabilities for each item, and further validate its performance through extensive simulations. The practical power of the proposed method is further illustrated via an application to survey data on social role performance, where it provides a parsimonious and clear characterisation of the resulting latent classes. The code for implementing the proposed method is publicly available at https://github.com/florence07/Sparse-LCA-Refinement.

stat.ME

Simultaneous Factors Selection and Fusion of Their Levels in Penalized Logistic Regression

Nowadays, several data analysis problems require for complexity reduction, mainly meaning that they target at removing the non-influential covariates from the model and at delivering a sparse model. When categorical covariates are present, with their levels being dummy coded, the number of parameters included in the model grows rapidly, fact that emphasizes the need for reducing the number of parameters to be estimated. In this case, beyond variable selection, sparsity is also achieved through fusion of levels of covariates which do not differentiate significantly in terms of their influence on the response variable. In this work a new regularization technique is introduced, called $L_{0}$-Fused Group Lasso ($L_{0}$-FGL) for binary logistic regression. It uses a group lasso penalty for factor selection and for the fusion part it applies an $L_{0}$ penalty on the differences among the levels' parameters of a categorical predictor. Using adaptive weights, the adaptive version of $L_{0}$-FGL method is derived. Theoretical properties, such as the existence, $\sqrt{n}$ consistency and oracle properties under certain conditions, are established. In addition, it is shown that even in the diverging case where the number of parameters $p_{n}$ grows with the sample size $n$, $\sqrt{n}$ consistency and a consistency in variable selection result are achieved. Two computational methods, PIRLS and a block coordinate descent (BCD) approach using quasi Newton, are developed and implemented. A simulation study supports that $L_{0}$-FGL shows an outstanding performance, especially in the high dimensional case.

math.ST

Exact mediation analysis for ordinal outcome and binary mediator

With reference to a single mediator context, this brief report presents a model-based strategy to estimate counterfactual direct and indirect effects when the response variable is ordinal and the mediator is binary. Postulating a logistic regression model for the mediator and a cumulative logit model for the outcome, the exact parametric formulation of the causal effects is presented, thereby extending previous work that only contained approximated results. The identification conditions are equivalent to the ones already established in the literature. The effects can be estimated by making use of standard statistical software and standard errors can be computed via a bootstrap algorithm. To make the methodology accessible, routines to implement the proposal in R are presented in the Appendix. A natural effect model coherent with the postulated data generating mechanism is also derived.

stat.ME

A Metropolized adaptive subspace algorithm for high-dimensional Bayesian variable selection

A simple and efficient adaptive Markov Chain Monte Carlo (MCMC) method, called the Metropolized Adaptive Subspace (MAdaSub) algorithm, is proposed for sampling from high-dimensional posterior model distributions in Bayesian variable selection. The MAdaSub algorithm is based on an independent Metropolis-Hastings sampler, where the individual proposal probabilities of the explanatory variables are updated after each iteration using a form of Bayesian adaptive learning, in a way that they finally converge to the respective covariates' posterior inclusion probabilities. We prove the ergodicity of the algorithm and present a parallel version of MAdaSub with an adaptation scheme for the proposal probabilities based on the combination of information from multiple chains. The effectiveness of the algorithm is demonstrated via various simulated and real data examples, including a high-dimensional problem with more than 20,000 covariates.

stat.ME

High-dimensional variable selection via low-dimensional adaptive learning

A stochastic search method, the so-called Adaptive Subspace (AdaSub) method, is proposed for variable selection in high-dimensional linear regression models. The method aims at finding the best model with respect to a certain model selection criterion and is based on the idea of adaptively solving low-dimensional sub-problems in order to provide a solution to the original high-dimensional problem. Any of the usual $\ell_0$-type model selection criteria can be used, such as Akaike's Information Criterion (AIC), the Bayesian Information Criterion (BIC) or the Extended BIC (EBIC), with the last being particularly suitable for high-dimensional cases. The limiting properties of the new algorithm are analysed and it is shown that, under certain conditions, AdaSub converges to the best model according to the considered criterion. In a simulation study, the performance of AdaSub is investigated in comparison to alternative methods. The effectiveness of the proposed method is illustrated via various simulated datasets and a high-dimensional real data example.

stat.CO

An extended class of RC association models: estimation and main properties

The extended class of multiplicative row-column (RC) association models, introduced in this paper for two-way contingency tables, allows users to select both the type of logit (local, global, continuation, reverse continuation) suitable for the row and column classification variables and the scale on which interactions are measured. As in \cite{Kateri95} for the case of local logits, our extended class of bivariate interactions is linked to divergence measures and, by means of a representation theorem, we provide reconstruction formulas for the joint probabilities depending on pairs of logit types. These results are the key to show that, given marginal logits, our extended interactions determine uniquely the bivariate distribution. We also determine the kind of positive association which is implied by our extended interactions being non negative. Quick model selection within this wide class can be performed by an efficient algorithm for computing maximum likelihood estimates which exploits the properties of a reduced rank constraint imposed on the matrix of extended interactions and allows for additional linear constraint on marginal logits. An application to social mobility data is presented and discussed.

stat.CO

Modeling long-term capacity degradation of lithium-ion batteries

Capacity degradation of lithium-ion batteries under long-term cyclic aging is modelled via a flexible sigmoidal-type regression set-up, where the regression parameters can be interpreted. Different approaches known from the literature are discussed and compared with the new proposal. Statistical procedures, such as parameter estimation, confidence and prediction intervals are presented and applied to real data. The long-term capacity degradation model may be applied in second-life scenarios of batteries. Using some prior information or training data on the complete degradation path, the model can be fitted satisfactorily even if only short-term degradation data is available. The training data may arise from a single battery.

stat.AP

A Family of Quasisymmetry Models

We present a one-parameter family of models for square contingency tables that interpolates between the classical quasisymmetry model and its Pearsonian analogue. Algebraically, this corresponds to deformations of toric ideals associated with graphs. Our discussion of the statistical issues centers around maximum likelihood estimation.

math.ST