SearcharxivSearch

arXiv subjects

Anestis Touloumis

Publications and source records attributed to Anestis Touloumis.

6 recordsLinked to original sources

Bias-Reduced GEE via Adjusted Estimating Equations, with Odds-Ratio Extensions

Generalized estimating equations (GEE) are widely used for correlated data, but with small to moderate numbers of independent clusters the ordinary GEE regression estimators can be substantially biased. We develop a first-order bias-reduction principle for GEE by viewing the estimator as a clustered-data $M$-estimator and deriving an adjustment to the estimating equations that targets the leading bias term while accounting for the dependence of the working covariance on the mean parameters. The resulting class includes three bias-reduced estimators and three one-step bias-corrected analogs, nesting the bias-corrected estimator of Lunardon and Scharfstein (2017) and the bias-reduced and bias-corrected estimators of Paul and Zhang (2014) as special cases. The framework applies to general response types through correlation-coefficient parameterizations for the association structure and extends to correlated binary data through pairwise odds-ratio parameterizations, yielding the first bias-reduced and bias-corrected GEE estimators under this parameterization, for which the marginal-mean compatibility constraints are far less restrictive than those of correlation-coefficient parameterizations, making them better suited for small-sample settings. Under standard regularity conditions, all six estimators share the same asymptotic distribution as the ordinary GEE. Simulation studies show that the proposed estimators reduce bias while maintaining efficiency and coverage close to those of ordinary GEE across a range of settings, and a clinical trial analysis illustrates the proposed estimators in practice. Software is available in the R package geer.

stat.ME

Jeffreys-Type Penalized GEE for Correlated Binary Data with an Odds-Ratio Parameterization

Generalized estimating equations (GEE) are widely used for population-averaged inference on correlated binary responses, but ordinary GEE can fail under separation, a situation that is more likely in small-sample, sparse, or rare-event settings, leading to nonconvergence, infinite or extreme estimates, and unreliable inference. Existing penalized GEE (PGEE) approaches mitigate some of these problems but do not generally guarantee finite estimates under nonindependence working structures and often rely on correlation-coefficient parameterizations whose admissible range shrinks as fitted probabilities approach zero or one, forcing the working association toward independence under separation. We propose a PGEE framework that combines a Jeffreys-prior penalty with marginalized odds-ratio working parameterizations. The odds-ratio parameterization avoids this failure, while the penalty, with tunable strength $\delta$ and default $\delta = 1/2$, stabilizes estimation under separation. Under working independence, PGEE reduces to the Jeffreys-prior penalized maximum-likelihood estimator, yielding finite estimates for logit, probit, complementary log-log, and cauchit links. Under nonindependence odds-ratio structures, where a formal finiteness guarantee is unavailable, PGEE achieves near-complete empirical convergence even in separated settings. We also propose one-step and hybrid variants, OPGEE and HPGEE, that reduce computational cost. Simulations show that all three variants substantially outperform ordinary GEE under separation while retaining the performance of ordinary GEE in regular settings. We illustrate the method using a respiratory-illness trial in which ordinary GEE fails, and provide an implementation in the R package geer.

stat.ME

Testing the Mean Matrix in High-Dimensional Transposable Data

The structural information in high-dimensional transposable data allows us to write the data recorded for each subject in a matrix such that both the rows and the columns correspond to variables of interest. One important problem is to test the null hypothesis that the mean matrix has a particular structure without ignoring the potential dependence structure among and/or between the row and column variables. To address this, we develop a simple and computationally efficient nonparametric testing procedure to assess the hypothesis that, in each predefined subset of columns (rows), the column (row) mean vector remains constant. In simulation studies, the proposed testing procedure seems to have good performance and unlike traditional approaches, it is powerful without leading to inflated nominal sizes. Finally, we illustrate the use of the proposed methodology via two empirical examples from gene expression microarrays.

stat.ME

Hypothesis Testing for the Covariance Matrix in High-Dimensional Transposable Data with Kronecker Product Dependence Structure

The matrix-variate normal distribution is a popular model for high-dimensional transposable data because it decomposes the dependence structure of the random matrix into the Kronecker product of two covariance matrices: one for each of the row and column variables. We develop tests for assessing the form of the row (column) covariance matrix in high-dimensional settings while treating the column (row) dependence structure as a nuisance. Our tests are robust to normality departures provided that the Kronecker product dependence structure holds. In simulations, we observe that the proposed tests maintain the nominal level and are powerful against the alternative hypotheses tested. We illustrate the utility of our approach by examining whether genes associated with a given signalling network show correlated patterns of expression in different tissues and by studying correlation patterns within measurements of brain activity collected using electroencephalography.

stat.ME

R Package multgee: A Generalized Estimating Equations Solver for Multinomial Responses

The R package multgee implements the local odds ratios generalized estimating equations (GEE) approach proposed by Touloumis et al. (2013), a GEE approach for correlated multinomial responses that circumvents theoretical and practical limitations of the GEE method. A main strength of multgee is that it provides GEE routines for both ordinal (ordLORgee) and nominal (nomLORgee) responses, while relevant softwares in R and SAS are restricted to ordinal responses under a marginal cumulative link model specification. In addition, multgee offers a marginal adjacent categories logit model for ordinal responses and a marginal baseline category logit model for nominal. Further, utility functions are available to ease the local odds ratios structure selection (intrinsic.pars) and to perform a Wald type goodness-of-fit test between two nested GEE models (waldts). We demonstrate the application of multgee through a clinical trial with clustered ordinal multinomial responses.

stat.CO

Nonparametric Stein-type Shrinkage Covariance Matrix Estimators in High-Dimensional Settings

Estimating a covariance matrix is an important task in applications where the number of variables is larger than the number of observations. Shrinkage approaches for estimating a high-dimensional covariance matrix are often employed to circumvent the limitations of the sample covariance matrix. A new family of nonparametric Stein-type shrinkage covariance estimators is proposed whose members are written as a convex linear combination of the sample covariance matrix and of a predefined invertible target matrix. Under the Frobenius norm criterion, the optimal shrinkage intensity that defines the best convex linear combination depends on the unobserved covariance matrix and it must be estimated from the data. A simple but effective estimation process that produces nonparametric and consistent estimators of the optimal shrinkage intensity for three popular target matrices is introduced. In simulations, the proposed Stein-type shrinkage covariance matrix estimator based on a scaled identity matrix appeared to be up to 80% more efficient than existing ones in extreme high-dimensional settings. A colon cancer dataset was analyzed to demonstrate the utility of the proposed estimators. A rule of thumb for adhoc selection among the three commonly used target matrices is recommended.

stat.ME