SearcharxivSearch

arXiv subjects

Cristiano Varin

Publications and source records attributed to Cristiano Varin.

8 recordsLinked to original sources

ML, PL, QL in Markov chain models

In many spatial and spatial-temporal models, and more generally in models with complex dependencies, it may be too difficult to carry out full maximum likelihood (ML) analysis. Remedies include the use of pseudo-likelihood (PL) and quasi-likelihood (QL) (also called the composite likelihood). The present article studies the ML, the PL and the QL methods for general Markov chain models, partly motivated by the desire to understand the precise behaviour of PL and QL methods in settings where this can be analysed. We present limiting normality results and compare performances in different settings. The PL and QL methods can be seen as maximum penalised likelihood methods. We find that the QL strategy is typically preferable to the PL, and that it loses very little to the ML, while earning in model robustness. It has also appeal and potential as a modelling tool. Our methods are illustrated for analysis of DNA sequence evolution type models.

stat.ME

Consistent and Scalable Composite Likelihood Estimation of Probit Models with Crossed Random Effects

Estimation of crossed random effects models commonly requires computational costs that grow faster than linearly in the sample size $N$, often as fast as $Ω(N^{3/2})$, making them unsuitable for large data sets. For non-Gaussian responses, integrating out the random effects to get a marginal likelihood brings significant challenges, especially for high dimensional integrals where the Laplace approximation might not be accurate. We develop a composite likelihood approach to probit models that replaces the crossed random effects model with some hierarchical models that require only one-dimensional integrals. We show how to consistently estimate the crossed effects model parameters from the hierarchical model fits. We find that the computation scales linearly in the sample size. We illustrate the method on about five million observations from Stitch Fix where the crossed effects formulation would require an integral of dimension larger than $700{,}000$.

stat.ME

Tractable Ridge Regression for Paired Comparisons

Paired comparison models, such as Bradley-Terry and Thurstone-Mosteller, are commonly used to estimate relative strengths of pairwise compared items in tournament-style data. We discuss estimation of paired comparison models with a ridge penalty. A new approach is derived which combines empirical Bayes and composite likelihoods without any need to re-fit the model, as a convenient alternative to cross-validation of the ridge tuning parameter. Simulation studies demonstrate much better predictive accuracy of the new approach relative to ordinary maximum likelihood. A widely used alternative, the application of a standard bias-reducing penalty, is also found to improve appreciably the performance of maximum likelihood; but the ridge penalty, with tuning as developed here, yields greater accuracy still. The methodology is illustrated through application to 28 seasons of English Premier League football.

stat.ME

Pairwise likelihood estimation of latent autoregressive count models

Latent autoregressive models are useful time series models for the analysis of infectious disease data. Evaluation of the likelihood function of latent autoregressive models is intractable and its approximation through simulation-based methods appears as a standard practice. Although simulation methods may make the inferential problem feasible, they are often computationally intensive and the quality of the numerical approximation may be difficult to assess. We consider instead a weighted pairwise likelihood approach and explore several computational and methodological aspects including estimation of robust standard errors and the role of numerical integration. The suggested approach is illustrated using monthly data on invasive meningococcal disease infection in Greece and Italy.

stat.ME

Improving the accuracy of likelihood-based inference in meta-analysis and meta-regression

Random-effects models are frequently used to synthesise information from different studies in meta-analysis. While likelihood-based inference is attractive both in terms of limiting properties and of implementation, its application in random-effects meta-analysis may result in misleading conclusions, especially when the number of studies is small to moderate. The current paper shows how methodology that reduces the asymptotic bias of the maximum likelihood estimator of the variance component can also substantially improve inference about the mean effect size. The results are derived for the more general framework of random-effects meta-regression, which allows the mean effect size to vary with study-specific covariates.

stat.ME

Statistical Modelling of Citation Exchange Between Statistics Journals

Rankings of scholarly journals based on citation data are often met with skepticism by the scientific community. Part of the skepticism is due to disparity between the common perception of journals' prestige and their ranking based on citation counts. A more serious concern is the inappropriate use of journal rankings to evaluate the scientific influence of authors. This paper focuses on analysis of the table of cross-citations among a selection of Statistics journals. Data are collected from the Web of Science database published by Thomson Reuters. Our results suggest that modelling the exchange of citations between journals is useful to highlight the most prestigious journals, but also that journal citation data are characterized by considerable heterogeneity, which needs to be properly summarized. Inferential conclusions require care in order to avoid potential over-interpretation of insignificant differences between journal ratings. Comparison with published ratings of institutions from the UK's Research Assessment Exercise shows strong correlation at aggregate level between assessed research quality and journal citation `export scores' within the discipline of Statistics.

stat.AP

Beta regression for time series analysis of bounded data, with application to Canada Google${}^\circledR$ Flu Trends

Bounded time series consisting of rates or proportions are often encountered in applications. This manuscript proposes a practical approach to analyze bounded time series, through a beta regression model. The method allows the direct interpretation of the regression parameters on the original response scale, while properly accounting for the heteroskedasticity typical of bounded variables. The serial dependence is modeled by a Gaussian copula, with a correlation matrix corresponding to a stationary autoregressive and moving average process. It is shown that inference, prediction, and control can be carried out straightforwardly, with minor modifications to standard analysis of autoregressive and moving average models. The methodology is motivated by an application to the influenza-like-illness incidence estimated by the Google${}^\circledR$ Flu Trends project.

stat.AP

The ranking lasso and its application to sport tournaments

Ranking a vector of alternatives on the basis of a series of paired comparisons is a relevant topic in many instances. A popular example is ranking contestants in sport tournaments. To this purpose, paired comparison models such as the Bradley-Terry model are often used. This paper suggests fitting paired comparison models with a lasso-type procedure that forces contestants with similar abilities to be classified into the same group. Benefits of the proposed method are easier interpretation of rankings and a significant improvement of the quality of predictions with respect to the standard maximum likelihood fitting. Numerical aspects of the proposed method are discussed in detail. The methodology is illustrated through ranking of the teams of the National Football League 2010-2011 and the American College Hockey Men's Division I 2009-2010.

stat.AP