SearcharxivSearch

arXiv subjects

Thierry Duchesne

Publications and source records attributed to Thierry Duchesne.

9 recordsLinked to original sources

A Leave-One-Out Influence Statistic for Density-Based Outlier Detection

We propose a density-based leave-one-out influence score for unsupervised outlier detection. The motivation is that outliers are naturally associated with regions of very small probability density, but direct leave-one-out density refitting can be computationally prohibitive. We use the Linear-Blend Frequency Polygon (LBFP) estimator and define a score that compares the full-sample fitted density at an observation with the fitted density obtained after removing that observation, while keeping the grid and bandwidth fixed. The resulting statistic measures a relative density perturbation at the observation's own location. For the LBFP estimator, this score has an exact closed-form update, so the density estimator does not need to be refitted for each observation. This preserves a direct density interpretation while making the method computationally efficient for large samples. We study the score under contamination and show that regular positive-density observations and contamination-driven observations have distinct asymptotic orders. Simulations over a broad range of contamination models illustrate these theoretical regimes, show competitive performance relative to standard benchmarks, and document computing time. A credit-card fraud application with 29 variables illustrates that the method works well on a large real data set.

stat.ME

Mixtures of Transparent Local Models

The predominance of machine learning models in many spheres of human activity has led to a growing demand for their transparency. The transparency of models makes it possible to discern some factors, such as security or non-discrimination. In this paper, we propose a mixture of transparent local models as an alternative solution for designing interpretable (or transparent) models. Our approach is designed for the situations where a simple and transparent function is suitable for modeling the label of instances in some localities/regions of the input space, but may change abruptly as we move from one locality to another. Consequently, the proposed algorithm is to learn both the transparent labeling function and the locality of the input space where the labeling function achieves a small risk in its assigned locality. By using a new multi-predictor (and multi-locality) loss function, we established rigorous PAC-Bayesian risk bounds for the case of binary linear classification problem and that of linear regression. In both cases, synthetic data sets were used to illustrate how the learning algorithms work. The results obtained from real data sets highlight the competitiveness of our approach compared to other existing methods as well as certain opaque models. Keywords: PAC-Bayes, risk bounds, local models, transparent models, mixtures of local transparent models.

cs.LG

Hypothesis tests for structured rank correlation matrices

Joint modeling of a large number of variables often requires dimension reduction strategies that lead to structural assumptions of the underlying correlation matrix, such as equal pair-wise correlations within subsets of variables. The underlying correlation matrix is thus of interest for both model specification and model validation. In this paper, we develop tests of the hypothesis that the entries of the Kendall rank correlation matrix are linear combinations of a smaller number of parameters. The asymptotic behavior of the proposed test statistics is investigated both when the dimension is fixed and when it grows with the sample size. We pay special attention to the restricted hypothesis of partial exchangeability, which contains full exchangeability as a special case. We show that under partial exchangeability, the test statistics and their large-sample distributions simplify, which leads to computational advantages and better performance of the tests. We propose various scalable numerical strategies for implementation of the proposed procedures, investigate their behavior through simulations and power calculations under local alternatives, and demonstrate their use on a real dataset of mean sea levels at various geographical locations.

stat.ME

Fully Bayesian inference for spatiotemporal data with the multi-resolution approximation

Large spatiotemporal datasets are a challenge for conventional Bayesian models because of the cubic computational complexity of the algorithms for obtaining the Cholesky decomposition of the covariance matrix in the multivariate normal density. Moreover, standard numerical algorithms for posterior estimation, such as Markov Chain Monte Carlo (MCMC), are intractable in this context, as they require thousands, if not millions, of costly likelihood evaluations. To overcome those limitations, we propose IS-MRA (Importance sampling - Multi-Resolution Approximation), which takes advantage of the sparse inverse covariance structure produced by the Multi-Resolution Approximation (MRA) approach. IS-MRA is fully Bayesian and facilitates the approximation of the hyperparameter marginal posterior distributions. We apply IS-MRA to large MODIS Level 3 Land Surface Temperature (LST) datasets, sampled between May 18 and May 31, 2012 in the western part of the state of Maharashtra, India. We find that IS-MRA can produce realistic prediction surfaces over regions where concentrated missingness, caused by sizable cloud cover, is observed. Through a validation analysis and simulation study, we also find that predictions tend to be very accurate.

stat.CO

Detection of Block-Exchangeable Structure in Large-Scale Correlation Matrices

Correlation matrices are omnipresent in multivariate data analysis. When the number d of variables is large, the sample estimates of correlation matrices are typically noisy and conceal underlying dependence patterns. We consider the case when the variables can be grouped into K clusters with exchangeable dependence; this assumption is often made in applications, e.g., in finance and econometrics. Under this partial exchangeability condition, the corresponding correlation matrix has a block structure and the number of unknown parameters is reduced from d(d-1)/2 to at most K(K+1)/2. We propose a robust algorithm based on Kendall's rank correlation to identify the clusters without assuming the knowledge of K a priori or anything about the margins except continuity. The corresponding block-structured estimator performs considerably better than the sample Kendall rank correlation matrix when K < d. The new estimator can also be much more efficient in finite samples even in the unstructured case K = d, although there is no gain asymptotically. When the distribution of the data is elliptical, the results extend to linear correlation matrices and their inverses. The procedure is illustrated on financial stock returns.

math.ST

A scalable and efficient covariate selection criterion for mixed effects regression models with unknown random effects structure

We propose a new model selection criterion for mixed effects regression models that is computable when the model is fitted with a two-step method, even when the structure and the distribution of the random effects are unknown. The criterion is especially useful in the early stage of the model building process when one needs to decide which covariates should be included in a mixed effects regression model but has no knowledge of the random effect structure. This is particularly relevant in substantive fields where variable selection is guided by information criteria rather than regularization. The calculation of the criterion requires only the evaluation of cluster-level log-likelihoods and does not rely on heavy numerical integration. We provide theoretical and numerical arguments to justify the method and we illustrate its usefulness by analyzing data on a socio-economic study of young American Indians.

stat.ME

Prediction of the margin of victory only from team rankings for regular season games in NCAA men's basketball

The main objective of this paper is to investigate the extent to which the margin of victory can be predicted solely by the rankings of the opposing teams in NCAA Division I men's basketball games. Several past studies have modeled this relationship for the games played during the March Madness tournament, and this work aims at verifying if the models advocated in these papers still perform well for regular season games. Indeed, most previous articles have shown that a simple quadratic regression model provides fairly accurate predictions of the margin of victory when team rankings only range from 1 to 16. Does that still hold true when team rankings can go as high as 351? Do the model assumptions hold? Can we find semi- or non-parametric methods that yield even better results (i.e. predicted margins of victory that more closely resemble actual results)? The analyses presented in this paper suggest that the answer is "yes" on all three counts!

stat.AP

A Multi-State Conditional Logistic Regression Model for the Analysis of Animal Movement

A multi-state version of an animal movement analysis method based on conditional logistic regression, called Step Selection Function (SSF), is proposed. In ecology SSF is developed from a comparison between the observed location of an animal and randomly sampled locations at each time step. Interpretation of the parameters in the multi-state model and the impact of different sampling schemes for the random locations are discussed. We prove the equivalence between the new model and a random walk model on the plane. This equivalence allows one to use both pure movement and local discrete choice behaviors in identifying the model's hidden states. The new method is used to model the movement behavior of GPS-collared bison in Prince Albert National Park, Canada. The multi-state SSF successfully teases apart areas used to forage and to travel. The analysis thus provides valuable insights into how bison adjust their movement to habitat features, thereby revealing spatial determinants of functional connectivity in heterogeneous landscapes.

stat.ME

A General Hidden State Random Walk Model for Animal Movement

In this paper, we propose a general hidden state random walk model to describe the movement of an animal that takes into account movement taxis with respect to features of the environment. A circular-linear process models the direction and distance between two consecutive localizations of the animal. A hidden process structure accounts for the animal's change in movement behavior. The originality of the proposed approach is that several environmental targets can be included in the directional model. An EM algorithm is devised to fit this model and an application to the analysis of the movement of caribou in Canada's boreal forest is presented

stat.ME