SearcharxivSearch

arXiv subjects

Jean Peyhardi

Publications and source records attributed to Jean Peyhardi.

8 recordsLinked to original sources

Zero-inflated binary Tree P\'olya splitting regression for multivariate count data

Species distribution models (SDMs) are widely used to assess the effects of environmental factors on species distributions. However, classical SDMs ignore inter-species dependencies. Multivariate SDMs (MSDMs), especially those based on latent Gaussian fields such as the multivariate Poisson log-normal (MPLN), address this limitation but face challenges related to computation, dimensionality, and interpretability. P\'olya-splitting (PS) distributions offer an alternative, combining a model for total abundance with a multivariate allocation structure, and have natural interpretations from ecological process models. Yet, they lack flexibility in modeling correlation structures. Tree P\'olya-splitting (TPS) distributions overcome this by introducing hierarchical structure such as a phylogenetic tree. In this paper, we extend TPS to account for zero-inflation, leading to the zero-inflated tree P\'olya-splitting (Z-TPS) family. We detail its statistical properties, show how standard software enables efficient inference, and illustrate its ecological relevance using tree abundance data from over 180 genera across the Congo Basin tropical rainforest.

stat.AP

Tree P\'olya Splitting distributions for multivariate count data

In this article, we develop a new class of multivariate distributions adapted for count data, called Tree P\'olya Splitting. This class results from the combination of a univariate distribution and singular multivariate distributions along a fixed partition tree. Known distributions, including the Dirichlet-multinomial, the generalized Dirichlet-multinomial and the Dirichlet-tree multinomial, are particular cases within this class. As we will demonstrate, these distributions are flexible, allowing for the modeling of complex dependence structures (positive, negative, or null) at the observation level. Specifically, we present the theoretical properties of Tree P\'olya Splitting distributions by focusing primarily on marginal distributions, factorial moments, and dependence structures (covariance and correlations). A dataset of abundance of Trichoptera is used, on one hand, as a benchmark to illustrate the theoretical properties developed in this article, and on the other hand, to demonstrate the interest of these types of models, notably by comparing them to other approaches for fitting multivariate data, such as the Poisson-lognormal model in ecology or singular multivariate distributions used in microbiome.

math.ST

Asymptotic tail properties of Poisson mixture distributions

Count data are omnipresent in many applied fields, often with overdispersion. With mixtures of Poisson distributions representing an elegant and appealing modelling strategy, we focus here on how the tail behaviour of the mixing distribution is related to the tail of the resulting Poisson mixture. We define five sets of mixing distributions and we identify for each case whenever the Poisson mixture is in, close to or far from a domain of attraction of maxima. We also characterize how the Poisson mixture behaves similarly to a standard Poisson distribution when the mixing distribution has a finite support. Finally, we study, both analytically and numerically, how goodness-of-fit can be assessed with the inspection of tail behaviour.

math.ST

Choice of mixture Poisson models based on Extreme value theory

Count data are omnipresent in many applied fields, often with overdispersion due to an excess of zeroes or extreme values. With mixtures of Poisson distributions representing an elegant and appealing modelling strategy, we focus here on the challenging problem of identifying a suitable mixing distribution and study how extreme value theory can be used. We propose an original strategy to select the most appropriate candidate among three categories: Fr{é}chet, Gumbel and pseudo-Gumbel. Such an approach is presented with the aid of a decision tree and evaluated with numerical simulations.

math.ST

Splitting models for multivariate count data

Considering discrete models, the univariate framework has been studied in depth compared to the multivariate one. This paper first proposes two criteria to define a sensu stricto multivariate discrete distribution. It then introduces the class of splitting distributions that encompasses all usual multivariate discrete distributions (multinomial, negative multinomial, multivariate hypergeometric, multivariate neg- ative hypergeometric, etc . . . ) and contains several new. Many advantages derive from the compound aspect of split- ting distributions. It simplifies the study of their characteris- tics, inferences, interpretations and extensions to regression models. Moreover, splitting models can be estimated only by combining existing methods, as illustrated on three datasets with reproducible studies.

math.ST

Item response models for the longitudinal analysis of health-related quality of life in cancer clinical trials

Statistical research regarding health-related quality of life (HRQoL) is a major challenge to better evaluate the impact of the treatments on their everyday life and to improve patients' care. Among the models that are used for the longitudinal analysis of HRQoL, we focused on the mixed models from the item response theory to analyze directly the raw data from questionnaires. Using a recent classification of generalized linear models for categorical data, we discussed about a conceptual selection of these models for the longitudinal analysis of HRQoL in cancer clinical trials. Through methodological and practical arguments, the adjacent and cumulative models seem particularly suitable for this {context}. Specially in cancer clinical trials and for the comparison between two groups, the cumulative models has the advantage of providing intuitive illustrations of results. To complete the comparison studies already performed in literature, a simulation study based on random part of the mixed models is then carried out to compare the linear mixed model classically used to the discussed item response models. As expected, the sensitivity of item response models to detect random effect with lower variance is better than the linear mixed model sensitivity. Finally, a longitudinal analysis of HRQoL data from cancer clinical trial is carried out using an item response cumulative model.

stat.AP

Partitioned conditional generalized linear models for categorical data

In categorical data analysis, several regression models have been proposed for hierarchically-structured response variables, e.g. the nested logit model. But they have been formally defined for only two or three levels in the hierarchy. Here, we introduce the class of partitioned conditional generalized linear models (PCGLMs) defined for any numbers of levels. The hierarchical structure of these models is fully specified by a partition tree of categories. Using the genericity of the (r,F,Z) specification, the PCGLM can handle nominal, ordinal but also partially-ordered response variables.

stat.ME

A new specification of generalized linear models for categorical data

Regression models for categorical data are specified in heterogeneous ways. We propose to unify the specification of such models. This allows us to define the family of reference models for nominal data. We introduce the notion of reversible models for ordinal data that distinguishes adjacent and cumulative models from sequential ones. The combination of the proposed specification with the definition of reference and reversible models and various invariance properties leads to a new view of regression models for categorical data.

stat.ME