SearcharxivSearch

arXiv subjects

Fabio Rapallo

Publications and source records attributed to Fabio Rapallo.

At least 19 recordsLinked to original sources

The Fisher score on the closed simplex

We extend classical analytic tools for finite-state statistical models to allow zero probabilities. Using methods from algebraic statistics and information geometry, we develop a framework in which a smooth statistical model could hit the boundary of the simplex, for example, in contingency tables with non-structural zeros. The central object of our approach is the vector bundle whose fibres are the $p$-contrasts associated to each probability distribution $p$. In this framework, Fisher score and other key statistical concepts, such as entropy for one-dimensional statistical models, admit an algebraic representation also on the boundary of the simplex.

math.ST

Characterization of multi-way binary tables with uniform margins and fixed correlations

In many applications involving binary variables, only pairwise dependence measures, such as correlations, are available. However, for multi-way tables involving more than two variables, these quantities do not uniquely determine the joint distribution, but instead define a family of admissible distributions that share the same pairwise dependence while potentially differing in higher-order interactions. In this paper, we introduce a geometric framework to describe the entire feasible set of such joint distributions with uniform margins. We show that this admissible set forms a convex polytope, analyze its symmetry properties, and characterize its extreme rays. These extremal distributions provide fundamental insights into how higher-order dependence structures may vary while preserving the prescribed pairwise information. Unlike traditional methods for table generation, which return a single table, our framework makes it possible to explore and understand the full admissible space of dependence structures, enabling more flexible choices for modeling and simulation. We illustrate the usefulness of our theoretical results through examples and a real case study on rater agreement.

stat.ME

How many outbreaks before an epidemic?

In this work, we study the finite-population behaviour of the Reed-Frost epidemic model. Our analysis relies on the exact expression for the final epidemic size, replaced by Monte Carlo simulations in cases where the exact formula becomes numerically unstable. When the initial reproduction number is greater than a critical threshold, the distribution of the final size becomes bimodal. We therefore define the probabilities of small and large outbreaks, providing an intuitive answer to the question posed in the title through simple arguments based on the geometric distribution. Finally, an agent-based simulation confirms that the Reed-Frost model offers a good approximation in the case of the COVID-19 outbreak.

q-bio.PE

Testing for a common subspace in compositional datasets with structural zeros

In this paper, we consider the problem of testing for a common principal subspace in compositional datasets with structural zeros. In particular, we address the problem in the general setting in which two groups are compared, each characterized by its own pattern of structural zeros. This situation prevents the direct use of standard logratio-based principal subspace comparison methods, since the two groups cannot be represented in a common logratio coordinate system in the usual way. We thus define a test for the presence of a common subspace that is fully compatible with the Aitchison geometry and logratio analysis. The proposed construction allows the principal subspaces associated with the two groups to be compared despite the different patterns of structural zeros. Under logratio-normality, we derive an analytical approximation to the null distribution of the test statistic. Alongside this parametric version, we also introduce a nonparametric bootstrap procedure that does not rely on distributional assumptions. The finite-sample behaviour of the test is investigated through simulations. The method is finally illustrated on a reproducible microbiome dataset available through Bioconductor.

stat.ME

Zero patterns in multi-way binary contingency tables with uniform margins

We study the problem of transforming a multi-way contingency table into an equivalent table with uniform margins and same dependence structure. This is an old question which relates to recent advances in copula modeling for discrete random vectors. In this work, we focus on multi-way binary tables and develop novel theory to show how the zero patterns affect the existence of the transformation as well as its statistical interpretability in terms of dependence structure. The implementation of the theory relies on combinatorial and linear programming techniques, which can also be applied to arbitrary multi-way tables. In addition, we investigate which odds ratios characterize the unique solution in relation to specific zero patterns. Several examples are described to illustrate the approach and point to interesting future research directions.

math.ST

Multi-way contingency tables with uniform margins

We study the problem of transforming a multi-way contingency table into an equivalent table with uniform margins and same dependence structure. Such a problem relates to recent developments in copula modeling for discrete random vectors. Here, we focus on three-way binary tables and show that, even in such a simple case, the situation is quite different than for two-way tables. Many more constraints are needed to ensure a unique solution to the problem. Therefore, the uniqueness of the transformed table is subject to arbitrary choices of the practitioner. We illustrate the theory through some examples, and conclude with a discussion on the topic and future research directions.

stat.ME

Robustness against data loss with Algebraic Statistics

The paper describes an algorithm that, given an initial design $\mathcal{F}_n$ of size $n$ and a linear model with $p$ parameters, provides a sequence $\mathcal{F}_n \supset \ldots \supset \mathcal{F}_{n-k} \supset \ldots \supset \mathcal{F}_p$ of nested \emph{robust} designs. The sequence is obtained by the removal, one by one, of the runs of $\mathcal{F}_n$ till a $p$-run \emph{saturated} design $\mathcal{F}_p$ is obtained. The potential impact of the algorithm on real applications is high. The initial fraction $\mathcal{F}_n$ can be of any type and the output sequence can be used to organize the experimental activity. The experiments can start with the runs corresponding to $\mathcal{F}_p$ and continue adding one run after the other (from $\mathcal{F}_{n-k}$ to $\mathcal{F}_{n-k+1}$) till the initial design $\mathcal{F}_n$ is obtained. In this way, if for some unexpected reasons the experimental activity must be stopped before the end when only $n-k$ runs are completed, the corresponding $\mathcal{F}_{n-k}$ has a high value of robustness for $k \in \{1, \ldots, n-p\}$. The algorithm uses the circuit basis, a special representation of the kernel of a matrix with integer entries. The effectiveness of the algorithm is demonstrated through the use of simulations.

stat.CO

Circuits for robust designs

This paper continues the application of circuit theory to experimental design started by the first two authors. The theory gives a very special and detailed representation of the kernel of the design model matrix. This representation turns out to be an appropriate way to study the optimality criteria referred to as robustness: the sensitivity of the design to the removal of design points. Many examples are given, from classical combinatorial designs to two-level factorial design including interactions. The complexity of the circuit representations are useful because the large range of options they offer, but conversely require the use of dedicated software. Suggestions for speed improvement are made.

stat.CO

Circuit bases for randomisation

After a rich history in medicine, randomisation control trials both simple and complex are in increasing use in other areas such as web-based AB testing and planning and design decisions. A main objective is to be able to measure parameters, and contrasts in particular, while guarding against biases from hidden confounders. After careful definitions of classical entities such as contrasts, an algebraic method based on circuits is introduced which gives a wide choice of randomisation schemes.

math.ST

Finite space Kantorovich problem with an MCMC of table moves

In Optimal Transport (OT) on a finite metric space, one defines a distance on the probability simplex that extends the distance on the ground space. The distance is the value of a Linear Programming (LP) problem on the set of non-negative-valued 2-way tables with assigned probability functions as margins. We apply to this case the methodology of moves from Algebraic Statistics (AS) and use it to derive a Monte Carlo Markov Chain (MCMC) solution algorithm.

stat.ME

Analysis of the weighted kappa and its maximum with Markov moves

In this paper the notion of Markov move from Algebraic Statistics is used to analyze the weighted kappa indices in rater agreement problems. In particular, the problem of the maximum kappa and its dependence on the choice of the weighting schemes are discussed. The Markov moves are also used in a simulated annealing algorithm to actually find the configuration of maximum agreement.

stat.ME

Modal operators and toric ideals

In the present paper we consider modal propositional logic and look for the constraints that are imposed to the propositions of the special type $\Box a$ by the structure of the relevant finite Kripke frame. We translate the usual language of modal propositional logic in terms of notions of commutative algebra, namely polynomial rings, ideals, and bases of ideals. We use extensively the perspective obtained in previous works in Algebraic Statistics. We prove that the constraints on $\Box a$ can be derived through a binomial ideal containing a toric ideal and we give sufficient conditions under which the toric ideal fully describes the constraints.

math.LO

On the aberrations of mixed level Orthogonal Arrays with removed runs

Given an Orthogonal Array we analyze the aberrations of the sub-fractions which are obtained by the deletion of some of its points. We provide formulae to compute the Generalized Word-Length Pattern of any sub-fraction. In the case of the deletion of one single point, we provide a simple methodology to find which the best sub-fractions are according to the Generalized Minimum Aberration criterion. We also study the effect of the deletion of 1, 2 or 3 points on some examples. The methodology does not put any restriction on the number of levels of each factor. It follows that any mixed level Orthogonal Array can be considered.

math.ST

Algebra and geometry of tensors for modeling rater agreement data

We study three different quasi-symmetry models and three different mixture models of $n\times n\times n$ tensors for modeling rater agreement data. For these models we give a geometric description of the associated varieties and we study their invariants distinguishing between the case $n=2$ and the case $n>2$. Finally, for the two models for pairwise agreement we state some results about the pairwise Cohen's $κ$ coefficients.

math.ST

Unions of Orthogonal Arrays and their aberrations via Hilbert bases

We generate all the Orthogonal Arrays (OAs) of a given size n and strength t as the union of a collection of OAs which belong to an inclusion-minimal set of OAs. We derive a formula for computing the (Generalized) Word Length Pattern of a union of OAs that makes use of their polynomial counting functions. In this way the best OAs according to the Generalized Minimum Aberration criterion can be found by simply exploring a relatively small set of counting functions. The classes of OAs with 5 binary factors, strength 2, and sizes 16 and 20 are fully described.

math.ST

Algebraic characterization of regular fractions under level permutations

In this paper we study the behavior of the fractions of a factorial design under permutations of the factor levels. We focus on the notion of regular fraction and we introduce methods to check whether a given symmetric orthogonal array can or can not be transformed into a regular fraction by means of suitable permutations of the factor levels. The proposed techniques take advantage of the complex coding of the factor levels and of some tools from polynomial algebra. Several examples are described, mainly involving factors with five levels.

stat.ME

Exact tests to compare contingency tables under quasi-independence and quasi-symmetry

In this work we define log-linear models to compare several square contingency tables under the quasi-independence or the quasi-symmetry model, and the relevant Markov bases are theoretically characterized. Through Markov bases, an exact test to evaluate if two or more tables fit a common model is introduced. Two real-data examples illustrate the use of these models in different fields of applications.

math.ST

Detection of outlying proportions

In this paper we introduce a new method for detecting outliers in a set of proportions. It is based on the construction of a suitable two-way contingency table and on the application of an algorithm for the detection of outlying cells in such table. We exploit the special structure of the relevant contingency table to increase the efficiency of the method. The main properties of our algorithm, together with a guide for the choice of the parameters, are investigated through simulations, and in simple cases some theoretical justifications are provided. Several examples on synthetic data and an example based on pseudo-real data from biological experiments demonstrate the good performances of our algorithm.

stat.ME