SearcharxivSearch

arXiv subjects

Irene Crimaldi

Publications and source records attributed to Irene Crimaldi.

At least 19 recordsLinked to original sources

Design and inference for multi-arm clinical trials with informational borrowing: the interacting urns design

This paper deals with a new design methodology for stratified comparative experiments based on a system of interacting urns. The key idea is to model the interaction between urns for borrowing information across strata and to use it in the design phase in order to i) enhance the information exchange at the beginning of the study, when only few subjects have been enrolled and the stratum-specific information on treatments' efficacy could be scarce, ii) let the information sharing adaptively evolve via an update mechanism based on the observed outcomes, for skewing at each step the allocations towards the stratum-specific most promising treatment and iii) make the contribution of the strata with different treatment efficacy vanishing as the stratum information grows. In particular, we introduce the Interacting Urns Design, namely a new Covariate-Adjusted Response-Adaptive procedure, that randomizes the treatment allocations according to the evolution of the urn system. The theoretical properties of this proposal are described and the corresponding asymptotic inference is provided. Moreover, by a functional central limit theorem, we obtain the asymptotic joint distribution of the Wald-type sequential test statistics, which allows to sequentially monitor the suggested design in the clinical practice

stat.ME

A multi-factorial innovation model with mean-field interaction

We introduce an Indian-buffet-type model for multi-factorial innovation in which each arriving agent may exhibit both previously observed and new features. The number of new features follows a power-law behavior, while the probability of selecting an old feature combines self-reinforcement, depending on the feature-specific popularity, with a mean-field interaction term depending on the average popularity of all observed features. The model is governed by the usual innovation parameters (mass, discount and concentration), together with two additional parameters: one controlling the strength of reinforcement against a forcing input toward zero, and one regulating the intensity of the mean-field interaction. Although the growth of the total number of distinct observed features has the same behavior as in the three-parameter Indian buffet process, the mean-field interaction mechanism produces new asymptotic regimes. For aggregate quantities, including the predictive mean, the averaged number of features per agent, the mean inclusion probability, and the mean feature popularity, the phase transition is determined by the comparison between the discount parameter and the weight of the forcing input. For feature-specific quantities, a further transition appears according to the comparison between the interaction level and a critical threshold. In particular, high interaction leads to an asymptotic synchronization of feature-specific inclusion probabilities. We establish strong laws and second-order asymptotic results, including central limit theorems in regimes where martingale fluctuations compete with deterministic or random terms. The analysis relies on novel general results for recursive stochastic dynamics, which may be useful beyond the present framework.

math.ST

Granger Causality in Expectiles: an M-vine copula test

A model-free measure of Granger causality in expectiles is proposed, generalizing the traditional mean-based measure to arbitrary positions of the conditional distribution. Expectiles are the only law-invariant risk measures that are both coherent and elicitable, making them particularly well-suited for studying distributional Granger causality where risk quantification and forecast evaluation are both relevant. Based on this measure, a test is developed using M-vine copula models that accounts for multivariate Granger causality with $d+1$ series under non-linear and non-Gaussian dependence, without imposing parametric assumptions on the joint distribution. Strong consistency of the test statistic is established under some regularity conditions. In finite samples, simulations show accurate size control and power increasing with sample size. A key advantage is the joint testing capability: causal relationships invisible to pairwise tests can be detected, as demonstrated both theoretically and empirically. Two applications to international stock market indices at the global and Asian regional level illustrate the practical relevance of the proposed framework.

econ.EM

Triggered urn models for frequently asked questions (FAQ)

We investigate a nonclassic urn model with triggers that increase the number of colors. The scheme has emerged as a model for web services that set up frequently asked questions (FAQ). We present a thorough asymptotic analysis of the FAQ urn scheme in generality that covers a large number of special cases, such as Simon urn. For instance, we consider time dependent triggering probabilities. We identify regularity conditions on these probabilities that classify the schemes into those where the number of colors in the urn remains almost surely finite or increases to infinity and conditions that tell us whether all the existing colors are observed infinitely often or not. We determine the rank curve, too. In view of the broad generality of the trigger probabilities, a spectrum of limit distributions appears, from central limit theorems to Poisson approximation, to power-laws, revealing connections to Heap's exponent and Zipf's law. A combinatorial approach to the Simon urn is presented to indicate the possibility of such exact analysis, which is important for short-term predictions. Extensive simulations on real datasets (from Amazon sales) as well as computer-generated data clearly indicate that the asymptotic and exact theory developed agrees with practice.

math.PR

Modeling Innovation Ecosystem Dynamics through Interacting Reinforced Bernoulli Processes

Innovation is cumulative and interdependent: successful inventions build on prior knowledge within technological fields and may also affect success across related ones. Yet these dimensions are often studied separately in the innovation literature. This paper asks whether patent success across technological categories can be represented within a single dynamic framework that jointly captures within-category reinforcement, cross-category spillovers, and a set of aggregate regularities observed in patent data. To address this question, we propose a model of interacting reinforced Bernoulli processes in which the probability of success in a given category depends on past successes both within that category and across other categories. The framework yields joint predictions for success probabilities, cumulative successes, relative success shares, and cross-category dependence. We implement the model using granted US patent families from GLOBAL PATSTAT (1980-2018), defining category-specific success through a cohort-normalized forward-citation index. The empirical analysis shows that successful innovations continue to accumulate, but less than proportionally to the growth in patent opportunities, while technological categories remain interdependent without becoming homogeneous. Under a mean-field restriction, the model-based inferential exercise yields an estimated interaction intensity of 0.643, pointing to positive but non-maximal interaction across technological categories.

stat.AP

Non-linear dependence and Granger causality: A vine copula approach

Inspired by Jang et al. (2022), we propose a Granger causality-in-the-mean test for bivariate $k-$Markov stationary processes based on a recently introduced class of non-linear models, i.e., vine copula models. By means of a simulation study, we show that the proposed test improves on the statistical properties of the original test in Jang et al. (2022), and also of other previous methods, constituting an excellent tool for testing Granger causality in the presence of non-linear dependence structures. Finally, we apply our test to study the pairwise relationships between energy consumption, GDP and investment in the U.S. and, notably, we find that Granger-causality runs two ways between GDP and energy consumption.

econ.EM

Central limit theorems for interacting innovation processes, related statistical tools and general results

We study a networked system of innovation processes, where each process is modeled as an urn with infinitely many colors-a classical framework for capturing the emergence of novelties. Extending this paradigm, we analyze a model of interacting urns, where the probability of generating or reusing elements in one process is influenced by the histories of others. This interaction is governed by two matrices that control innovation triggering and reinforcement dynamics across the system. The core contribution of this work is a detailed analysis of the second-order asymptotic behavior of the model. Building on these theoretical results, we develop statistical tools to infer the structure and strength of inter-process influence. The methodology is framed in a general setting, making it broadly applicable. We validate our approach with applications to two real-world datasets from Reddit discussions and Gutenberg text corpora.

stat.ME

Networks of reinforced stochastic processes: probability of asymptotic polarization and related general results

In a network of reinforced stochastic processes, for certain values of the parameters, all the agents' inclinations synchronize and converge almost surely toward a certain random variable. The present work aims at clarifying when the agents can asymptotically polarize, i.e. when the common limit inclination can take the extreme values, 0 or 1, with probability zero, strictly positive, or equal to one. Moreover, we present a suitable technique to estimate this probability that, along with the theoretical results, has been framed in the more general setting of a class of martingales taking values in [0, 1] and following a specific dynamics.

math.PR

Interacting Innovation processes: case studies from Reddit and Gutenberg

In this work, we introduce an extremely general model for a collection of innovation processes in order to model and analyze the interaction among them. We provide theoretical results, analytically proven, and we show how the proposed model fits the behaviors observed in some real data sets (from Reddit and Gutenberg). It is worth mentioning that the given applications are only examples of the potentialities of the proposed model and related results: due to its abstractness and generality, it can be applied to many interacting innovation processes.

stat.AP

Networks of reinforced stochastic processes: a complete description of the first-order asymptotics

We consider a finite collection of reinforced stochastic processes with a general network-based interaction among them. We provide sufficient and necessary conditions in order to have some form of almost sure asymptotic synchronization, which could be roughly defined as the almost sure long-run uniformization of the behavior of interacting processes. Specifically, we detect a regime of complete synchronization, where all the processes converge toward the same random variable, a second regime where the system almost surely converges, but there exists no form of almost sure asymptotic synchronization, and another regime where the system does not converge with a strictly positive probability. In this latter case, partitioning the system in cyclic classes according to the period of the interaction matrix, we have an almost sure asymptotic synchronization within the cyclic classes, and, with a strictly positive probability, an asymptotic periodic behavior of these classes.

math.PR

Statistical test for an urn model with random multidrawing and random addition

We complete the study of the model introduced in [11]. It is a two-color urn model with multiple drawing and random (non-balanced) time-dependent reinforcement matrix. The number of sampled balls at each time-step is random. We identify the exact rates at which the number of balls of each color grows to infinity and define two strongly consistent estimators for the limiting reinforcement averages. Then we prove a Central Limit Theorem, which allows to design a statistical test for such averages.

math.ST

The Rescaled Polya Urn and the Wright-Fisher process with mutation

In [arXiv:1906.10951 (forthcoming on Advances in Applied Probability),arXiv:2011.05933 (published on PLOS ONE)] the authors introduce, study and apply a new variant of the Eggenberger-Polya urn, called the "Rescaled" Polya urn, which, for a suitable choice of the model parameters, is characterized by the following features: (i) a "local" reinforcement, i.e. a reinforcement mechanism mainly based on the last observations, (ii) a random persistent fluctuation of the predictive mean, and (iii) a long-term almost sure convergence of the empirical mean to a deterministic limit, together with a chi-squared goodness of fit result for the limit probabilities. In this work, motivated by some empirical evidences in [arXiv:2011.05933 (published on PLOS ONE)], we show that the multidimensional Wright-Fisher diffusion with mutation can be obtained as a suitable limit of the predictive means associated to a family of rescaled Polya urns

math.PR

Causal Effects with Hidden Treatment Diffusion on Observed or Partially Observed Networks

In randomized experiments, interactions between units might generate a treatment diffusion process. This is common when the treatment of interest is an actual object or product that can be shared among peers (e.g., flyers, booklets, videos). For instance, if the intervention of interest is an information campaign realized through the distribution of a video to targeted individuals, some of these treated individuals might share the video they received with their friends. Such a phenomenon is usually unobserved, causing a misallocation of individuals in the two treatment arms: some of the initially untreated units might have actually received the treatment by diffusion. Treatment misclassification can, in turn, introduce a bias in the estimation of the causal effect. Inspired by a recent field experiment on the effect of different types of school incentives aimed at encouraging students to attend cultural events, we present a novel approach to deal with a hidden diffusion process on observed or partially observed networks.Specifically, we develop a simulation-based sensitivity analysis that assesses the robustness of the estimates against the possible presence of a treatment diffusion. We simulate several diffusion scenarios within a plausible range of sensitivity parameters and we compare the treatment effect which is estimated in each scenario with the one that is obtained while ignoring the diffusion process. Results suggest that even a treatment diffusion parameter of small size may lead to a significant bias in the estimation of the treatment effect.

stat.ME

Urn models with random multiple drawing and random addition

We consider an urn model with multiple drawing and random time-dependent addition matrix. The model is very general with respect to previous literature: the number of sampled balls at each time-step is random, the addition matrix has general random entries. For the proportion of balls of a given color, we prove almost sure convergence results and fluctuation theorems (through CLTs in the sense of stable convergence and of almost sure conditional convergence, which are stronger than convergence in distribution). Asymptotic confidence intervals are given for the limit proportion, whose distribution is generally unknown.

math.PR

Damping effect in innovation processes: case studies from Twitter

Understanding the innovation process, that is the underlying mechanisms through which novelties emerge, diffuse and trigger further novelties is undoubtedly of fundamental importance in many areas (biology, linguistics, social science and others). The models introduced so far satisfy the Heaps' law, regarding the rate at which novelties appear, and the Zipf's law, that states a power law behavior for the frequency distribution of the elements. However, there are empirical cases far from showing a pure power law behavior and such a deviation is present for elements with high frequencies. We explain this phenomenon by means of a suitable "damping" effect in the probability of a repetition of an old element. While the proposed model is extremely general and may be also employed in other contexts, it has been tested on some Twitter data sets and demonstrated great performances with respect to Heaps' law and, above all, with respect to the fitting of the frequency-rank plots for low and high frequencies.

physics.soc-ph

Generalized Rescaled Polya urn and its statistical applications

We introduce the Generalized Rescaled Polya (GRP) urn, that provides a generative model for a chi-squared test of goodness of fit for the long-term probabilities of clustered data, with independence between clusters and correlation, due to a reinforcement mechanism, inside each cluster. We apply the proposed test to a data set of Twitter posts about COVID-19 pandemic: in a few words, for a classical chi-squared test the data result strongly significant for the rejection of the null hypothesis (the daily long-run sentiment rate remains constant), but, taking into account the correlation among data, the introduced test leads to a different conclusion. Beside the statistical application, we point out that the GRP urn is a simple variant of the standard Eggenberger-Polya urn, that, with suitable choices of the parameters, shows "local" reinforcement, almost sure convergence of the empirical mean to a deterministic limit and different asymptotic behaviours of the predictive mean. Moreover, the study of this model provides the opportunity to analyze stochastic approximation dynamics, that are unusual in the related literature.

math.ST

A model for the Twitter sentiment curve

Twitter is among the most used online platforms for the political communications, due to the concision of its messages (which is particularly suitable for political slogans) and the quick diffusion of messages. Especially when the argument stimulate the emotionality of users, the content on Twitter is shared with extreme speed and thus studying the tweet sentiment if of utmost importance to predict the evolution of the discussions and the register of the relative narratives. In this article, we present a model able to reproduce the dynamics of the sentiments of tweets related to specific topics and periods and to provide a prediction of the sentiment of the future posts based on the observed past. The model is a recent variant of the Pólya urn, introduced and studied in arXiv:1906.10951 and arXiv:2010.06373, which is characterized by a "local" reinforcement, i.e. a reinforcement mechanism mainly based on the most recent observations, and by a random persistent fluctuation of the predictive mean. In particular, this latter feature is capable of capturing the trend fluctuations in the sentiment curve. While the proposed model is extremely general and may be also employed in other contexts, it has been tested on several Twitter data sets and demonstrated greater performances compared to the standard Pólya urn model. Moreover, the different performances on different data sets highlight different emotional sensitivities respect to a public event.

cs.SI

Interacting non-linear reinforced stochastic processes: synchronization and no-synchronization

'Rich get richer' rule comforts previously often chosen actions. What is happening to the evolution of individual inclinations to choose an action when agents do interact ? Interaction tends to homogenize while each individual dynamics tends to reinforce its own position. Interacting stochastic systems of reinforced processes were recently considered in many papers, where the asymptotic behavior was proven to exhibit a.s. synchronization. We consider in this paper models where, even if interaction among agents is present, absence of synchronization may happen due to the choice of an individual non-linear reinforcement. We show how these systems can naturally be considered as models for coordination games, technological or opinion dynamics.

math.PR