Independence properties for tree beta models
This paper is devoted to introduce new probabilistic tree-based models complementing already known Matsumoto-Yor and Hamza-Vallois tree models with two tree beta-type models.
arXiv subjects
Publications and source records attributed to Jacek Wesołowski.
This paper is devoted to introduce new probabilistic tree-based models complementing already known Matsumoto-Yor and Hamza-Vallois tree models with two tree beta-type models.
A map $F\colon\mathcal X\times\mathcal Y\to \mathcal U\times \mathcal V$ is said to be independence preserving (IP) if there exists a pair of independent random variables $(X,Y)$ valued in $\mathcal X\times\mathcal Y$ such that the two coordinates of $(U,V)=F(X,Y)$ are also independent. Recently, Sasada and Uozumi (2024) observed that a hierarchy of quadrirational Yang-Baxter maps gives rise to independence preserving transformations, and identified corresponding families of probability distributions. In view of the limiting properties of these IP maps, the newly defined generalized second kind beta ($\mathrm{GB}_{II}$) model stands at the top of the hierarchy: for independent random variables $X$ and $Y$ following a $\mathrm{GB}_{II}$ distribution, Sasada and Uozumi (2024) showed that when a special quadrirational Yang-Baxter map $F^{(α,β)}$, parameterized by $(α,β)\in(0,\infty)^2$, is applied to the pair $(X,Y)$, it produces another pair $(U,V)$ of independent $\mathrm{GB}_{II}$-distributed random variables. The aim of this paper is to show that the IP property of $F^{(α,β)}$ uniquely identifies distributions of $X,Y,U$ and $V$ as belonging to the $\mathrm{GB}_{II}$ family. To this end, we introduce specially designed Laplace-type transforms. First, we carefully explain the connection between the results from Sasada and Uozumi (2024) and Koudou and Vallois (2012). Next, we focus on the characterization of the second kind beta and the generalized second kind beta distributions through the IP map $F^{(α,\infty)}$. Finally, extending considerably the methodology developed for the case $(α,\infty)$, we prove the characterization of $\mathrm{GB}_{II}$ distributions in the case $(α,β)\in(0,\infty)^2$ with $α\neqβ$, which implies uniqueness in the ultimate missing case of the quadrirational Yang-Baxter hierarchy of IP models.
Sasada and Uozumi, \cite{SasUoz2024}, identified independence preserving $[2:2]$ quadrirational parametric Yang-Baxter maps, see \eqref{YBEQ}, on $(0,\infty)$. In particular, the map denoted there by $H_{III,B}^{(α,β)}$, see \eqref{CS}, was connected to the independence preserving property of the GIG distributions on $(0,\infty)$. Remarkably, the property appears also naturally in probabilistic integrable models of discrete Korteweg de Vries type, as observed by Croydon and Sasada, \cite{CroSas2020}. In the case of $(α,β)=(1,0)$ the independence reduces to the classical Matsumoto-Yor property, \cite{MatYor2001}. In \cite{LetWes2024} we proposed an extension of $H_{III,B}^{(α,β)}$ to a map on the cone of symmetric positive definite matrices of a fixed dimension, showing that such extended map preserves independence of GIG random matrices. In the present paper we prove two results: (i) the matrix GIG distributions are characterized by the independence property governed by this map; (ii) the matrix variate extension of $H_{III,B}^{(α,β)}$ we use, is a parametric Yang-Baxter map.
Bayesian statistical graphical models are typically classified as either continuous and parametric (Gaussian, parameterized by the graph-dependent precision matrix with Wishart-type priors) or discrete and non-parametric (with graph-dependent structure of probabilities of cells and Dirichlet-type priors). We propose to break this dichotomy by introducing two discrete parametric graphical models on finite decomposable graphs: the graph negative multinomial and the graph multinomial distributions (the former related to the Cartier-Foata theorem for the graph genereted free quotient monoid). These models interpolate between the product of univariate negative binomial laws and the negative multinomial distribution, and between the product of binomial laws and the multinomial distribution, respectively. We derive their Markov decompositions and provide related probabilistic representations. We also introduce graphical versions of the Dirichlet and inverted Dirichlet distributions, which serve as conjugate priors for the two discrete graphical Markov models. We derive explicit normalizing constants for both graphical Dirichlet laws and establish their independence structure (a graphical version of neutrality), which yields a strong hyper Markov property for both Bayesian models. We also provide characterization theorems for graphical Dirichlet laws via respective graphical versions of neutrality, which extends previously known results.
For a TASEP on $\mathbb Z$ with the step initial condition we identify limits as $t\to\infty$ of the expected total number of jumps until time $t>0$ and the expected number of active particles at a time $t$. We also connect the two quantities proving that non-asymptotically, that is as a function of $t>0$, the latter is the derivative of the former. Our approach builds on asymptotics derived by Rost and intensive use of the fact that the rightmost particle evolves according to the Poisson process.
In this paper the relations between independence preserving (IP) involutions and reversible Markov kernels are investigated. We introduce an involutive augmentation H = (f, g_f) of a measurable function f and relate the IP property of H to f-generated reversible Markov kernels. Various examples appeared in the literature are presented as particular cases of the construction. In particular, we prove that the IP property generated by the (reversible) Markov kernel of random walk with a reflecting barrier at the origin characterizes geometric-type laws
The optimum sample allocation in stratified sampling is one of the basic issues of survey methodology. It is a procedure of dividing the overall sample size into strata sample sizes in such a way that for given sampling designs in strata the variance of the stratified $π$ estimator of the population total (or mean) for a given study variable assumes its minimum. In this work, we consider the optimum allocation of a sample, under lower and upper bounds imposed jointly on sample sizes in strata. We are concerned with the variance function of some generic form that, in particular, covers the case of the simple random sampling without replacement in strata. The goal of this paper is twofold. First, we establish (using the Karush-Kuhn-Tucker conditions) a generic form of the optimal solution, the so-called optimality conditions. Second, based on the established optimality conditions, we derive an efficient recursive algorithm, named RNABOX, which solves the allocation problem under study. The RNABOX can be viewed as a generalization of the classical recursive Neyman allocation algorithm, a popular tool for optimum allocation when only upper bounds are imposed on sample strata-sizes. We implement RNABOX in R as a part of our package stratallo which is available from the Comprehensive R Archive Network (CRAN) repository.
Quadratic harnesses are time-inhomogeneous Markov polynomial processes with linear conditional expectations and quadratic conditional variances with respect to the past-future filtrations. Typically they are determined by five numerical constants hidden in the form of conditional variances. In this paper we derive infinitesimal generators of such processes, extending previously known results. The infinitesimal generators are identified through a solution of a q-commutation equation in the algebra Q of infinite sequences of polynomials in one variable. The solution is a special element in Q, whose coordinates satisfy a three-term recurrence and thus define a system of orthogonal polynomials. It turns out that the respective orthogonality measure uniquely determines the infinitesimal generator (acting on polynomials or bounded functions with bounded continuous second derivative) as an integro-differential operator with the explicit kernel, where the integration is with respect to this measure.
We prove that if $X,Y$ are positive, independent, non-Dirac random variables and if for $α,β\ge 0$, $α\neq β$, $$ ψ_{α,β}(x,y)=\left(y\,\tfrac{1+β(x+y)}{1+αx+βy},\;x\,\tfrac{1+α(x+y)}{1+αx+βy}\right), $$ then the random variables $U$ and $V$ defined by $(U,V)=ψ_{α,β}(X,Y)$ are independent if and only if $X$ and $Y$ follow Kummer distributions with suitably related parameters. In other words, any invariant measure for a lattice recursion model governed by $ψ_{α,β}$ in the scheme introduced by Croydon and Sasada in \cite{CS2020} is necessarily a product measure with Kummer marginals. The result extends earlier characterizations of Kummer and gamma laws by independence of $$ U=\tfrac{Y}{1+X}\quad\mbox{and}\quad V= X\left(1+\tfrac{Y}{1+X}\right), $$ which corresponds to the case of $ψ_{1,0}$. We also show that this independence property of Kummer laws covers, as limiting cases, several independence models known in the literature: the Lukacs, the Kummer-Gamma, the Matsumoto-Yor and the discrete Korteweg de Vries models.
If $α,β>0$ are distinct and if $A$ and $B$ are independent non-degenerate positive random variables such that $$S=\tfrac{1}{B}\,\tfrac{βA+B}{αA+B}\quad \mbox{and}\quad T=\tfrac{1}{A}\,\tfrac{βA+B}{αA+B} $$ are independent, we prove that this happens if and only if the $A$ and $B$ have generalized inverse Gaussian distributions with suitable parameters. Essentially, this has already been proved in Bao and Noack (2021) with supplementary hypothesis on existence of smooth densities. The sources of these questions are an observation about independence properties of the exponential Brownian motion due to Matsumoto and Yor (2001) and a recent work of Croydon and Sasada (2000) on random recursion models rooted in the discrete Korteweg - de Vries equation, where the above result was conjectured. We also extend the direct result to random matrices proving that a matrix variate analogue of the above independence property is satisfied by independent matrix-variate GIG variables. The question of characterization of GIG random matrices through this independence property remains open.
We introduce one-way flows in near algebras and two-way flows in double near algebras with two interrelated multiplications. We establish parametric representations of the one-way and two-way flows in terms of a single element of the algebra that we call a flow generator. We indicate probabilistic applications of the one-way flows to a study of polynomial stochastic processes. We apply our results on the two-way flows to harnesses and quadratic harnesses in probability theory, generalizing some previous results.
We point out to a connection between a problem of invariance of power series families of probability distributions under binomial thinning and functional equations which generalize both the Cauchy and an additive form of the Gołab-Schinzel equation. We solve these equations in several settings with no or mild regularity assumptions imposed on unknown functions.
We derive a formula for the optimal sample allocation in a general stratified scheme under upper bounds on the sample strata-sizes. Such a general scheme includes SRSWOR within strata as a special case. The solution is given in terms of V-allocation with V being the set of take-all strata. We use V-allocation to give a formal proof of optimality of the popular recursive Neyman algorithm, rNa. This approach is convenient also for a quick proof of optimality of the algorithm of Stenger and Gabler (2005), SGa, as well as of its modification, coma, we propose here. Finally, we compare running times of rNa, SGa and coma. Ready-to-use R-implementations of these algorithms are available on CRAN repository at https://cran.r-project.org/web/packages/stratallo.
We use here a recent idea of studying functions of free random variables using Boolean cumulants. We develop idea of explicit calculations of conditional expectation using Boolean cumulants. We demonstrate Boolean cumulants approach allows to calculate explicitly some conditional expectations of functions in free random variables. We present how Boolean cumulants together with subordination simplify proofs of some results which are known in research literature.
In Sabot and Tarrès (2015), the authors have explicitly computed the integral $$STZ_n=\int \exp( -\langle x,y\rangle)(\det M_x)^{-1/2}dx$$ where $M_x$ is a symmetric matrix of order $n$ with fixed non positive off-diagonal coefficients and with diagonal $(2x_1,\ldots,2x_n)$. The domain of integration is the part of $\mathbb{R}^n$ for which $M_x$ is positive definite. We calculate more generally for $ b_1\geq 0,\ldots b_n\geq 0$ the integral $$GSTZ_n=\int \exp \left(-\langle x,y\rangle-\frac{1}{2}b^*M_x^{-1}b\right)(\det M_x)^{-1/2}dx,$$ we show that it leads to a natural family of distributions in $\mathbb{R}^n$, called the $GSTZ_n$ probability laws. This family is stable by marginalization and by conditioning, and it has number of properties which are multivariate versions of familiar properties of univariate reciprocal inverse Gaussian distribution. We also show that if the graph with the set of vertices $V=\{1,\ldots,n\}$ and the set $E$ of edges $\{i,j\}'$ s of non zero entries of $M_x$ is a tree, then the integral $$\int \exp( -\langle x,y\rangle)(\det M_x)^{q-1}dx$$ where $q>0,$ is computable in terms of the MacDonald function $K_q.$
Consider a number, finite or not, of urns each with fixed capacity $r$ and balls randomly distributed among them. An overflow is the number of balls that are assigned to urns that already contain $r$ balls. When $r=1$, using analytic methods, Hwang and Janson gave conditions under which the overflow (which in this case is just the number of balls landing in non--empty urns) has an asymptotically Poisson distribution as the number of balls grows to infinity. Our aim here is to systematically study the asymptotics of the overflow in general situation, i.~e. for arbitrary $r$. In particular, we provide sufficient conditions for both Poissonian and normal asymptotics for general $r$, thus extending Hwang--Janson's work. Our approach relies on purely probabilistic methods.
In this paper we are interested in the joint distribution of two order statistics from overlapping samples. We give an explicit formula for the distribution of such a pair of random variables under the assumption that the parent distribution is absolutely continuous (with respect to the Lebesgue measure on the real line). The distribution is identified through the form of the density with respect to a measure which is a sum of the bivariate Lebesgue measure on $\R^2$ and the univariate Lebesgue measure on the diagonal $\{(x,x):\,x\in\R\}$. We are also interested in the question to what extent conditional expectation of one of such order statistic given another determines the parent distribution. In particular, we provide a new characterization by linearity of regression of an order statistic from the extended sample given the one from the original sample, special case of which solves a problem explicitly stated in the literature. It appears that to describe the correct parent distribution it is convenient to use quantile density functions. In several other cases of regressions of order statistics we provide new results regarding uniqueness of the distribution in the sample. Nevertheless the general question of identifiability of the parent distribution by regression of order statistics from overlapping samples remains open.
If $X$ and $Y$ are independent random variables with distributions $μ$ and $ν$ then $U=ψ(X,Y)$ and $V=ϕ(X,Y)$ are also independent for some $ψ$ and $ϕ$. Properties of this type are known for many important probability distributions $μ$ and $ν$. Also related characterization questions have been widely investigated: Let $X$ and $Y$ be independent and let $U$ and $V$ be independent. Are the distributions of $X$ and $Y$ $μ$ and $ν$, respectively? Recently two new properties and characterizations of this kind involving the Kummer distribution appeared in the literature. For independent $X$ and $Y$ with gamma and Kummer distributions Koudou and Vallois observed that $U=(1+(X+Y)^{-1})/(1+X^{-1})$ and $V=X+Y$ are also independent, and Hamza and Vallois observed that $U=Y/(1+X)$ and $V=X(1+Y/(1+X))$ are independent. In 2011 and 2012 Koudou, Vallois characterizations related to the first property were proved, while the characterizations in the second setting have been recently given in Piliszek, Wesołowski (2016). In both cases technical assumptions on smoothness properties of densities of $X$ and $Y$ were needed. In 2015, the assumption of independence of $U$ and $V$ in the first setting was weakened to constancy of regressions of $U$ and $U^{-1}$ given $V$ with no density assumptions. However, the additional assumption $\mathbb{E} X^{-1}<\infty$ was introduced. In the present paper we provide a complete answer to the characterization question in both settings without any additional technical assumptions regarding smoothness or existence of moments. The approach is, first, via characterizations exploiting some conditions imposed on regressions of $U$ given $V$, which are weaker than independence, but for which moment assumptions are necessary. Second, using a technique of change of measure we show that the moment assumptions can be avoided.