SearcharxivSearch

arXiv subjects

Xiao-yan Xue

Publications and source records attributed to Xiao-yan Xue.

4 recordsLinked to original sources

A Unified Exact Factorial-Moment Theory for Multi-set Allocation Occupancy (MAO) in Finite Populations

Let $A_1,\ldots,A_T$ be independent uniformly selected subsets of a finite population of size $n$, with prescribed cardinalities $m_1,\ldots,m_T$. For each population element, define its occupancy level as the number of selected subsets containing it. Let $x_t$ and $x_{\geq t}$ denote the numbers of elements with occupancy exactly $t$ and at least $t$, respectively. The 2025 work introduced a general multi-set allocation occupancy (MAO) representation for higher-order occupancy moments. The present paper establishes its joint-probabilistic interpretation and provides a rigorous unified derivation for arbitrary joint occupancy categories and moment orders. Specifically, for arbitrary $B_1,\ldots,B_\ell\subseteq\{0,1,\ldots,T\}$, we prove the exact representation $F_\ell(B_1,\ldots,B_\ell)=G_T(B_1,\ldots,B_\ell)/(n)_\ell^{T-1}$, where $G_T(B_1,\ldots,B_\ell)$ is the corresponding generalized MAO transversal sum. This identity unifies the joint factorial moments of arbitrary occupancy categories within a single exact finite-population framework. In particular, writing $B_{\geq t}=\{t,t+1,\ldots,T\}$, the factorial moments of exact- and threshold-occupancy counts are obtained as the specializations $\mathbb{E}[(x_t)_\ell]=F_\ell(\{t\},\ldots,\{t\})$ and $\mathbb{E}[(x_{\geq t})_\ell]=F_\ell(B_{\geq t},\ldots,B_{\geq t})$. Mixed factorial moments, raw moments, variances, and covariances follow from the same representation through standard transformations. The formulas are verified by exhaustive enumeration over feasible parameter ranges and by Monte Carlo simulation. The resulting theory provides a rigorous and unified finite-population foundation for exact and threshold multi-set occupancy statistics.

math.PR

The Multi-set Allocation Occupancy function and inequality (MAO function and MAO inequality): the foundation of Generalized hypergeometric distribution theory

In our previous work, we studied the Generalized Hypergeometric Distribution (GHGD), which we refer to as the Multi-set Allocation Occupancy (MAO) distribution. We derived formulas for its expectation and variance for any number of subsets $T$ and overlap count $t$ ($1 \le t \le T$), and established an asymptotic property. However, these formulas were complex, and higher moments were not derived. Through further study, we have established a novel function that describes all higher moments of the MAO distribution with a unified, elegant formula. The core definitions are the MAO function $g(A_1, A_2, \dots, A_r) = \prod_{i=1}^{T} (m_i)_{k_i} \cdot (n-m_i)_{r-k_i}$ and the MAO norm $\|(p_1, \dots, p_r)\|_T = \frac{\sum_{A_1, \dots, A_r \subseteq [T] \; : \; |A_j|=p_j} g(A_1, \dots, A_r)}{((n)_r)^{T-1}}$, where $p_i$ is the size of subset $A_i$, $m_i < n$, and $(x)_r$ is the falling factorial. Using these definitions, the intricate moment relations simplify into a unified form: the $ν$-th raw moment of $p(x_{=t})$ and $p(x_{\ge t})$ can be calculated as $E(x_{=t}^ν) = \sum_{1 \le i \le ν} s_{ν,i} \|t^i\|$ and $E(x_{\ge t}^ν) = \sum_{1 \le i \le ν} s_{ν,i} \|[t, T]^i\|$, where $s_{ν,i}$ are Stirling numbers of the second kind and $[t,T] = \{t, t+1, \dots, T\}$. Furthermore, based on the MAO norm, we formulate a novel MAO inequality under the proximity condition $\max(p_i) - \min(p_i) \le 1$: $\prod_{1\le i \le r} \|(p_i)\|_T \ge \|(p_1, \dots, p_r)\|_T$. A direct corollary is the asymptotic property of the MAO distribution: $E(X) > \text{Var}(X)$ and $E(X) - \text{Var}(X) = o(E(X))$ as $E(X) \to 0$.

math.PR

Mean, Variance and Asymptotic Property for General Hypergeometric Distribution

General hypergeometric distribution (GHGD) definition: from a finite space $N$ containing $n$ elements, randomly select totally $T$ subsets $M_i$ (each contains $m_i$ elements, $1 \geq i \geq T$), what is the probability that exactly $x$ elements are overlapped exactly $t$ times or at least $t$ times ($x_t$ or $x_{\geq t}$)? The GHGD described the distribution of random variables $x_t$ and $x_{\geq t}$. In our previous results, we obtained the formulas of mathematical expectation and variance for special situations ($T \leq 7$), and not provided proofs. Here, we completed the exact formulas of mean and variance for $x_t$ and $x_{\geq t}$ for any situation, and provided strict mathematical proofs. In addition, we give the asymptotic property of the variables. When the mean approaches to 0, the variance fast approaches to the value of mean, and actually, their difference is a higher order infinitesimal of mean. Therefore, when the mean is small enough ($<1$), it can be used as a fairly accurate approximation of variance.

math.PR

General hypergeometric distribution: A basic statistical distribution for the number of overlapped elements in multiple subsets drawn from a finite population

General hypergeometric distribution (GHGD) describes the following distribution: from a finite space containing N elements, select T subsets with each subset contains M[i] (T-1 >= i >= 0) elements, what is the probability that exactly x elements are overlapped exactly t times or at least t times (XLO=t or XLO>=t, T >= t >= 0, here LO is level of overlap)? The classical hypergeometric distribution (HGD) describes the situation of two subsets, while the general situation has not been resolved, despite the overlapped elements has been visualized with the Venn diagram method for about 140 years. GHGD described not only the distribution of XLO=t or XLO>=t that are overlapped in all of the subsets (XLO=T), but also the XLO=t or XLO>=t that are overlapped in a portion of the subsets (LO = t or LO >= t, T >= t >= 0). Here, we developed algorithms to calculate the GHGD and discovered graceful formulas of the essential statistics for the GHGD, including mathematical expectation, variance, and high order moments. In addition, statistical theory to infer a statistically reliable gene set from multiple datasets based on these formulas was established by applying Chebyshev's inequalities.

math.ST