SearcharxivSearch

arXiv subjects

Donald Richards

Publications and source records attributed to Donald Richards.

At least 19 recordsLinked to original sources

Omnibus Goodness-of-Fit Testing for Distributions on Stiefel Manifolds

In this article, a comprehensive framework for goodness-of-fit testing for distributions on Stiefel manifolds is developed. The approach is based on integrals of the squared differences between empirical and theoretical characteristic functions, yielding test statistics that are consistent against all fixed alternatives. For the Fisher-Bingham family of distributions, explicit computable forms of the test statistic are derived. Simplified expressions for important special cases, including the matrix Fisher, matrix Bingham, and uniform distributions are provided. In the case of testing uniformity on hyperspheres, we obtain the complete asymptotic distribution of the test statistic, enabling computationally efficient asymptotic testing. For general Fisher-Bingham distributions, we establish theoretically justified Monte Carlo testing procedures for both simple and composite hypotheses. Simulation studies demonstrate accurate Type I error control and strong power across a wide range of alternatives. The practical relevance of the proposed methodology is illustrated by an application to data on the orbits of comets.

math.ST

Stein's method for the Wishart distribution

In this work, we develop Stein's method for the Wishart distribution on the cone of positive definite matrices. We establish the basic ingredients of a Wishart Stein framework: we derive an extended-generator-based Stein characterization from the Wishart diffusion process, identify the corresponding transition semigroup through the noncentral Wishart law, provide an explicit semigroup representation for the solution of the Stein equation, and obtain regularity estimates for the solution. The new methodology is demonstrated in four applications: (i) an order $n^{-1}$ bound, for smooth test functions, for the Wishart approximation of uncentered group-mean scatter matrices in MANOVA; (ii) a quantitative multivariate Satterthwaite approximation; (iii) local/integrated De Bruijn identities and logarithmic Sobolev inequalities for the Wishart measure; and (iv) Stein's method of moments for the shape and scale parameters, including structured scale estimation.

math.PR

Stein's method for the matrix normal distribution

This work presents the first systematic development of Stein's method for matrix distributions. We establish the basic essential ingredients of Stein's method for matrix normal approximation: we derive an extended-generator-based Stein identity from a matrix Ornstein-Uhlenbeck diffusion with two-sided scales, provide an explicit semigroup representation for the solution of the Stein equation, and obtain regularity estimates for the solution. The new methodology is demonstrated in three examples: (i) smooth Wasserstein distance bounds to quantify the matrix central limit theorem (a didactic example), (ii) a Wasserstein distance bound for the matrix normal approximation of the centered matrix $T$ distribution, and (iii) a Stein's method-of-moments approach to estimating the row and column covariance factors of the matrix normal, yielding a flexible class of weighted flip-flop Stein estimators that generalize Dutilleul's classical flip-flop algorithm and naturally accommodate row/column importance weights, systematic missingness, and projection onto structured covariance families. The latter two examples are intrinsically matrix-valued and cannot be treated using naive vectorization.

math.ST

Wishart kernel density estimation for strongly mixing time series on the cone of positive definite matrices

A Wishart kernel density estimator (KDE) is introduced for density estimation in the cone of positive definite matrices. The estimator is boundary-aware and mitigates the boundary bias suffered by conventional KDEs, while remaining simple to implement. Its mean squared error, uniform strong consistency on expanding compact sets, and asymptotic normality are established under the Lebesgue measure and suitable mixing conditions. This work represents the first study of density estimation for dependent data on this space under any metric. For independent observations, an asymptotic upper bound on the mean absolute error is also derived. A simulation study compares the performance of the Wishart KDE with that of the log-Gaussian KDE, another boundary-aware estimator based on the matrix-variate lognormal distribution proposed by Schwartzman [Int. Stat. Rev., 2016, 84(3), 456--486], and with the naive Gaussian KDE on the ambient Euclidean space. When estimating the stationary marginal density of a Wishart autoregressive process for several autoregressive coefficient matrices and innovation covariance matrices, the Wishart KDE exhibits the best overall accuracy and stability. The practical utility of the Wishart KDE is illustrated by estimating the marginal density of a one-year time series of realized covariance matrices computed from 5-minute intra-day returns on Amazon Corp. shares and on the Standard & Poor's 500 exchange-traded fund. All code is publicly available via the R package ksm to facilitate implementation of the method and reproducibility of the findings.

stat.ME

A Tribute to Richard Askey, and an Expansive View of Some of his Favorite Beta Integrals

This article represents a personal tribute to Richard Askey together with a new look at some of his favorite integrals, including the Cauchy beta integral. The article also provides some new multidimensional extensions of Cauchy's beta integral in which the domain of integration is the space of real symmetric matrices, and these multidimensional integrals are used to obtain some special cases of the Cauchy--Selberg integrals.

math.CA

Random Linear Modulation with Spherically Symmetric Modulators

We consider the modulation of data given by random vectors $X_n \in \mathbb{R}^{d_n}$, $n \in \mathbb{N}$. For each $X_n$, one chooses an independent modulating random vector $\Xi_n \in \mathbb{R}^{d_n}$ and forms the projection $Y_n = \Xi_n'X_n$. It is shown, under regularity conditions on $X_n$ and $\Xi_n$, that $Y_n|\Xi_n$ converges weakly in probability to a normal distribution. More broadly, the conditional joint distribution of a family of projections constructed from random samples from $X_n$ and $\Xi_n$ is shown to converge weakly to a matrix normal distribution. We derive, $via$ G. P\'olya's characterization of the normal distribution, a necessary and sufficient condition on $Y_n$ for $\Xi_n$ to be normally distributed. When $\Xi_n$ has a spherically symmetric distribution we deduce, through I. J. Schoenberg's characterization of the spherically symmetric characteristic functions on Hilbert spaces, that the probability density function of $Y_n|\Xi_n$ converges pointwise in certain $p$th means to a mixture of normal densities and the rate of convergence is quantified, resulting in uniform convergence. The cumulative distribution function of $Y_n|\Xi_n$ is shown to converge uniformly in those $p$th means to the distribution function of the same mixture, and a Lipschitz property is obtained. Examples of distributions satisfying our results are provided; these include Bingham distributions on hyperspheres of random radii, uniform distributions on hyperspheres and hypercubes of random volumes, and multivariate normal distributions; and examples of such $\Xi_n$ include the multivariate $t$-, multivariate Laplace, and spherically symmetric stable distributions.

math.ST

An explicit Wishart moment formula for the product of two disjoint principal minors

This paper provides the first explicit formula for the expectation of the product of two disjoint principal minors of a Wishart random matrix, solving a part of a broader problem put forth by Samuel S. Wilks in 1934 in the Annals of Mathematics. The proof makes crucial use of hypergeometric functions of matrix argument and their Laplace transforms. Additionally, a Wishart generalization of the Gaussian product inequality conjecture is formulated and a stronger quantitative version is proved to hold in the case of two minors.

math.PR

On Wilks' joint moment formulas for embedded principal minors of Wishart random matrices

In 1934, the American statistician Samuel S. Wilks derived remarkable formulas for the joint moments of embedded principal minors of sample covariance matrices in multivariate Gaussian populations, and he used them to compute the moments of sample statistics in various applications related to multivariate linear regression. These important but little-known moment results were extended in 1963 by the Australian statistician A. Graham Constantine using Bartlett's decomposition. In this note, a new proof of Wilks' results is derived using the concept of iterated Schur complements, thereby bypassing Bartlett's decomposition. Furthermore, Wilks' open problem of evaluating joint moments of disjoint principal minors of Wishart random matrices is related to the Gaussian product inequality conjecture.

math.ST

Complete Asymptotic Expansions for the Normalizing Constants of High-Dimensional Matrix Bingham and Matrix Langevin Distributions

For positive integers $d$ and $p$ such that $d \ge p$, let $\mathbb{R}^{d \times p}$ denote the set of $d \times p$ real matrices, $I_p$ be the identity matrix of order $p$, and $V_{d,p} = \{x \in \mathbb{R}^{d \times p} \mid x'x = I_p\}$ be the Stiefel manifold in $\mathbb{R}^{d \times p}$. Complete asymptotic expansions as $d \to \infty$ are obtained for the normalizing constants of the matrix Bingham and matrix Langevin probability distributions on $V_{d,p}$. The accuracy of each truncated expansion is strictly increasing in $d$; also, for sufficiently large $d$, the accuracy is strictly increasing in $m$, the number of terms in the truncated expansion. Lower bounds are obtained for the truncated expansions when the matrix parameters of the matrix Bingham distribution are positive definite and when the matrix parameter of the matrix Langevin distribution is of full rank. These results are applied to obtain the rates of convergence of the asymptotic expansions as both $d \to \infty$ and $p \to \infty$. Values of $d$ and $p$ arising in numerous data sets are used to illustrate the rate of convergence of the truncated approximations as $d$ or $m$ increases. These results extend recently-obtained asymptotic expansions for the normalizing constants of the high-dimensional Bingham distributions.

math.ST

EM Estimation of the B-Spline Copula with Penalized Pseudo-Likelihood Functions

The B-spline copula function is defined by a linear combination of elements of the normalized B-spline basis. We develop a modified EM algorithm, to maximize the penalized pseudo-likelihood function, wherein we use the smoothly clipped absolute deviation (SCAD) penalty function for the penalization term. We conduct simulation studies to demonstrate the stability of the proposed numerical procedure, show that penalization yields estimates with smaller mean-square errors when the true parameter matrix is sparse, and provide methods for determining tuning parameters and for model selection. We analyze as an example a data set consisting of birth and death rates from 237 countries, available at the website, ''Our World in Data,'' and we estimate the marginal density and distribution functions of those rates together with all parameters of our B-spline copula model.

stat.ME

On the Gaussian product inequality conjecture for disjoint principal minors of Wishart random matrices

This paper extends various results related to the Gaussian product inequality (GPI) conjecture to the setting of disjoint principal minors of Wishart random matrices. This includes product-type inequalities for matrix-variate analogs of completely monotone functions and Bernstein functions of Wishart disjoint principal minors, respectively. In particular, the product-type inequalities apply to inverse determinant powers. Quantitative versions of the inequalities are also obtained when there is a mix of positive and negative exponents. Furthermore, an extended form of the GPI is shown to hold for the eigenvalues of Wishart random matrices by virtue of their law being multivariate totally positive of order 2 (MTP${}_2$). A new, unexplored avenue of research is presented to study the GPI from the point of view of elliptical distributions.

math.ST

Complete Asymptotic Expansions and the High-Dimensional Bingham Distributions

For $d \ge 2$, let $X$ be a random vector having a Bingham distribution on $\mathcal{S}^{d-1}$, the unit sphere centered at the origin in $\R^d$, and let $\Sigma$ denote the symmetric matrix parameter of the distribution. Let $\Psi(\Sigma)$ be the normalizing constant of the distribution and let $\nabla \Psi_d(\Sigma)$ be the matrix of first-order partial derivatives of $\Psi(\Sigma)$ with respect to the entries of $\Sigma$. We derive complete asymptotic expansions for $\Psi(\Sigma)$ and $\nabla \Psi_d(\Sigma)$, as $d \to \infty$; these expansions are obtained subject to the growth condition that $\|\Sigma\|$, the Frobenius norm of $\Sigma$, satisfies $\|\Sigma\| \le \gamma_0 d^{r/2}$ for all $d$, where $\gamma_0 > 0$ and $r \in [0,1)$. Consequently, we obtain for the covariance matrix of $X$ an asymptotic expansion up to terms of arbitrary degree in $\Sigma$. Using a range of values of $d$ that have appeared in a variety of applications of high-dimensional spherical data analysis we tabulate the bounds on the remainder terms in the expansions of $\Psi(\Sigma)$ and $\nabla \Psi_d(\Sigma)$ and we demonstrate the rapid convergence of the bounds to zero as $r$ decreases.

math.ST

Non-Steepness and Maximum Likelihood Estimation Properties of the Truncated Multivariate Normal Distributions

This article considers exponential families of truncated multivariate normal distributions with one-sided truncation for some or all coordinates. We observe that if all components are one-sided truncated then this family is not full. The family of truncated multivariate normal distributions is extended to a full family, and the extended family is investigated in detail. We identify the canonical parameter space of the extended family and establish that the family is not regular and not even steep. We also consider maximum likelihood estimation for the location vector parameter and the positive definite (symmetric) matrix dispersion parameter of a truncated non-singular multivariate normal distribution. It is shown that if the sample size is sufficiently large then, almost surely, the maximizer of the likelihood function is unique, provided that it exists. It is also shown that each solution to the score equations for the location and dispersion parameters satisfies the method-of-moments equations. Finally, it is observed that similar results arise in the case of an arbitrary number of truncated components.

math.ST

Hoffmann-J{\o}rgensen Inequalities for Random Walks on the Cone of Positive Definite Matrices

We consider random walks on the cone of $m \times m$ positive definite matrices, where the underlying random matrices have orthogonally invariant distributions on the cone and the Riemannian metric is the measure of distance on the cone. By applying results of Khare and Rajaratnam (Ann. Probab., 45 (2017), 4101--4111), we obtain inequalities of Hoffmann-J{\o}rgensen type for such random walks on the cone. In the case of the Wishart distribution $W_m(a,I_m)$, with index parameter $a$ and matrix parameter $I_m$, the identity matrix, we derive explicit and computable bounds for each term appearing in the Hoffmann-J{\o}rgensen inequalities.

math.PR

A Continuous-Time Markov Chain Model for the Spread of COVID-19

Since late 2019 the novel coronavirus, also known as COVID-19, has caused a pandemic that persists. This paper shows how a continuous-time Markov chain model for the spread of COVID-19 can be used to explain, and justify to undergraduate students, strategies now being used in attempts to control the virus. The material in the paper is written at the level of students who are taking an introductory course on the theory and applications of stochastic processes.

stat.AP

A Basic Treatment of the Distance Covariance

The distance covariance of Sz\'ekely, et al. [23] and Sz\'ekely and Rizzo [21], a powerful measure of dependence between sets of multivariate random variables, has the crucial feature that it equals zero if and only if the sets are mutually independent. Hence the distance covariance can be applied to multivariate data to detect arbitrary types of non-linear associations between sets of variables. We provide in this article a basic, albeit rigorous, introductory treatment of the distance covariance. Our investigations yield an approach that can be used as the foundation for presentation of this important and timely topic even in advanced undergraduate- or junior graduate-level courses on mathematical statistics.

math.ST

Product Inequalities for Multivariate Gaussian, Gamma, and Positively Upper Orthant Dependent Distributions

The Gaussian product inequality is an important conjecture concerning the moments of Gaussian random vectors. While all attempts to prove the Gaussian product inequality in full generality have been unsuccessful to date, numerous partial results have been derived in recent decades and we provide here further results on the problem. Most importantly, we establish a strong version of the Gaussian product inequality for multivariate gamma distributions in the case of nonnegative correlations, thereby extending a result recently derived by Genest and Ouimet [5]. Further, we show that the Gaussian product inequality holds with nonnegative exponents for all random vectors with positive components whenever the underlying vector is positively upper orthant dependent. Finally, we show that the Gaussian product inequality with negative exponents follows directly from the Gaussian correlation inequality.

math.PR

Shrinkage Estimation for the Diagonal Multivariate Exponential Families

We study shrinkage estimation of the mean parameters of a class of multivariate distributions for which the diagonal entries of the corresponding covariance matrix are certain quadratic functions of the mean parameter. This class of distributions includes the diagonal multivariate natural exponential families. We propose two classes of semi-parametric shrinkage estimators for the mean and construct unbiased estimators of the corresponding risk. We establish the asymptotic consistency and convergence rates for these shrinkage estimators under squared error loss as both $n$, the sample size, and $p$, the dimension, tend to infinity. Next, we specialize these results to the diagonal multivariate natural exponential families, which have been classified as consisting of the normal, Poisson, gamma, multinomial, negative multinomial, and hybrid classes of distributions. We establish the consistency of our estimators in the normal, gamma, and negative multinomial cases subject to the condition that $p n^{-1/3} (\log{n})^{4/3} \to 0$, and in the Poisson and multinomial cases if $p n^{-1/2} \to 0$, as $n,p \to \infty$. Simulation studies are provided to evaluate the performance of our estimators and we illustrate that, in the gamma and Poisson cases, our estimators achieve lower risk than the maximum likelihood estimator, thereby demonstrating the superiority of our estimators over the maximum likelihood estimator.

math.ST