SearcharxivSearch

arXiv subjects

Satoshi Kuriki

Publications and source records attributed to Satoshi Kuriki.

At least 19 recordsLinked to original sources

Expected number density of critical points of smooth Gaussian random fields in arbitrary dimensions

We obtain explicit formulas for the expected number and height distribution of critical points of smooth isotropic Gaussian random fields on $\mathbb{R}^d$. The expected number density formula is expressed in terms of at most one-dimensional integrals, regardless of the dimension $d$. To obtain the formulas, we provide a variant of de Bruijn's theorem, as well as Weierstrass' convolution formula with a Gaussian random variable and its inversion.

math.ST

Rigorous Formulation of Finite-Sample and Finite-Window Effects in Galaxy Clustering

Galaxy surveys provide finite catalogs of objects observed within bounded volumes, yet clustering statistics are often interpreted using theoretical frameworks developed for infinite point processes. In this work, we formulate key statistical quantities directly for finite point processes and examine the structural consequences of finite-number and finite-window constraints. We show that several well-known features of galaxy survey analysis arise naturally from finiteness alone. In particular, non-vanishing higher-order connected correlations can occur even in statistically independent samples when the total number of points is fixed, and the integral constraint in two-point statistics appears as an exact identity implied by the finite-number condition rather than as an estimator artifact. We further demonstrate that counts-in-cells and point-centered environmental measures correspond to distinct statistical ensembles. Using Palm conditioning, we derive an exact relation between random-cell and point-centered statistics, showing that the latter probe a tilted version of the underlying distribution. These results provide a probabilistic framework for separating structural effects imposed by finite sampling from correlations reflecting genuine astrophysical processes. The formulation presented here remains valid for realistic survey geometries and finite data sets and clarifies the interpretation of commonly used clustering statistics in galaxy surveys.

astro-ph.GA

Tube formula for spherically contoured random fields with subexponential marginals

It is widely known that the tube method, or equivalently the Euler characteristic heuristic, provides a very accurate approximation for the tail probability that the supremum of a smooth Gaussian random field exceeds a threshold value $c$. The relative approximation error $\Delta(c)$ is exponentially small as a function of $c$ when $c$ tends to infinity. On the other hand, little is known about non-Gaussian random fields. In this paper, we obtain the approximation error of the tube method applied to the canonical isotropic random fields on a unit sphere defined by $u\mapsto\langle u,\xi\rangle$, $u\in M\subset\mathbb{S}^{n-1}$, where $\xi$ is a spherically contoured random vector. These random fields have statistical applications in multiple testing and simultaneous regression inference when the unknown variance is estimated. The decay rate of the relative error $\Delta(c)$ depends on the tail of the distribution of $\|\xi\|^2$ and the critical radius of the index set $M$. If this distribution is subexponential but not regularly varying, $\Delta(c)\to 0$ as $c\to\infty$. However, in the regularly varying case, $\Delta(c)$ does not vanish and hence is not negligible. To address this limitation, we provide simple upper and lower bounds for $\Delta(c)$ and for the tube formula itself. Numerical studies are conducted to assess the accuracy of the asymptotic approximation.

math.PR

Constructing Simultaneous Confidence Bands for Errors-in-variables Curves with Application to the Lorenz Curve

Errors-in-variables curves are curves where errors exist not only in the independent variable but also in the dependent variable. We address the challenge of constructing simultaneous confidence bands (SCBs) for such curves. Our method finds application in the Lorenz curve, which represents the concentration of income or wealth. Unlike ordinary regression curves, the Lorenz curve incorporates errors in its explanatory variable and requires a fundamentally different treatment. To the best of our knowledge, the development of SCBs for such curves has not been explored in previous research. Using the Lorenz curve as a case study, this paper proposes a novel approach to address this challenge.

stat.AP

Integrated empirical measures and generalizations of classical goodness-of-fit statistics

Based on $m$-fold integrated empirical measures, we study three new classes of goodness-of-fits tests, generalizing Anderson-Darling, Cramér-von Mises, and Watson statistics, respectively, and examine the corresponding limiting stochastic processes. The limiting null distributions of the statistics all lead to explicitly solvable cases with closed-form expressions for the corresponding Karhunen-Loève expansions and covariance kernels. In particular, the eigenvalues are shown to be $\frac1{k(k+1)\cdots (k+2m-1)}$ for the generalized Anderson-Darling, $\frac1{(πk)^{2m}}$ for the generalized Cramér-von Mises, and $\frac1{2π\lceil k/2\rceil^{2m}}$ for the generalized Watson statistics, respectively. The infinite products of the resulting moment generating functions are further simplified to finite ones so as to facilitate efficient numerical calculations. These statistics are capable of detecting different features of the distributions and thus provide a useful toolbox for goodness-of-fit testing.

math.ST

On the Limits of Topological Data Analysis for Statistical Inference

Topological data analysis has emerged as a powerful tool for extracting the metric, geometric and topological features underlying the data as a multi-resolution summary statistic, and has found applications in several areas where data arises from complex sources. In this paper, we examine the use of topological summary statistics through the lens of statistical inference. We investigate necessary and sufficient conditions under which \textit{valid statistical inference} is possible using {topological summary statistics}. Additionally, we provide examples of models that demonstrate invariance with respect to topological summaries.

math.PR

EM Estimation of the B-Spline Copula with Penalized Pseudo-Likelihood Functions

The B-spline copula function is defined by a linear combination of elements of the normalized B-spline basis. We develop a modified EM algorithm, to maximize the penalized pseudo-likelihood function, wherein we use the smoothly clipped absolute deviation (SCAD) penalty function for the penalization term. We conduct simulation studies to demonstrate the stability of the proposed numerical procedure, show that penalization yields estimates with smaller mean-square errors when the true parameter matrix is sparse, and provide methods for determining tuning parameters and for model selection. We analyze as an example a data set consisting of birth and death rates from 237 countries, available at the website, ''Our World in Data,'' and we estimate the marginal density and distribution functions of those rates together with all parameters of our B-spline copula model.

stat.ME

Expected Euler characteristic method for the largest eigenvalue: (Skew-)orthogonal polynomial approach

The expected Euler characteristic (EEC) method is an integral-geometric method used to approximate the tail probability of the maximum of a random field on a manifold. Noting that the largest eigenvalue of a real-symmetric or Hermitian matrix is the maximum of the quadratic form of a unit vector, we provide EEC approximation formulas for the tail probability of the largest eigenvalue of orthogonally invariant random matrices of a large class. For this purpose, we propose a version of a skew-orthogonal polynomial by adding a side condition such that it is uniquely defined, and describe the EEC formulas in terms of the (skew-)orthogonal polynomials. In addition, for the classical random matrices (Gaussian, Wishart, and multivariate beta matrices), we analyze the limiting behavior of the EEC approximation as the matrix size goes to infinity under the so-called edge-asymptotic normalization. It is shown that the limit of the EEC formula approximates well the Tracy-Widom distributions in the upper tail area, as does the EEC formula when the matrix size is finite.

math.PR

Random eigenvalues of graphenes and the triangulation of plane

We analyse the numbers of closed paths of length $k\in\mathbb{N}$ on two important regular lattices: the hexagonal lattice (also called $\textit{graphene}$ in chemistry) and its dual triangular lattice. These numbers form a moment sequence of specific random variables connected to the distance of a position of a planar random flight (in three steps) from the origin. Here, we refer to such a random variable as a $\textit{random eigenvalue}$ of the underlying lattice. Explicit formulas for the probability density and characteristic functions of these random eigenvalues are given for both the hexagonal and the triangular lattice. Furthermore, it is proven that both probability distributions can be approximated by a functional of the random variable uniformly distributed on increasing intervals $[0,b]$ as $b\to\infty$. This yields a straightforward method to simulate these random eigenvalues without generating graphene and triangular lattice graphs. To demonstrate this approximation, we first prove a key integral identity for a specific series containing the third powers of the modified Bessel functions $I_n$ of $n$th order, $n\in\mathbb{Z}$. Such series play a crucial role in various contexts, in particular, in analysis, combinatorics, and theoretical physics.

math.SP

Asymptotic expansion of the expected Minkowski functional for isotropic central limit random fields

The Minkowski functionals, including the Euler characteristic statistics, are standard tools for morphological analysis in cosmology. Motivated by cosmic research, we examine the Minkowski functional of the excursion set for an isotropic central limit random field, the $k$-point correlation functions ($k$th order cumulants) of which have the same structure as that assumed in cosmic research. Using 3- and 4-point correlation functions, we derive the asymptotic expansions of the Euler characteristic density, which is the building block of the Minkowski functional. The resulting formula reveals the types of non-Gaussianity that cannot be captured by the Minkowski functionals. As an example, we consider an isotropic chi-square random field and confirm that the asymptotic expansion accurately approximates the true Euler characteristic density.

math.ST

Robust Persistence Diagrams using Reproducing Kernels

Persistent homology has become an important tool for extracting geometric and topological features from data, whose multi-scale features are summarized in a persistence diagram. From a statistical perspective, however, persistence diagrams are very sensitive to perturbations in the input space. In this work, we develop a framework for constructing robust persistence diagrams from superlevel filtrations of robust density estimators constructed using reproducing kernels. Using an analogue of the influence function on the space of persistence diagrams, we establish the proposed framework to be less sensitive to outliers. The robust persistence diagrams are shown to be consistent estimators in bottleneck distance, with the convergence rate controlled by the smoothness of the kernel. This, in turn, allows us to construct uniform confidence bands in the space of persistence diagrams. Finally, we demonstrate the superiority of the proposed approach on benchmark datasets.

math.ST

Adversarially Robust Topological Inference

The distance function to a compact set plays a crucial role in the paradigm of topological data analysis. In particular, the sublevel sets of the distance function are used in the computation of persistent homology -- a backbone of the topological data analysis pipeline. Despite its stability to perturbations in the Hausdorff distance, persistent homology is highly sensitive to outliers. In this work, we develop a framework of statistical inference for persistent homology in the presence of outliers. Drawing inspiration from recent developments in robust statistics, we propose a \textit{median-of-means} variant of the distance function (\textsf{MoM Dist}) and establish its statistical properties. In particular, we show that, even in the presence of outliers, the sublevel filtrations and weighted filtrations induced by \textsf{MoM Dist} are both consistent estimators of the true underlying population counterpart and exhibit near minimax-optimal performance in adversarial settings. Finally, we demonstrate the advantages of the proposed methodology through simulations and applications.

math.ST

The volume-of-tube method for Gaussian random fields with inhomogeneous variance

The tube method or the volume-of-tube method approximates the tail probability of the maximum of a smooth Gaussian random field with zero mean and unit variance. This method evaluates the volume of a spherical tube about the index set, and then transforms it to the tail probability. In this study, we generalize the tube method to a case in which the variance is not constant. We provide the volume formula for a spherical tube with a non-constant radius in terms of curvature tensors, and the tail probability formula of the maximum of a Gaussian random field with inhomogeneous variance, as well as its Laplace approximation. In particular, the critical radius of the tube is generalized for evaluation of the asymptotic approximation error. As an example, we discuss the approximation of the largest eigenvalue distribution of the Wishart matrix with a non-identity matrix parameter. The Bonferroni method is the tube method when the index set is a finite set. We provide the formula for the asymptotic approximation error for the Bonferroni method when the variance is not constant.

math.PR

Existence and Uniqueness of the Kronecker Covariance MLE

In matrix-valued datasets the sampled matrices often exhibit correlations among both their rows and their columns. A useful and parsimonious model of such dependence is the matrix normal model, in which the covariances among the elements of a random matrix are parameterized in terms of the Kronecker product of two covariance matrices, one representing row covariances and one representing column covariance. An appealing feature of such a matrix normal model is that the Kronecker covariance structure allows for standard likelihood inference even when only a very small number of data matrices is available. For instance, in some cases a likelihood ratio test of dependence may be performed with a sample size of one. However, more generally the sample size required to ensure boundedness of the matrix normal likelihood or the existence of a unique maximizer depends in a complicated way on the matrix dimensions. This motivates the study of how large a sample size is needed to ensure that maximum likelihood estimators exist, and exist uniquely with probability one. Our main result gives precise sample size thresholds in the paradigm where the number of rows and the number of columns of the data matrices differ by at most a factor of two. Our proof uses invariance properties that allow us to consider data matrices in canonical form, as obtained from the Kronecker canonical form for matrix pencils.

math.ST

Minkowski functionals and the nonlinear perturbation theory in the large-scale structure: second-order effects

The second-order formula of Minkowski functionals in weakly non-Gaussian fields is compared with the numerical $N$-body simulations. Recently, weakly non-Gaussian formula of Minkowski functionals is extended to include the second-order effects of non-Gaussianity in general dimensions. We apply this formula to the three-dimensional density field in the large-scale structure of the Universe. The parameters of the second-order formula include several kinds of skewness and kurtosis parameters. We apply the tree-level nonlinear perturbation theory to estimate these parameters. First we compare the theoretical values with those of numerical simulations on the basis of parameter values, and next we test the performance of the analytic formula combined with the perturbation theory. The second-order formula outperforms the first-order formula in general. The performance of the perturbation theory depends on the smoothing radius applied in defining the Minkowski functionals. The quantitative comparisons are presented in detail.

astro-ph.CO

Weakly non-Gaussian formula for the Minkowski functionals in general dimensions

The Minkowski functionals are useful statistics to quantify the morphology of various random fields. They have been applied to numerous analyses of geometrical patterns, including various types of cosmic fields, morphological image processing, etc. In some cases, including cosmological applications, small deviations from the Gaussianity of the distribution are of fundamental importance. Analytic formulas for the expectation values of Minkowski functionals with small non-Gaussianity have been derived in limited cases to date. We generalize these previous works to derive an analytic expression for expectation values of Minkowski functionals up to second-order corrections of non-Gaussianity in a space of general dimensions. The derived formula has sufficient generality to be applied to any random fields with weak non-Gaussianity in a statistically homogeneous and isotropic space of any dimensions.

astro-ph.CO

Computation of the Expected Euler Characteristic for the Largest Eigenvalue of a Real Non-central Wishart Matrix

We give an approximate formula for the distribution of the largest eigenvalue of real Wishart matrices by the expected Euler characteristic method for the general dimension. The formula is expressed in terms of a definite integral with parameters. We derive a differential equation satisfied by the integral for the $2 \times 2$ matrix case and perform a numerical analysis of it.

math.ST

Optimal experimental design that minimizes the width of simultaneous confidence bands

We propose an optimal experimental design for a curvilinear regression model that minimizes the band-width of simultaneous confidence bands. Simultaneous confidence bands for curvilinear regression are constructed by evaluating the volume of a tube about a curve that is defined as a trajectory of a regression basis vector (Naiman, 1986). The proposed criterion is constructed based on the volume of a tube, and the corresponding optimal design that minimizes the volume of tube is referred to as the tube-volume optimal (TV-optimal) design. For Fourier and weighted polynomial regressions, the problem is formalized as one of minimization over the cone of Hankel positive definite matrices, and the criterion to minimize is expressed as an elliptic integral. We show that the Möbius group keeps our problem invariant, and hence, minimization can be conducted over cross-sections of orbits. We demonstrate that for the weighted polynomial regression and the Fourier regression with three bases, the tube-volume optimal design forms an orbit of the Möbius group containing D-optimal designs as representative elements.

math.ST