SearcharxivSearch

arXiv subjects

Koji Tsukuda

Publications and source records attributed to Koji Tsukuda.

17 recordsLinked to original sources

Asymptotic bias of the plug-in Shannon entropy estimator under a regularly varying occupancy model

Estimating the Shannon entropy of discrete distributions with countably infinite support is a challenging problem. In this paper, we investigate the bias of the plug-in estimator $\hat{H}_n$ for the Shannon entropy $H(\boldsymbol{p})$ under an occupancy model whose frequency sequence exhibits regular variation with tail index $α\in(0,1)$. Using Poissonization and the theory of regular variation, we establish the asymptotic relation $|\mathsf{E}[\hat H_n] - H(\boldsymbol{p})| \sim n^{α-1}L(n)C_α$, where $L$ is a slowly varying function and $C_α$ is an explicit constant depending only on $α$ that admits an integral representation. Our result shows that the asymptotic behavior of the bias of the plug-in estimator under power-law frequency distributions is determined by the tail behavior of the underlying distribution.

math.PR

Hybrid principal component analysis in multivariate allometric regression

In biological data from allometry studies, the largest eigenvalue is typically dominant, and the gaps between minor eigenvalues are often narrow. Such proximity among small minor eigenvalues can lead to instability in statistics based on their corresponding eigenvectors. This study derives the asymptotic normality of the hybrid principal component analysis estimator of the leading principal eigenvector in the multivariate allometric regression model and proposes a test based on a geometric statistic for the parallelism between the regression direction and the principal component direction that avoids this instability. Using the hybrid principal component analysis framework, we analyze the well-known painted turtle carapace data and confirm previously reported results on the allometric extension relationship between female and male turtles.

stat.ME

High-dimensional linear regression inference via $\ell^2$ weak convergence

We prove weak convergence in a separable Hilbert space for estimators of high-dimensional regression coefficients, which yields asymptotic normality and enables direct use of standard asymptotic tools such as the continuous mapping theorem. The approach permits diverging sparsity with many small nonzero coefficients, while requiring that only finitely many have moderate magnitude. As applications, we develop a test for finitely many linear hypotheses and, via a Scheffé-type approach, simultaneous inference for infinitely many linear hypotheses, yielding both a global test and simultaneous confidence bands for the regression function. The limiting distributions are given by weighted sums of independent chi-squared variables, and plug-in critical values achieve asymptotically correct size.

math.ST

On spectral clustering under non-isotropic Gaussian mixture models

We evaluate the misclustering probability of a spectral clustering algorithm under a Gaussian mixture model with a general covariance structure. The algorithm partitions the data into two groups based on the sign of the first principal component score. As a corollary of the main result, the clustering procedure is shown to be consistent in a high-dimensional regime.

math.ST

Estimating the Shannon Entropy Using the Pitman--Yor Process

The Shannon entropy is a fundamental measure for quantifying diversity and model complexity in fields such as information theory, ecology, and genetics. However, many existing studies assume that the number of species is known, an assumption that is often unrealistic in practice. In recent years, efforts have been made to relax this restriction. Motivated by these developments, this study proposes an entropy estimation method based on the Pitman--Yor process, a representative approach in Bayesian nonparametrics. By approximating the true distribution as an infinite-dimensional process, the proposed method enables stable estimation even when the number of observed species is smaller than the true number of species. This approach provides a principled way to deal with the uncertainty in species diversity and enhances the reliability and robustness of entropy-based diversity assessment. In addition, we investigate the convergence property of the Shannon entropy for regularly varying distributions and use this result to establish the consistency of the proposed estimator. Finally, we demonstrate the effectiveness of the proposed method through numerical experiments.

stat.ME

Robust $M$-Estimation of Scatter Matrices via Precision Structure Shrinkage

Maronna's and Tyler's $M$-estimators are among the most widely used robust estimators for scatter matrices. However, when the dimension of observations is relatively high, their performance can substantially deteriorate in certain situations, particularly in the presence of clustered outliers. To address this issue, we propose an estimator that shrinks the estimated precision matrix toward the identity matrix. We derive a sufficient condition for its existence, discuss its statistical interpretation, and establish upper and lower bounds for its additive finite sample breakdown point. Numerical experiments confirm the robustness of the proposed method.

stat.ME

Equality between two general ridge estimators and equivalence of their residual sums of squares

General ridge estimators are typical linear estimators in a general linear model. The class of them includes some shrinkage estimators in addition to classical linear unbiased estimators such as the ordinary least squares estimator and the weighted least squares estimator. We derive necessary and sufficient conditions under which two general ridge estimators coincide. In particular, two noteworthy conditions are added to those from previous studies. The first condition is given as a seemingly column space relationship to the covariance matrix of the error term, and the second one is based on the biases of general ridge estimators. Another problem studied in this paper is to derive an equivalence condition such that equality between two residual sums of squares holds when general ridge estimators are considered.

math.ST

Two step estimations via the Dantzig selector for models of stochastic processes with high-dimensional parameters

We consider the sparse estimation for stochastic processes with possibly infinite-dimensional nuisance parameters, by using the Dantzig selector which is a sparse estimation method similar to $Z$-estimation. When a consistent estimator for a nuisance parameter is obtained, it is possible to construct an asymptotically normal estimator for the parameter of interest under appropriate conditions. Motivated by this fact, we establish the asymptotic behavior of the Dantzig selector for models of ergodic stochastic processes with high-dimensional parameters of interest and possibly infinite-dimensional nuisance parameters. Moreover, we construct an asymptotically normal estimator by the two step estimation with help of the variable selection through the Dantzig selector and a consistent estimator of the nuisance parameter. Applications to ergodic time series models including integer-valued autoregressive models and ergodic diffusion processes are presented.

math.ST

Estimators for multivariate allometric regression model

In a regression model with multiple response variables and multiple explanatory variables, if the difference of the mean vectors of the response variables for different values of explanatory variables is always in the direction of the first principal eigenvector of the covariance matrix of the response variables, then it is called a multivariate allometric regression model. This paper studies the estimation of the first principal eigenvector in the multivariate allometric regression model. A class of estimators that includes conventional estimators is proposed based on weighted sum-of-squares matrices of regression sum-of-squares matrix and residual sum-of-squares matrix. We establish an upper bound of the mean squared error of the estimators contained in this class, and the weight value minimizing the upper bound is derived. Sufficient conditions for the consistency of the estimators are discussed in weak identifiability regimes under which the difference of the largest and second largest eigenvalues of the covariance matrix decays asymptotically and in ``large $p$, large $n$" regimes, where $p$ is the number of response variables and $n$ is the sample size. Several numerical results are also presented.

math.ST

Spectral clustering algorithm for the allometric extension model

The spectral clustering algorithm is often used as a binary clustering method for unclassified data by applying the principal component analysis. To study theoretical properties of the algorithm, the assumption of conditional homoscedasticity is often supposed in existing studies. However, this assumption is restrictive and often unrealistic in practice. Therefore, in this paper, we consider the allometric extension model, that is, the directions of the first eigenvectors of two covariance matrices and the direction of the difference of two mean vectors coincide, and we provide a non-asymptotic bound of the error probability of the spectral clustering algorithm for the allometric extension model. As a byproduct of the result, we obtain the consistency of the clustering method in high-dimensional settings.

math.ST

Evaluating moments of length of Pitman partition

The Pitman sampling formula has been intensively studied as a distribution of random partitions. One of the objects of interest is the length $K (= K_{n,θ,α})$ of a random partition that follows the Pitman sampling formula, where $n\in\mathbb{N}$, $α\in(0,\infty)$ and $θ> -α$ are parameters. This paper presents asymptotic evaluations of its $r$-th moment $\mathsf{E}[K^r]$ ($r=1,2,\ldots$) under two asymptotic regimes. In particular, the goals of this study are to provide a finer approximate evaluation of $\mathsf{E}[K^r]$ as $n\to\infty$ than has previously been developed and to provide an approximate evaluation of $\mathsf{E}[K^r]$ as the parameters $n$ and $θ$ simultaneously tend to infinity with $θ/n \to 0$. The results presented in this paper will provide a more accurate understanding of the asymptotic behavior of $K$.

math.PR

Limit theorem associated with Wishart matrices with application to hypothesis testing for common principal components

This study derives a new property of the Wishart distribution when the degree-of-freedom and the size of the matrix parameter of the distribution grow simultaneoulsy. Particularly, the asymptotic normality of the product of four independent Wishart matrices is shown under a high dimensional asymptotic regime. As an application of the result, a statistical test procedure for the common principal components hypothesis is proposed. For this problem, the proposed test statistic is asymptotically normal under the null hypothesis. In addition, the proposed test statistic diverges to positive infinity in probability under the alternative hypothesis.

math.ST

Error bounds for the normal approximation to the length of a Ewens partition

Let $K(=K_{n,θ})$ be a positive integer-valued random variable whose distribution is given by ${\rm P}(K = x) = \bar{s}(n,x) θ^x/(θ)_n$ $(x=1,\ldots,n) $, where $θ$ is a positive number, $n$ is a positive integer, $(θ)_n=θ(θ+1)\cdots(θ+n-1)$ and $\bar{s}(n,x)$ is the coefficient of $θ^x$ in $(θ)_n$ for $x=1,\ldots,n$. This formula describes the distribution of the length of a Ewens partition, which is a standard model of random partitions. As $n$ tends to infinity, $K$ asymptotically follows a normal distribution. Moreover, as $n$ and $θ$ simultaneously tend to infinity, if $n^2/θ\to\infty$, $K$ also asymptotically follows a normal distribution. In this paper, error bounds for the normal approximation are provided. The result shows that the decay rate of the error changes due to asymptotic regimes.

math.ST

Weak convergences of marked empirical processes in a Hilbert space and their applications

In this paper, weak convergences of marked empirical processes in $L^2(\mathbb{R},ν)$ and their applications to statistical goodness-of-fit tests are provided, where $L^2(\mathbb{R},ν)$ is the set of equivalence classes of the square integrable functions on $\mathbb{R}$ with respect to a finite Borel measure $ν$. The results obtained in our framework of weak convergences are, in the topological sense, weaker than those in the Skorokhod topology on a space of cádlág functions or the uniform topology on a space of bounded functions, which have been well studied in previous works. However, our results have the following merits: (1) avoiding conditions which do not suit for our purpose; (2) treating a weight function which makes us possible to propose an Anderson--Darling type test statistics for goodness-of-fit tests. Indeed, the applications presented in this paper are novel.

math.ST

A reversal phenomenon in estimation based on multiple samples from the Poisson--Dirichlet distribution

Consider two forms of sampling from a population: (i) drawing $s$ samples of $n$ elements with replacement and (ii) drawing a single sample of $ns$ elements. In this paper, under the setting where the descending order population frequency follows the Poisson--Dirichlet distribution with parameter $θ$, we report that the magnitude relation of the Fisher information, which sample partitions converted from samples (i) and (ii) possess, can change depending on the parameters, $n$, $s$, and $θ$. Roughly speaking, if $θ$ is small relative to $n$ and $s$, the Fisher information of (i) is larger than that of (ii); on the contrary, if $θ$ is large relative to $n$ and $s$, the Fisher information of (ii) is larger than that of (i). The result represents one aspect of random distributions.

math.ST

Covariance structure associated with an equality between two general ridge estimators

In a general linear model, this paper derives a necessary and sufficient condition under which two general ridge estimators coincide with each other. The condition is given as a structure of the dispersion matrix of the error term. Since the class of estimators considered here contains linear unbiased estimators such as the ordinary least squares estimator and the best linear unbiased estimator, our result can be viewed as a generalization of the well-known theorems on the equality between these two estimators, which have been fully studied in the literature. Two related problems are also considered: equality between two residual sums of squares, and classification of dispersion matrices by a perturbation approach.

math.ST

On Poisson approximations for the Ewens sampling formula when the mutation parameter grows with the sample size

The Ewens sampling formula was firstly introduced in the context of population genetics by Warren John Ewens in 1972, and has appeared in a lot of other scientific fields. There are abundant approximation results associated with the Ewens sampling formula especially when one of the parameters, the sample size $n$ or the mutation parameter $θ$ which denotes the scaled mutation rate, tends to infinity while the other is fixed. By contrast, the case that $θ$ grows with $n$ has been considered in a relatively small number of works, although this asymptotic setup is also natural. In this paper, when $θ$ grows with $n$, we advance the study concerning the asymptotic properties of the total number of alleles and of the counts of components in the allelic partition assuming the Ewens sampling formula from the viewpoint of Poisson approximations.

math.PR