SearcharxivSearch

arXiv subjects

Yo Sheena

Publications and source records attributed to Yo Sheena.

13 recordsLinked to original sources

Convergence of Estimative Density to Information Projection for Misspecified Normal Distribution Model

This paper investigates the convergence of an estimative multivariate normal density when the true distribution is a misspecified multivariate t-distribution. The statistical model is the family of k-dimensional normal distributions (N_k(\mu,\Sigma)), whereas the observations are assumed to follow (t_k(0,I_k,\nu)), with (\nu>6). The information projection of the true distribution onto the normal model is first identified as the normal distribution with mean zero and covariance matrix (\nu/(\nu-2)I_k). The main objective is to evaluate the expected Kullback-Leibler divergence between this information projection and the normal density obtained by substituting the maximum likelihood estimator into the model. Using a general asymptotic expansion for estimative densities, the paper derives explicit first- and second-order terms of the risk as functions of the sample size (n), the dimension (k), and the degrees of freedom (\nu). To obtain the second-order term, the paper calculates the required moments, information matrices, and higher-order cumulants under both the multivariate normal and multivariate t-distributions. In particular, the complicated third- and fourth-order cumulants involving quadratic sufficient statistics are classified according to their index patterns, and their values and multiplicities are systematically derived.

stat.ME

Criterion for the resemblance between the mother and the model distribution

If the probability distribution model aims to approximate the hidden mother distribution, it is imperative to establish a useful criterion for the resemblance between the mother and the model distributions. This study proposes a criterion that measures the Hellinger distance between discretized (quantized) samples from both distributions. Unlike information criteria such as AIC, this criterion does not require the probability density function of the model distribution, which cannot be explicitly obtained for a complicated model such as a deep learning machine. Second, it can draw a positive conclusion (i.e., both distributions are sufficiently close) under a given threshold, whereas a statistical hypothesis test, such as the Kolmogorov-Smirnov test, cannot genuinely lead to a positive conclusion when the hypothesis is accepted. In this study, we establish a reasonable threshold for the criterion deduced from the Bayes error rate and also present the asymptotic bias of the estimator of the criterion. From these results, a reasonable and easy-to-use criterion is established that can be directly calculated from the two sets of samples from both distributions.

math.ST

MLE convergence speed to information projection of exponential family: Criterion for model dimension and sample size -- complete proof version--

For a parametric model of distributions, the closest distribution in the model to the true distribution located outside the model is considered. Measuring the closeness between two distributions with the Kullback-Leibler (K-L) divergence, the closest distribution is called the "information projection." The estimation risk of the maximum likelihood estimator (MLE) is defined as the expectation of K-L divergence between the information projection and the predictive distribution with plugged-in MLE. Here, the asymptotic expansion of the risk is derived up to $n^{-2}$-order, and the sufficient condition on the risk for the Bayes error rate between the true distribution and the information projection to be lower than a specified value is investigated. Combining these results, the "$p-n$ criterion" is proposed, which determines whether the MLE is sufficiently close to the information projection for the given model and sample. In particular, the criterion for an exponential family model is relatively simple and can be used for a complex model with no explicit form of normalizing constant. This criterion can constitute a solution to the sample size or model acceptance problem. Use of the $p-n$ criteria is demonstrated for two practical datasets. The relationship between the results and information criteria is also studied.

math.ST

Efficiency of maximum likelihood estimation for a multinomial distribution with known probability sums

For a multinomial distribution, suppose that we have prior knowledge of the sum of the probabilities of some categories. This allows us to construct a submodel in a full (i.e., no-restriction) model. Maximum likelihood estimation (MLE) under this submodel is expected to have better estimation efficiency than MLE under the full model. This article presents the asymptotic expansion of the risk of MLE with respect to Kullback--Leibler divergence for both the full model and submodel. The results reveal that, using the submodel, the reduction of the risk is quite small in some cases. Furthermore, when the sample size is small, the use of the subomodel can increase the risk.

math.ST

Asymptotic efficiency of M.L.E. using prior survey in multinomial distributions

Incorporating information from a prior survey is generally supposed to decrease the estimation risk of the present survey. This paper aims to show how the risk changes by incorporating the information of a prior survey through watching the first and the second-order terms of the asymptotic expansion of the risk. We recognize that the prior information is of some help for risk reduction when we can acquire samples of a sufficient size for both surveys. Interestingly, when the sample size of the present survey is small, the use of the prior survey can increase the risk. In other words, blending information from both surveys can have a negative effect on the risk. Based on these observations, we give some suggestions on whether or not to use the results of the prior survey and the sample size to use in the surveys for a reliable estimation.

math.ST

Estimation of a Continuous Distribution on a Real Line by Discretization Methods -- Complete Version--

For an unknown continuous distribution on a real line, we consider the approximate estimation by the discretization. There are two methods for the discretization. First method is to divide the real line into several intervals before taking samples ("fixed interval method") . Second method is dividing the real line using the estimated percentiles after taking samples ("moving interval method"). In either way, we settle down to the estimation problem of a multinomial distribution. We use (symmetrized) $f$-divergence in order to measure the discrepancy of the true distribution and the estimated one. Our main result is the asymptotic expansion of the risk (i.e. expected divergence) up to the second-order term in the sample size. We prove theoretically that the moving interval method is asymptotically superior to the fixed interval method. We also observe how the presupposed intervals (fixed interval method) or percentiles (moving interval method) affect the asymptotic risk.

math.ST

Asymptotic Expansion of Risk for a Regression Model with respect to $\alpha$-Divergence with an Application to the Sample Size Problem -- Complete Version

For a regression model, we consider the risk of the maximum likelihood estimator with respect to $\alpha$-divergence, which includes the special cases of Kullback-Leibler divergence, Hellinger distance and $\chi^2$ divergence. The asymptotic expansion of the risk with respect to the sample size $n$ is given up to the order $n^{-2}$. We are interested in how the risk convergence speed (to zero) is affected by the error term distributions of the regression model and the magnitude of the joint moments of the standardized explanatory variables. Besides the general result (which is given by Mathematica program), we consider three concrete error term distributions; a normal distribution, a t-distribution and a skew-normal distribution. We use the (approximated) risk of m.l.e. as a measure of the difficulty of estimation for the regression model. Especially comparing the value of the (approximated) risk with that of a binomial distribution, we can give a certain standard for the sample size required to estimate the regression model.

math.ST

Asymptotic expansion of the risk of maximum likelihood estimator with respect to $\alpha$-divergence as a measure of the difficulty of specifying a parametric model -- with detailed proof

For a given parametric probability model, we consider the risk of the maximum likelihood estimator with respect to $\alpha$-divergence, which includes the special cases of Kullback--Leibler divergence, the Hellinger distance and $\chi^2$ divergence. The asymptotic expansion of the risk is given with respect to sample sizes of up to order $n^{-2}$. Each term in the expansion is expressed with the geometrical properties of the Riemannian manifold formed by the parametric probability model. We attempt to measure the difficulty of specifying a model through this expansion.

math.ST

Inference on the eigenvalues of the covariance matrix of a multivariate normal distribution--geometrical view--

We consider an inference on the eigenvalues of the covariance matrix of a multivariate normal distribution. The family of multivariate normal distributions with a fixed mean is seen as a Riemannian manifold with Fisher information metric. Two submanifolds naturally arises; one is the submanifold given by fixed eigenvectors of the covariance matrix, the other is the one given by fixed eigenvalues. We analyze the geometrical structures of these manifolds such as metric, embedding curvature under $e$-connection or $m$-connection. Based on these results, we study 1) the bias of the sample eigenvalues, 2)the information loss caused by neglecting the sample eigenvectors, 3)new estimators that are naturally derived from the geometrical view.

math.ST

Modified estimator of the contribution rates of population eigenvalues

Modified estimators for the contribution rates of population eigenvalues are given under an elliptically contoured distribution. These estimators decrease the bias of the classical estimator, i.e. the sample contribution rates. The improvement of the modified estimators over the classical estimator are proved theoretically in view of their risks. We also checked numerically that the drawback of the classical estimator, namely the underestimation of the dimension in principal component analysis or factor analysis, are corrected in the modification.

math.ST

Inference on Eigenvalues of Wishart Distribution Using Asymptotics with respect to the Dispersion of Population Eigenvalues

In this paper we derive some new and practical results on testing and interval estimation problems for the population eigenvalues of a Wishart matrix based on the asymptotic theory for block-wise infinite dispersion of the population eigenvalues. This new type of asymptotic theory has been developed by the present authors in Takemura and Sheena (2005) and Sheena and Takemura (2007a,b) and in these papers it was applied to point estimation problem of population covariance matrix in a decision theoretic framework. In this paper we apply it to some testing and interval estimation problems. We show that the approximation based on this type of asymptotics is generally much better than the traditional large-sample asymptotics for the problems.

math.ST

Asymptotic Distribution of Wishart Matrix for Block-wise Dispersion of Population Eigenvalues

This paper deals with the asymptotic distribution of Wishart matrix and its application to the estimation of the population matrix parameter when the population eigenvalues are block-wise infinitely dispersed. We show that the appropriately normalized eigenvectors and eigenvalues asymptotically generate two Wishart matrices and one normally distributed random matrix, which are mutually independent. For a family of orthogonally equivariant estimators, we calculate the asymptotic risks with respect to the entropy or the quadratic loss function and derive the asymptotically best estimator among the family. We numerically show 1) the convergence in both the distributions and the risks are quick enough for a practical use, 2) the asymptotically best estimator is robust against the deviation of the population eigenvalues from the block-wise infinite dispersion.

math.ST