SearcharxivSearch

arXiv subjects

Tomonari Sei

Publications and source records attributed to Tomonari Sei.

At least 19 recordsLinked to original sources

Minimum information Markov model

The analysis of high-dimensional time series data has become increasingly important across a wide range of fields. Recently, a method for constructing the minimum information Markov kernel on finite state spaces was established. In this study, we propose a statistical model based on a parametrization of its dependence function, which we call the \textit{Minimum Information Markov Model}. We show that its parametrization induces an orthogonal structure between the stationary distribution and the dependence function, and that the model arises as the optimal solution to a divergence rate minimization problem. In particular, for the Gaussian autoregressive case, we establish the existence of the optimal solution to this minimization problem, a nontrivial result requiring a rigorous proof. For parameter estimation, our approach exploits the conditional independence structure inherent in the model, which is supported by the orthogonality. Specifically, we develop several estimators, including conditional likelihood and pseudo likelihood estimators, for the minimum information Markov model in both univariate and multivariate settings. We demonstrate their practical performance through simulation studies and applications to real-world time series data.

stat.ME

Consistency of Nonparametric Density Estimators in CAT(0) Orthant Space

The inference of evolutionary histories is a central problem in evolutionary biology. The analysis of a sample of phylogenetic trees can be conducted in Billera-Holmes-Vogtmann tree space, which is a CAT(0) metric space of phylogenetic trees. The globally non-positively curved (CAT(0)) property of this space enables the extension of various statistical techniques. In the problem of nonparametric density estimation, two primary methods, kernel density estimation and log-concave maximum likelihood estimation, have been proposed, yet their theoretical properties remain largely unexplored. In this paper, we address this gap by proving the consistency of these estimators in a more general setting$\unicode{x2014}$CAT(0) orthant spaces, which include BHV tree space. We extend log-concave approximation techniques to this setting and establish consistency via the continuity of the log-concave projection map. We also modify the kernel density estimator to correct boundary bias and establish uniform consistency using empirical process theory.

math.ST

Shrinkage priors for circulant correlation structure models

We consider a new statistical model called the circulant correlation structure model, which is a multivariate Gaussian model with unknown covariance matrix and has a scale-invariance property. We construct shrinkage priors for the circulant correlation structure models and show that Bayesian predictive densities based on those priors asymptotically dominate Bayesian predictive densities based on Jeffreys priors under the Kullback-Leibler (KL) risk function. While shrinkage of eigenvalues of covariance matrices of Gaussian models has been successful, the proposed priors shrink a non-eigenvalue part of covariance matrices.

math.ST

Constructing Markov chains with given dependence and marginal stationary distributions

A method of constructing Markov chains on finite state spaces is provided. The chain is specified by three constraints: stationarity, dependence and marginal distributions. The generalized Pythagorean theorem in information geometry plays a central role in the construction. An algorithm for obtaining the desired Markov chain is described. Integer-valued autoregressive processes are considered for illustration.

math.ST

Relative local dependence of bivariate copulas

For a bivariate probability distribution, local dependence around a single point on the support is often formulated as the second derivative of the logarithm of the probability density function. However, this definition lacks the invariance under marginal distribution transformations, which is often required as a criterion for dependence measures. In this study, we examine the \textit{relative local dependence}, which we define as the ratio of the local dependence to the probability density function, for copulas. By using this notion, we point out that typical copulas can be characterised as the solutions to the corresponding partial differential equations, particularly highlighting that the relative local dependence of the Frank copula remains constant. The estimation and visualization of the relative local dependence are demonstrated using simulation data. Furthermore, we propose a class of copulas where local dependence is proportional to the $k$-th power of the probability density function, and as an example, we demonstrate a newly discovered relationship derived from the density functions of two representative copulas, the Frank copula and the Farlie-Gumbel-Morgenstern (FGM) copula.

stat.ME

Frank copula is minimum information copula under fixed Kendall's $τ$

In dependence modeling, various copulas have been utilized. Among them, the Frank copula has been one of the most typical choices due to its simplicity. In this work, we demonstrate that the Frank copula is the minimum information copula under fixed Kendall's $τ$ (MICK), both theoretically and numerically. First, we explain that both MICK and the Frank density follow the hyperbolic Liouville equation. Moreover, we show that the copula density satisfying the Liouville equation is uniquely the Frank copula. Our result asserts that selecting the Frank copula as an appropriate copula model is equivalent to using Kendall's $τ$ as the sole available information about the true distribution, based on the entropy maximization principle.

stat.ME

On the minimum information checkerboard copulas under fixed Kendall's rank correlation

Copulas have gained widespread popularity as statistical models to represent dependence structures between multiple variables in various applications. The minimum information copula, given a finite number of constraints in advance, emerges as the copula closest to the uniform copula when measured in Kullback-Leibler divergence. In prior research, the focus has predominantly been on constraints related to expectations on moments, including Spearman's $ρ$. This approach allows for obtaining the copula through convex programming. However, the existing framework for minimum information copulas does not encompass non-linear constraints such as Kendall's $τ$. To address this limitation, we introduce MICK, a novel minimum information copula under fixed Kendall's $τ$. We first characterize MICK by its local dependence property. Despite being defined as the solution to a non-convex optimization problem, we demonstrate that the uniqueness of this copula is guaranteed when the correlation is sufficiently small. Additionally, we provide numerical insights into applying MICK to real financial data.

stat.ME

Minimum information dependence modeling

We propose a method to construct a joint statistical model for mixed-domain data to analyze their dependence. Multivariate Gaussian and log-linear models are particular examples of the proposed model. It is shown that the functional equation defining the model has a unique solution under fairly weak conditions. The model is characterized by two orthogonal parameters: the dependence parameter and the marginal parameter. To estimate the dependence parameter, a conditional inference together with a sampling procedure is proposed and is shown to provide a consistent estimator. Illustrative examples of data analyses involving penguins and earthquakes are presented.

stat.ME

A branch cut approach to the probability density and distribution functions of a linear combination of central and non-central Chi-square random variables

The paper considers the distribution of a general linear combination of central and non-central chi-square random variables by exploring the branch cut regions that appear in the standard Laplace inversion process. Due to the original interest from the directional statistics, the focus of this paper is on the density function of such distributions and not on their cumulative distribution function. In fact, our results confirm that the latter is a special case of the former. Our approach provides new insight by generating alternative characterizations of the probability density function in terms of a finite number of feasible univariate integrals. In particular, the central cases seem to allow an interesting representation in terms of the branch cuts, while general degrees of freedom and non-centrality can be easily adopted using recursive differentiation. Numerical results confirm that the proposed approach works well while more transparency and therefore easier control in the accuracy is ensured.

stat.CO

Maximum Likelihood Estimation of Log-Concave Densities on Tree Space

Phylogenetic trees are key data objects in biology, and the method of phylogenetic reconstruction has been highly developed. The space of phylogenetic trees is a nonpositively curved metric space. Recently, statistical methods to analyze the set of trees on this space are being developed utilizing this property. Meanwhile, in Euclidean space, the log-concave maximum likelihood method has emerged as a new nonparametric method for probability density estimation. In this paper, we derive a sufficient condition for the existence and uniqueness of the log-concave maximum likelihood estimator on tree space. We also propose an estimation algorithm for one and two dimensions. Since various factors affect the inferred trees, it is difficult to specify the distribution of sample trees. The class of log-concave densities is nonparametric, and yet the estimation can be conducted by the maximum likelihood method without selecting hyperparameters. We compare the estimation performance with a previously developed kernel density estimator numerically. In our examples where the true density is log-concave, we demonstrate that our estimator has a smaller integrated squared error when the sample size is large. We also conduct numerical experiments of clustering using the Expectation-Maximization (EM) algorithm and compare the results with k-means++ clustering using Fréchet mean.

stat.ME

A proper scoring rule for minimum information copulas

Multi-dimensional distributions whose marginal distributions are uniform are called copulas. Among them, the one that satisfies given constraints on expectation and is closest to the independent distribution in the sense of Kullback-Leibler divergence is called the minimum information copula. The density function of the minimum information copula contains a set of functions called the normalizing functions, which are often difficult to compute. Although a number of proper scoring rules for probability distributions having normalizing constants such as exponential families are proposed, these scores are not applicable to the minimum information copulas due to the normalizing functions. In this paper, we propose the conditional Kullback-Leibler score, which avoids computation of the normalizing functions. The main idea of its construction is to use pairs of observations. We show that the proposed score is strictly proper in the space of copula density functions and therefore the estimator derived from it has asymptotic consistency. Furthermore, the score is convex with respect to the parameters and can be easily optimized by the gradient methods.

stat.ME

Improving Randomization Tests under Interference Based on Power Analysis

In causal inference, we can consider a situation in which treatment on one unit affects others, i.e., interference exists. In the presence of interference, we cannot perform a classical randomization test directly because a null hypothesis is not sharp. Instead, we need to perform the randomization test restricted to a subset of units and assignments that makes the null hypothesis sharp. A previous study constructed a useful testing method, a biclique test, by reducing the selection of the appropriate subsets to searching for bicliques in a bipartite graph. However, since the power depends on the features of selected subsets, there is still room to improve the power by refining the selection procedure. In this paper, we propose a method to improve the biclique test based on a power evaluation of the randomization test. We explicitly derived an expression for the power of the randomization test under several assumptions and found that a certain quantity calculated from a given assignment set characterizes the power. Based on this fact, we propose a method to improve the power of the biclique test by modifying the selection rule for subsets of units and assignments. Through a simulation with a spatial interference setting, we confirm that the proposed method has higher power than the existing method.

stat.ME

Holonomic extended least angle regression

One of the main problems studied in statistics is the fitting of models. Ideally, we would like to explain a large dataset with as few parameters as possible. There have been numerous attempts at automatizing this process. Most notably, the Least Angle Regression algorithm, or LARS, is a computationally efficient algorithm that ranks the covariates of a linear model. The algorithm is further extended to a class of distributions in the generalized linear model by using properties of the manifold of exponential families as dually flat manifolds. However this extension assumes that the normalizing constant of the joint distribution of observations is easy to compute. This is often not the case, for example the normalizing constant may contain a complicated integral. We circumvent this issue if the normalizing constant satisfies a holonomic system, a system of linear partial differential equations with a finite-dimensional space of solutions. In this paper we present a modification of the holonomic gradient method and add it to the extended LARS algorithm. We call this the holonomic extended least angle regression algorithm, or HELARS. The algorithm was implemented using the statistical software R, and was tested with real and simulated datasets.

stat.CO

Inconsistency of diagonal scaling under high-dimensional limit: a replica approach

In this note, we claim that diagonal scaling of a sample covariance matrix is asymptotically inconsistent if the ratio of the dimension to the sample size converges to a positive constant, where population is assumed to be Gaussian with a spike covariance model. Our non-rigorous proof relies on the replica method developed in statistical physics. In contrast to similar results known in literature on principal component analysis, the strong inconsistency is not observed. Numerical experiments support the derived formulas.

math.ST

Calculating the normalising constant of the Bingham distribution on the sphere using the holonomic gradient method

In this paper we implement the holonomic gradient method to exactly compute the normalising constant of Bingham distributions. This idea is originally applied for general Fisher-Bingham distributions in Nakayama et al. (2011). In this paper we explicitly apply this algorithm to show the exact calculation of the normalising constant; derive explicitly the Pfaffian system for this parametric case; implement the general approach for the maximum likelihood solution search and finally adjust the method for degenerate cases, namely when the parameter values have multiplicities.

stat.CO

Infinitely imbalanced binomial regression and deformed exponential families

The logistic regression model is known to converge to a Poisson point process model if the binary response tends to infinitely imbalanced. In this paper, it is shown that this phenomenon is universal in a wide class of link functions on binomial regression. The proof relies on the extreme value theory. For the logit, probit and complementary log-log link functions, the intensity measure of the point process becomes an exponential family. For some other link functions, deformed exponential families appear. A penalized maximum likelihood estimator for the Poisson point process model is suggested.

math.ST

Properties and applications of Fisher distribution on the rotation group

We study properties of Fisher distribution (von Mises-Fisher distribution, matrix Langevin distribution) on the rotation group SO(3). In particular we apply the holonomic gradient descent, introduced by Nakayama et al. (2011), and a method of series expansion for evaluating the normalizing constant of the distribution and for computing the maximum likelihood estimate. The rotation group can be identified with the Stiefel manifold of two orthonormal vectors. Therefore from the viewpoint of statistical modeling, it is of interest to compare Fisher distributions on these manifolds. We illustrate the difference with an example of near-earth objects data.

stat.ME