SearcharxivSearch

arXiv subjects

Fumiyasu Komaki

Publications and source records attributed to Fumiyasu Komaki.

At least 19 recordsLinked to original sources

Matrix norm shrinkage estimators and priors

We develop a class of minimax estimators for a normal mean matrix under the Frobenius loss, which generalizes the James--Stein and Efron--Morris estimators. It shrinks the Schatten norm towards zero and works well for low-rank matrices. We also propose a class of superharmonic priors based on the Schatten norm, which generalizes Stein's prior and the singular value shrinkage prior. The generalized Bayes estimators and Bayesian predictive densities with respect to these priors are minimax. We examine the performance of the proposed estimators and priors in simulation.

math.ST

Improved nearly minimax prediction for independent Poisson processes under Kullback-Leibler loss

Simultaneous predictive distributions for independent Poisson observables are investigated, and the performance of predictive distributions is evaluated using the Kullback-Leibler (K-L) loss. This study introduces intuitive sufficient conditions, based on superharmonicity of priors, to improve the Bayesian predictive distribution based on the Jeffreys prior. The sufficient conditions exhibit a certain analogy with those known for the multivariate normal distribution. Additionally, this study examines the case where the observed data and target variables to be predicted are independent Poisson processes with different durations. Examples that satisfy the sufficient conditions are provided, including point and subspace shrinkage priors. The K-L risk of the improved predictions is demonstrated to be less than 1.04 times a minimax lower bound.

math.ST

Bayesian Prediction in Gamma Models: Admissibility and Infinitesimal Prediction

We study estimation and prediction in the Gamma model $\mathrm{Ga}(α,β)$, where the shape parameter $α$ is known and the scale parameter $β$ is unknown, under the Kullback--Leibler loss. For $α\le1$, all scale-invariant estimators of $β$ have infinite risk, indicating a qualitative change in the estimation problem at the boundary $α=1$. Our main result is that the Bayesian predictive density based on the Jeffreys prior is admissible for all $α>0$. This resolves the admissibility problem for Bayesian predictive densities in Gamma models. As a related result, we also establish the admissibility of the corresponding Bayesian estimator for $α>1$. To prove the predictive admissibility result, we develop an infinitesimal prediction framework based on Gamma processes. This framework naturally leads to a Kullback--Leibler loss for Lévy densities and establishes a connection between predictive distributions and Lévy measures. Under the resulting loss, the Bayesian predictive Lévy density is shown to be the posterior mean Lévy density. Unlike the normal and Poisson models, infinitesimal prediction in the Gamma model does not reduce to parameter estimation. Instead, it reduces to the estimation of a Lévy density. We relate this phenomenon to mean mixture curvature and discuss it from an information-geometric viewpoint.

math.ST

Minimization of Functions on Dually Flat Spaces Using Geodesic Descent Based on Dual Connections

We propose geodesic-based optimization methods on dually flat spaces, where the geometric structure of the parameter manifold is closely related to the form of the objective function. A primary application is maximum likelihood estimation in statistical models, especially exponential families, whose model manifolds are dually flat. We show that an m-geodesic update, which directly optimizes the log-likelihood, can theoretically reach the maximum likelihood estimator in a single step. In contrast, an e-geodesic update has a practical advantage in cases where the parameter space is geodesically complete, allowing optimization without explicitly handling parameter constraints. We establish the theoretical properties of the proposed methods and validate their effectiveness through numerical experiments.

stat.CO

Asymptotic mixed normality of maximum likelihood estimator for Ewens--Pitman partition

This paper investigates the asymptotic properties of parameter estimation for the Ewens--Pitman partition with parameters $0<α<1$ and $θ>-α$. Especially, we show that the maximum likelihood estimator (MLE) of $α$ is $n^{α/2}$-consistent and converges to a variance mixture of normal distributions, where the variance is governed by the Mittag-Leffler distribution. Moreover, we show that a proper normalization involving a random statistic eliminates the randomness in the variance. Building on this result, we construct an approximate confidence interval for $α$. Our proof relies on a stable martingale central limit theorem, which is of independent interest.

math.ST

Double shrinkage priors for a normal mean matrix

We consider estimation of a normal mean matrix under the Frobenius loss. Motivated by the Efron--Morris estimator, a generalization of Stein's prior has been recently developed, which is superharmonic and shrinks the singular values towards zero. The generalized Bayes estimator with respect to this prior is minimax and dominates the maximum likelihood estimator. However, here we show that it is inadmissible by using Brown's condition. Then, we develop two types of priors that provide improved generalized Bayes estimators and examine their performance numerically. The proposed priors attain risk reduction by adding scalar shrinkage or column-wise shrinkage to singular value shrinkage. Parallel results for Bayesian predictive densities are also given.

math.ST

BHGNN-RT: Network embedding for directed heterogeneous graphs

Networks are one of the most valuable data structures for modeling problems in the real world. However, the most recent node embedding strategies have focused on undirected graphs, with limited attention to directed graphs, especially directed heterogeneous graphs. In this study, we first investigated the network properties of directed heterogeneous graphs. Based on network analysis, we proposed an embedding method, a bidirectional heterogeneous graph neural network with random teleport (BHGNN-RT), for directed heterogeneous graphs, that leverages bidirectional message-passing process and network heterogeneity. With the optimization of teleport proportion, BHGNN-RT is beneficial to overcome the over-smoothing problem. Extensive experiments on various datasets were conducted to verify the efficacy and efficiency of BHGNN-RT. Furthermore, we investigated the effects of message components, model layer, and teleport proportion on model performance. The performance comparison with all other baselines illustrates that BHGNN-RT achieves state-of-the-art performance, outperforming the benchmark methods in both node classification and unsupervised clustering tasks.

cs.LG

On High-Dimensional Asymptotic Properties of Model Averaging Estimators

When multiple models are considered in regression problems, the model averaging method can be used to weigh and integrate the models. In the present study, we examined how the goodness-of-prediction of the estimator depends on the dimensionality of explanatory variables when using a generalization of the model averaging method in a linear model. We specifically considered the case of high-dimensional explanatory variables, with multiple linear models deployed for subsets of these variables. Consequently, we derived the optimal weights that yield the best predictions. we also observe that the double-descent phenomenon occurs in the model averaging estimator. Furthermore, we obtained theoretical results by adapting methods such as the random forest to linear regression models. Finally, we conducted a practical verification through numerical experiments.

math.ST

Predictive densities for multivariate normal models based on extended models and shrinkage Bayes methods

We investigate predictive densities for multivariate normal models with unknown mean vectors and known covariance matrices. Bayesian predictive densities based on shrinkage priors often have complex representations, although they are effective in various problems. We consider extended normal models with mean vectors and covariance matrices as parameters, and adopt predictive densities that belong to the extended models including the original normal model. We adopt predictive densities that are optimal with respect to the posterior Bayes risk in the extended models. The proposed predictive density based on a superharmonic shrinkage prior is shown to dominate the Bayesian predictive density based on the uniform prior under a loss function based on the Kullback-Leibler divergence. Our method provides an alternative to the empirical Bayes method, which is widely used to construct tractable predictive densities.

stat.ME

Enriched standard conjugate priors and the right invariant prior for Wishart distributions

The prediction of the variance-covariance matrix of the multivariate normal distribution is important in the multivariate analysis. We investigated Bayesian predictive distributions for Wishart distributions under the Kullback-Leibler divergence. The conditional reducibility of the family of Wishart distributions enables us to decompose the risk of a Bayesian predictive distribution. We considered a recently introduced class of prior distributions, which is called the family of enriched standard conjugate prior distributions, and compared the Bayesian predictive distributions based on these prior distributions. Furthermore, we studied the performance of the Bayesian predictive distribution based on the reference prior distribution in the family and showed that there exists a prior distribution in the family that dominates the reference prior distribution. Our study provides new insight into the multivariate analysis when there exists an ordered inferential importance for the independent variables.

math.ST

Structured regularization based velocity structure estimation in local earthquake tomography for the adaptation to velocity discontinuities

We propose a local earthquake tomography method that applies a structured regularization technique to determine sharp changes in Earth's seismic velocity structure using arrival time data of direct waves. Our approach focuses on the ability to better image two common features that are observed in Earth's seismic velocity structure: sharp changes in velocities that correspond to material boundaries, such as the Conrad and Moho discontinuities; and gradual changes in velocity that are associated with pressure and temperature distributions in the crust and mantle. We employ different penalty terms in the vertical and horizontal directions to refine the earthquake tomography. We utilize a vertical-direction (depth) penalty that takes the form of the l1-sum of the l2-norms of the second-order differences of the horizontal units in the vertical direction. This penalty is intended to represent sharp velocity changes caused by discontinuities by creating a piecewise linear depth profile of seismic velocity. We set a horizontal-direction penalty term on the basis of the l2-norm to express gradual velocity tendencies in the horizontal direction, which has been often used in conventional tomography methods. We use a synthetic data set to demonstrate that our method provides significant improvements over velocity structures estimated using conventional methods by obtaining stable estimates of both steep and gradual changes in velocity. Furthermore, we apply our proposed method to real seismic data in central Japan and present the potential of our method for detecting velocity discontinuities using the observed arrival times from a small number of local earthquakes.

stat.AP

Shrinkage priors for nonparametric Bayesian prediction of nonhomogeneous Poisson processes

We consider nonparametric Bayesian estimation and prediction for nonhomogeneous Poisson process models with unknown intensity functions. We propose a class of improper priors for intensity functions. Nonparametric Bayesian inference with kernel mixture based on the class improper priors is shown to be useful, although improper priors have not been widely used for nonparametric Bayes problems. Several theorems corresponding to those for finite-dimensional independent Poisson models hold for nonhomogeneous Poisson process models with infinite-dimensional parameter spaces. Bayesian estimation and prediction based on the improper priors are shown to be admissible under the Kullback--Leibler loss. Numerical methods for Bayesian inference based on the priors are investigated.

math.ST

Singular Value Shrinkage Priors for Bayesian Prediction

We develop singular value shrinkage priors for the mean matrix parameters in the matrix-variate normal model with known covariance matrices. Our priors are superharmonic and put more weight on matrices with smaller singular values. They are a natural generalization of the Stein prior. Bayes estimators and Bayesian predictive densities based on our priors are minimax and dominate those based on the uniform prior in finite samples. In particular, our priors work well when the true value of the parameter has low rank.

math.ST

Shrinkage priors on complex-valued circular-symmetric autoregressive processes

We investigate shrinkage priors on power spectral densities for complex-valued circular-symmetric autoregressive processes. We construct shrinkage predictive power spectral densities, which asymptotically dominate (i) the Bayesian predictive power spectral density based on the Jeffreys prior and (ii) the estimative power spectral density with the maximal likelihood estimator, where the Kullback-Leibler divergence from the true power spectral density to a predictive power spectral density is adopted as a risk. Furthermore, we propose general constructions of objective priors for Kähler parameter spaces, utilizing a positive continuous eigenfunction of the Laplace-Beltrami operator with a negative eigenvalue. We present numerical experiments on a complex-valued stationary autoregressive model of order $1$.

math.ST

Bayes Extended Estimators for Curved Exponential Families

The Bayesian predictive density has complex representation and does not belong to any finite-dimensional statistical model except for in limited situations. In this paper, we introduce its simple approximate representation employing its projection onto a finite-dimensional exponential family. Its theoretical properties are established parallelly to those of the Bayesian predictive density when the model belongs to curved exponential families. It is also demonstrated that the projection asymptotically coincides with the plugin density with the posterior mean of the expectation parameter of the exponential family, which we refer to as the Bayes extended estimator. Information-geometric correspondence indicates that the Bayesian predictive density can be represented as the posterior mean of the infinite-dimensional exponential family. The Kullback--Leibler risk performance of the approximation is demonstrated by numerical simulations and it indicates that the posterior mean of the expectation parameter approaches the Bayesian predictive density as the dimension of the exponential family increases. It also suggests that approximation by projection onto an exponential family of reasonable size is practically advantageous with respect to risk performance and computational cost.

math.ST

Minimax Predictive Density for Sparse Count Data

This paper discusses predictive densities under the Kullback--Leibler loss for high-dimensional Poisson sequence models under sparsity constraints. Sparsity in count data implies zero-inflation. We present a class of Bayes predictive densities that attain asymptotic minimaxity in sparse Poisson sequence models. We also show that our class with an estimator of unknown sparsity level plugged-in is adaptive in the asymptotically minimax sense. For application, we extend our results to settings with quasi-sparsity and with missing-completely-at-random observations. The simulation studies as well as application to real data illustrate the efficiency of the proposed Bayes predictive densities.

math.ST

Learning partially ranked data based on graph regularization

Ranked data appear in many different applications, including voting and consumer surveys. There often exhibits a situation in which data are partially ranked. Partially ranked data is thought of as missing data. This paper addresses parameter estimation for partially ranked data under a (possibly) non-ignorable missing mechanism. We propose estimators for both complete rankings and missing mechanisms together with a simple estimation procedure. Our estimation procedure leverages a graph regularization in conjunction with the Expectation-Maximization algorithm. Our estimation procedure is theoretically guaranteed to have the convergence properties. We reduce a modeling bias by allowing a non-ignorable missing mechanism. In addition, we avoid the inherent complexity within a non-ignorable missing mechanism by introducing a graph regularization. The experimental results demonstrate that the proposed estimators work well under non-ignorable missing mechanisms.

stat.ME

Non-asymptotic Bayesian Minimax Adaptation

This paper studies a Bayesian approach to non-asymptotic minimax adaptation in nonparametric estimation. Estimating an input function on the basis of output functions in a Gaussian white-noise model is discussed. The input function is assumed to be in a Sobolev ellipsoid with an unknown smoothness and an unknown radius. Our purpose in this paper is to present a Bayesian approach attaining minimaxity up to a universal constant without any knowledge regarding the smoothness and the radius. Our Bayesian approach provides not only a rate-exact minimax adaptive estimator in large sample asymptotics but also a risk bound for the Bayes estimator quantifying the effects of both the smoothness and the ratio of the squared radius to the noise variance, where the smoothness and the ratio are the key parameters to describe the minimax risk in this model. Application to non-parametric regression models is also discussed.

math.ST