SearcharxivSearch

arXiv subjects

Yoshikazu Terada

Publications and source records attributed to Yoshikazu Terada.

At least 19 recordsLinked to original sources

Functional Multiple-Set Canonical Correlation Analysis Revisited: From Finite-Dimensional Samples to Infinite-Dimensional Populations

We develop a population-level formulation of functional multiple-set canonical correlation analysis (P-FMCCA) for multivariate functional data in an infinite-dimensional Hilbert space. Since covariance operators for functional data are typically compact, the inverse covariance operators that appear in the formal extension of multiple-set canonical correlation analysis (MCCA) are generally unbounded and are not defined on the whole Hilbert space. We therefore provide sufficient conditions under which the proposed population formulation is well-defined and show that the resulting constrained maximization problem is characterized by an eigenvalue problem for a Hilbert-Schmidt extension of the relevant correlation operator. We further establish a canonical decomposition induced by the P-FMCCA components and introduce the associated truncated canonical representation. In addition, we formulate functional homogeneity analysis at the population level and show that the finite-dimensional equivalence between homogeneity analysis and MCCA does not generally carry over to the infinite-dimensional setting. Finally, we prove that, for the finite-rank truncated canonical representation, functional homogeneity analysis admits an explicit characterization in terms of the P-FMCCA components, thereby providing a population-level counterpart of the classical finite-dimensional correspondence.

stat.ME

Exponential mixing properties of nonlinear functional autoregressive models

The importance of functional data analysis has increased substantially in recent years. In machine learning, nonlinear function regression based on deep neural networks is referred to as operator learning, and many of its applications involve functional time series data. However, the theoretical understanding of nonlinear models in functional time series analysis remains limited, as most existing works focus on linear models. In this paper, we derive basic properties for analyzing adaptive learning in nonlinear functional autoregressive (NFAR) models. Specifically, we derive sufficient conditions for NFAR models to be exponentially mixing. We provide an example with a Hammerstein operator under which these conditions are satisfied. As an application of exponential mixing, we consider operator learning for NFAR models with Urysohn operators and derive convergence rates for adaptive estimators based on deep neural networks.

math.ST

A Theory of Nonparametric Covariance Function Estimation for Discretely Observed Data

We study nonparametric covariance function estimation for functional data observed with noise at discrete locations on a $d$-dimensional domain. Estimating the covariance function from discretely observed data is a challenging nonparametric problem, particularly in multidimensional settings, since the covariance function is defined on a product domain and thus suffers from the curse of dimensionality. This motivates the use of adaptive estimators, such as deep learning estimators. However, existing theoretical results are largely limited to estimators with explicit analytic representations, and the properties of general learning-based estimators remain poorly understood. We establish an oracle inequality for a broad class of learning-based estimators that applies to both sparse and dense observation regimes in a unified manner, and derive convergence rates for deep learning estimators over several classes of covariance functions. The resulting rates suggest that structural adaptation can mitigate the curse of dimensionality, similarly to classical nonparametric regression. We further compare the convergence rates of learning-based estimators with several existing procedures. For a one-dimensional smoothness class, deep learning estimators are suboptimal, whereas local linear smoothing estimators achieve a faster rate. For a structured function class, however, deep learning estimators attain the minimax rate up to polylogarithmic factors, whereas local linear smoothing estimators are suboptimal. These results reveal a distinctive adaptivity-variance trade-off in covariance function estimation.

math.ST

Glitch noise classification in KAGRA O3GK observing data using unsupervised machine learning

Gravitational wave interferometers are disrupted by various types of nonstationary noise, referred to as glitch noise, that affect data analysis and interferometer sensitivity. The accurate identification and classification of glitch noise are essential for improving the reliability of gravitational wave observations. In this study, we demonstrated the effectiveness of unsupervised machine learning for classifying images with nonstationary noise in the KAGRA O3GK data. Using a variational autoencoder (VAE) combined with spectral clustering, we identified eight distinct glitch noise categories. The latent variables obtained from VAE were dimensionally compressed, visualized in three-dimensional space, and classified using spectral clustering to better understand the glitch noise characteristics of KAGRA during the O3GK period. Our results highlight the potential of unsupervised learning for efficient glitch noise classification, which may in turn potentially facilitate interferometer upgrades and the development of future third-generation gravitational wave observatories.

gr-qc

Regularized k-POD: Sparse k-means clustering for high-dimensional missing data

The classical k-means clustering, based on distances computed from all data features, cannot be directly applied to incomplete data with missing values. A natural extension of k-means to missing data, namely k-POD, uses only the observed entries for clustering and is both computationally efficient and flexible. However, for high-dimensional missing data including features irrelevant to the underlying cluster structure, the presence of such irrelevant features leads to the bias of k-POD in estimating cluster centers, thereby damaging its clustering effect. Nevertheless, the existing k-POD method performs well in low-dimensional cases, highlighting the importance of addressing the bias issue. To this end, in this paper, we propose a regularized k-POD clustering method that applies feature-wise regularization on cluster centers into the existing k-POD clustering. Such a penalty on cluster centers enables us to effectively reduce the bias of k-POD for high-dimensional missing data. To the best of our knowledge, our method is the first to mitigate bias in k-means-type clustering for high-dimensional missing data, while retaining the computational efficiency and flexibility. Simulation results verify that the proposed method effectively reduces bias and improves clustering performance. Applications to real-world single-cell RNA sequencing data further show the utility of the proposed method.

stat.ME

Statistical properties of matrix decomposition factor analysis

Numerous estimators have been proposed for factor analysis, and their statistical properties have been extensively studied. In the early 2000s, a novel matrix factorization-based approach, known as Matrix Decomposition Factor Analysis (MDFA), was introduced and has been actively developed in computational statistics. The MDFA estimator offers several advantages, including the guarantee of proper solutions (i.e., no Heywood cases) and the extensibility to $\ell_0$-sparse estimation. However, the MDFA estimator does not appear to be formulated as a classical M-estimator or a minimum discrepancy function (MDF) estimator, and the statistical properties of the MDFA estimator have remained largely unexplored. Although the MDFA estimator minimizes a loss function resembling that of principal component analysis (PCA), it empirically behaves more like consistent estimators used in factor analysis than like PCA itself. This raises a fundamental question: can matrix decomposition factor analysis truly be regarded as "factor analysis"? To address this issue, we establish consistency and asymptotic normality of the MDFA estimator. Recognizing that the MDFA estimator can be formulated as a semiparametric maximum likelihood estimator, we surprisingly find that the profile likelihood is given by the squared Bures-Wasserstein distance between the sample covariance matrix and the modeled covariance matrix. As a consequence, the MDFA estimator is ultimately an MDF estimator for factor analysis. Beyond MDFA, the same representation holds for a broad class of component analysis methods, including PCA, thereby offering a unified perspective on component analysis. Numerical experiments demonstrate that MDFA performs competitively with other established estimators, suggesting that it is a theoretically grounded and computationally appealing alternative for factor analysis.

math.ST

Nonparametric function-on-scalar regression using deep neural networks

We focus on nonlinear Function-on-Scalar regression, where the predictors are scalar variables, and the responses are functional data. Most existing studies approximate the hidden nonlinear relationships using linear combinations of basis functions, such as splines. However, in classical nonparametric regression, it is known that these approaches lack adaptivity, particularly when the true function exhibits high spatial inhomogeneity or anisotropic smoothness. To capture the complex structure behind data adaptively, we propose a simple adaptive estimator based on a deep neural network model. The proposed estimator is straightforward to implement using existing deep learning libraries, making it accessible for practical applications. Moreover, we derive the convergence rates of the proposed estimator for the anisotropic Besov spaces, which consist of functions with varying smoothness across dimensions. Our theoretical analysis shows that the proposed estimator mitigates the curse of dimensionality when the true function has high anisotropic smoothness, as shown in the classical nonparametric regression. Numerical experiments demonstrate the superior adaptivity of the proposed estimator, outperforming existing methods across various challenging settings. Moreover, the proposed method is applied to analyze ground reaction force data in the field of sports medicine, demonstrating more efficient estimation compared to existing approaches.

stat.ME

Tree-Guided $L_1$-Convex Clustering

Convex clustering is a modern clustering framework that guarantees globally optimal solutions and performs comparably to other advanced clustering methods. However, obtaining a complete dendrogram (clusterpath) for large-scale datasets remains computationally challenging due to the extensive costs associated with iterative optimization approaches. To address this limitation, we develop a novel convex clustering algorithm called Tree-Guided $L_1$-Convex Clustering (TGCC). We first focus on the fact that the loss function of $L_1$-convex clustering with tree-structured weights can be efficiently optimized using a dynamic programming approach. We then develop an efficient cluster fusion algorithm that utilizes the tree structure of the weights to accelerate the optimization process and eliminate the issue of cluster splits commonly observed in convex clustering. By combining the dynamic programming approach with the cluster fusion algorithm, the TGCC algorithm achieves superior computational efficiency without sacrificing clustering performance. Remarkably, our TGCC algorithm can construct a complete clusterpath for $10^6$ points in $\mathbb{R}^2$ within 15 seconds on a standard laptop without the need for parallel or distributed computing frameworks. Moreover, we extend the TGCC algorithm to develop biclustering and sparse convex clustering algorithms.

cs.LG

Nonparametric logistic regression with deep learning

Consider the nonparametric logistic regression problem. In the logistic regression, we usually consider the maximum likelihood estimator, and the excess risk is the expectation of the Kullback-Leibler (KL) divergence between the true and estimated conditional class probabilities. However, in the nonparametric logistic regression, the KL divergence could diverge easily, and thus, the convergence of the excess risk is difficult to prove or does not hold. Several existing studies show the convergence of the KL divergence under strong assumptions. In most cases, our goal is to estimate the true conditional class probabilities. Thus, instead of analyzing the excess risk itself, it suffices to show the consistency of the maximum likelihood estimator in some suitable metric. In this paper, using a simple unified approach for analyzing the nonparametric maximum likelihood estimator (NPMLE), we directly derive convergence rates of the NPMLE in the Hellinger distance under mild assumptions. Although our results are similar to the results in some existing studies, we provide simple and more direct proofs for these results. As an important application, we derive convergence rates of the NPMLE with fully connected deep neural networks and show that the derived rate nearly achieves the minimax optimal rate.

math.ST

Sparse factor models of high dimension

We consider the estimation of a sparse factor model where the factor loading matrix is assumed sparse. The estimation problem is reformulated as a penalized M-estimation criterion, while the restrictions for identifying the factor loading matrix accommodate a wide range of sparsity patterns. We prove the sparsistency property of the penalized estimator when the number of parameters is diverging, that is the consistency of the estimator and the recovery of the true zeros entries. These theoretical results are illustrated by finite-sample simulation experiments, and the relevance of the proposed method is assessed by applications to portfolio allocation and macroeconomic data prediction.

math.ST

$K$-means clustering for sparsely observed longitudinal data

In longitudinal data analysis, observation points of repeated measurements over time often vary among subjects except in well-designed experimental studies. Additionally, measurements for each subject are typically obtained at only a few time points. From such sparsely observed data, identifying underlying cluster structures can be challenging. This paper proposes a fast and simple clustering method that generalizes the classical $k$-means method to identify cluster centers in sparsely observed data. The proposed method employs the basis function expansion to model the cluster centers, providing an effective way to estimate cluster centers from fragmented data. We establish the statistical consistency of the proposed method, as with the classical $k$-means method. Through numerical experiments, we demonstrate that the proposed method performs competitively with, or even outperforms, existing clustering methods. Moreover, the proposed method offers significant gains in computational efficiency due to its simplicity. Applying the proposed method to real-world data illustrates its effectiveness in identifying cluster structures in sparsely observed data.

stat.ME

Some notes on the $k$-means clustering for missing data

The classical $k$-means clustering requires a complete data matrix without missing entries. As a natural extension of the $k$-means clustering for missing data, the $k$-POD clustering has been proposed, which ignores the missing entries in the $k$-means clustering. This paper shows the inconsistency of the $k$-POD clustering even under the missing completely at random mechanism. More specifically, the expected loss of the $k$-POD clustering can be represented as the weighted sum of the expected $k$-means losses with parts of variables. Thus, the $k$-POD clustering converges to the different clustering from the $k$-means clustering as the sample size goes to infinity. This result indicates that although the $k$-means clustering works well, the $k$-POD clustering may fail to capture the hidden cluster structure. On the other hand, for high-dimensional data, the $k$-POD clustering could be a suitable choice when the missing rate in each variable is low.

math.ST

Semiparametric adaptive estimation under informative sampling

In survey sampling, survey data do not necessarily represent the target population, and the samples are often biased. However, information on the survey weights aids in the elimination of selection bias. The Horvitz-Thompson estimator is a well-known unbiased, consistent, and asymptotically normal estimator; however, it is not efficient. Thus, this study derives the semiparametric efficiency bound for various target parameters by considering the survey weight as a random variable and consequently proposes a semiparametric optimal estimator with certain working models on the survey weights. The proposed estimator is consistent, asymptotically normal, and efficient in a class of the regular and asymptotically linear estimators. Further, a limited simulation study is conducted to investigate the finite sample performance of the proposed method. The proposed method is applied to the 1999 Canadian Workplace and Employee Survey data.

stat.ME

Convex Clustering through MM: An Efficient Algorithm to Perform Hierarchical Clustering

Convex clustering is a modern method with both hierarchical and $k$-means clustering characteristics. Although convex clustering can capture complex clustering structures hidden in data, the existing convex clustering algorithms are not scalable to large data sets with sample sizes greater than several thousands. Moreover, it is known that convex clustering sometimes fails to produce a complete hierarchical clustering structure. This issue arises if clusters split up or the minimum number of possible clusters is larger than the desired number of clusters. In this paper, we propose convex clustering through majorization-minimization (CCMM) -- an iterative algorithm that uses cluster fusions and a highly efficient updating scheme derived using diagonal majorization. Additionally, we explore different strategies to ensure that the hierarchical clustering structure terminates in a single cluster. With a current desktop computer, CCMM efficiently solves convex clustering problems featuring over one million objects in seven-dimensional space, achieving a solution time of 51 seconds on average.

stat.ML

Selective inference after feature selection via multiscale bootstrap

It is common to show the confidence intervals or $p$-values of selected features, or predictor variables in regression, but they often involve selection bias. The selective inference approach solves this bias by conditioning on the selection event. Most existing studies of selective inference consider a specific algorithm, such as Lasso, for feature selection, and thus they have difficulties in handling more complicated algorithms. Moreover, existing studies often consider unnecessarily restrictive events, leading to over-conditioning and lower statistical power. Our novel and widely-applicable resampling method via multiscale bootstrap addresses these issues to compute an approximately unbiased selective $p$-value for the selected features. As a simplification of the proposed method, we also develop a simpler method via the classical bootstrap. We prove that the $p$-value computed by our multiscale bootstrap method is more accurate than the classical bootstrap method. Furthermore, numerical experiments demonstrate that our algorithm works well even for more complicated feature selection methods such as non-convex regularization.

stat.ME

Forecasting temporal variation of aftershocks immediately after a main shock using Gaussian process regression

Uncovering the distribution of magnitudes and arrival times of aftershocks is a key to comprehend the characteristics of the sequence of earthquakes, which enables us to predict seismic activities and hazard assessments. However, identifying the number of aftershocks immediately after the main shock is practically difficult due to contaminations of arriving seismic waves. To overcome the difficulty, we construct a likelihood based on the detected data incorporating a detection function to which the Gaussian process regression (GPR) is applied. The GPR is capable of estimating not only the parameters of the distribution of aftershocks together with the detection function but also credible intervals for both of the parameters and the detection function. A property that distributions of both the Gaussian process and aftershocks are exponential functions leads to an efficient Bayesian computational algorithm to estimate the hyperparameters. After the validations through numerical tests, the proposed method is retrospectively applied to the catalog data related to the 2004 Chuetsu earthquake towards early forecasting of the aftershocks. The result shows that the proposed method stably estimates the parameters of the distribution simultaneously their credible intervals even within three hours after the main shock.

physics.geo-ph

More Powerful Selective Kernel Tests for Feature Selection

Refining one's hypotheses in the light of data is a common scientific practice; however, the dependency on the data introduces selection bias and can lead to specious statistical analysis. An approach for addressing this is via conditioning on the selection procedure to account for how we have used the data to generate our hypotheses, and prevent information to be used again after selection. Many selective inference (a.k.a. post-selection inference) algorithms typically take this approach but will "over-condition" for sake of tractability. While this practice yields well calibrated statistic tests with controlled false positive rates (FPR), it can incur a major loss in power. In our work, we extend two recent proposals for selecting features using the Maximum Mean Discrepancy and Hilbert Schmidt Independence Criterion to condition on the minimal conditioning event. We show how recent advances in multiscale bootstrap makes conditioning on the minimal selection event possible and demonstrate our proposal over a range of synthetic and real world experiments. Our results show that our proposed test is indeed more powerful in most scenarios.

cs.LG

Fast generalization error bound of deep learning without scale invariance of activation functions

In theoretical analysis of deep learning, discovering which features of deep learning lead to good performance is an important task. In this paper, using the framework for analyzing the generalization error developed in Suzuki (2018), we derive a fast learning rate for deep neural networks with more general activation functions. In Suzuki (2018), assuming the scale invariance of activation functions, the tight generalization error bound of deep learning was derived. They mention that the scale invariance of the activation function is essential to derive tight error bounds. Whereas the rectified linear unit (ReLU; Nair and Hinton, 2010) satisfies the scale invariance, the other famous activation functions including the sigmoid and the hyperbolic tangent functions, and the exponential linear unit (ELU; Clevert et al., 2016) does not satisfy this condition. The existing analysis indicates a possibility that a deep learning with the non scale invariant activations may have a slower convergence rate of $O(1/\sqrt{n})$ when one with the scale invariant activations can reach a rate faster than $O(1/\sqrt{n})$. In this paper, without the scale invariance of activation functions, we derive the tight generalization error bound which is essentially the same as that of Suzuki (2018). From this result, at least in the framework of Suzuki (2018), it is shown that the scale invariance of the activation functions is not essential to get the fast rate of convergence. Simultaneously, it is also shown that the theoretical framework proposed by Suzuki (2018) can be widely applied for analysis of deep learning with general activation functions.

stat.ML