SearcharxivSearch

arXiv subjects

Michio Yamamoto

Publications and source records attributed to Michio Yamamoto.

8 recordsLinked to original sources

Functional Multiple-Set Canonical Correlation Analysis Revisited: From Finite-Dimensional Samples to Infinite-Dimensional Populations

We develop a population-level formulation of functional multiple-set canonical correlation analysis (P-FMCCA) for multivariate functional data in an infinite-dimensional Hilbert space. Since covariance operators for functional data are typically compact, the inverse covariance operators that appear in the formal extension of multiple-set canonical correlation analysis (MCCA) are generally unbounded and are not defined on the whole Hilbert space. We therefore provide sufficient conditions under which the proposed population formulation is well-defined and show that the resulting constrained maximization problem is characterized by an eigenvalue problem for a Hilbert-Schmidt extension of the relevant correlation operator. We further establish a canonical decomposition induced by the P-FMCCA components and introduce the associated truncated canonical representation. In addition, we formulate functional homogeneity analysis at the population level and show that the finite-dimensional equivalence between homogeneity analysis and MCCA does not generally carry over to the infinite-dimensional setting. Finally, we prove that, for the finite-rank truncated canonical representation, functional homogeneity analysis admits an explicit characterization in terms of the P-FMCCA components, thereby providing a population-level counterpart of the classical finite-dimensional correspondence.

stat.ME

Algebraic Approach for Orthomax Rotations

In exploratory factor analysis, rotation techniques are employed to derive interpretable factor loading matrices. Factor rotations deal with equality-constrained optimization problems aimed at determining a loading matrix based on measure of simplicity, such as ``perfect simple structure'' and ``Thurstone simple structure.'' Numerous criteria have been proposed, since the concept of simple structure is fundamentally ambiguous and involves multiple distinct aspects. However, most rotation criteria may fail to consistently yield a simple structure that is optimal for analytical purposes, primarily due to two challenges. First, existing optimization techniques, including the gradient projection descent method, exhibit strong dependence on initial values and frequently become trapped in suboptimal local optima. Second, multifaceted nature of simple structure complicates the ability of any single criterion to ensure interpretability across all aspects. In certain cases, even when a global optimum is achieved, other rotations may exhibit simpler structures in specific aspects. To address these issues, obtaining all equality-constrained stationary points -- including both global and local optima -- is advantageous. Fortunately, many rotation criteria are expressed as algebraic functions, and the constraints in the optimization problems in factor rotations are formulated as algebraic equations. Therefore, we can employ computational algebra techniques that utilize operations within polynomial rings to derive exact all equality-constrained stationary points. Unlike existing optimization methods, the computational algebraic approach can determine global optima and all stationary points, independent of initial values. We conduct Monte Carlo simulations to examine the properties of the orthomax rotation criteria, which generalizes various orthogonal rotation methods.

math.ST

$K$-means clustering for sparsely observed longitudinal data

In longitudinal data analysis, observation points of repeated measurements over time often vary among subjects except in well-designed experimental studies. Additionally, measurements for each subject are typically obtained at only a few time points. From such sparsely observed data, identifying underlying cluster structures can be challenging. This paper proposes a fast and simple clustering method that generalizes the classical $k$-means method to identify cluster centers in sparsely observed data. The proposed method employs the basis function expansion to model the cluster centers, providing an effective way to estimate cluster centers from fragmented data. We establish the statistical consistency of the proposed method, as with the classical $k$-means method. Through numerical experiments, we demonstrate that the proposed method performs competitively with, or even outperforms, existing clustering methods. Moreover, the proposed method offers significant gains in computational efficiency due to its simplicity. Applying the proposed method to real-world data illustrates its effectiveness in identifying cluster structures in sparsely observed data.

stat.ME

Causal Discovery with Multi-Domain LiNGAM for Latent Factors

Discovering causal structures among latent factors from observed data is a particularly challenging problem. Despite some efforts for this problem, existing methods focus on the single-domain data only. In this paper, we propose Multi-Domain Linear Non-Gaussian Acyclic Models for Latent Factors (MD-LiNA), where the causal structure among latent factors of interest is shared for all domains, and we provide its identification results. The model enriches the causal representation for multi-domain data. We propose an integrated two-phase algorithm to estimate the model. In particular, we first locate the latent factors and estimate the factor loading matrix. Then to uncover the causal structure among shared latent factors of interest, we derive a score function based on the characterization of independence relations between external influences and the dependence relations between multi-domain latent factors and latent factors of interest. We show that the proposed method provides locally consistent estimators. Experimental results on both synthetic and real-world data demonstrate the efficacy and robustness of our approach.

cs.LG

Model-based clustering of multivariate binary data with dimension reduction

Clustering methods with dimension reduction have been receiving considerable wide interest in statistics lately and a lot of methods to simultaneously perform clustering and dimension reduction have been proposed. This work presents a novel procedure for simultaneously determining the optimal cluster structure for multivariate binary data and the subspace to represent that cluster structure. The method is based on a finite mixture model of multivariate Bernoulli distributions, and each component is assumed to have a low-dimensional representation of the cluster structure. This method can be considered an extension of the traditional latent class analysis model. Sparsity is introduced to the loading values, which produces the low-dimensional subspace, for enhanced interpretability and more stable extraction of the subspace. An EM-based algorithm is developed to efficiently solve the proposed optimization problem. We demonstrate the effectiveness of the proposed method by applying it to a simulation study and real datasets.

stat.ME

Functional Factorial K-means Analysis

A new procedure for simultaneously finding the optimal cluster structure of multivariate functional objects and finding the subspace to represent the cluster structure is presented. The method is based on the $k$-means criterion for projected functional objects on a subspace in which a cluster structure exists. An efficient alternating least-squares algorithm is described, and the proposed method is extended to a regularized method for smoothness of weight functions. To deal with the negative effect of the correlation of coefficient matrix of the basis function expansion in the proposed algorithm, a two-step approach to the proposed method is also described. Analyses of artificial and real data demonstrate that the proposed method gives correct and interpretable results compared with existing methods, the functional principal component $k$-means (FPCK) method and tandem clustering approach. It is also shown that the proposed method can be considered complementary to FPCK.

stat.ME

Sparse estimation via nonconcave penalized likelihood in a factor analysis model

We consider the problem of sparse estimation in a factor analysis model. A traditional estimation procedure in use is the following two-step approach: the model is estimated by maximum likelihood method and then a rotation technique is utilized to find sparse factor loadings. However, the maximum likelihood estimates cannot be obtained when the number of variables is much larger than the number of observations. Furthermore, even if the maximum likelihood estimates are available, the rotation technique does not often produce a sufficiently sparse solution. In order to handle these problems, this paper introduces a penalized likelihood procedure that imposes a nonconvex penalty on the factor loadings. We show that the penalized likelihood procedure can be viewed as a generalization of the traditional two-step approach, and the proposed methodology can produce sparser solutions than the rotation technique. A new algorithm via the EM algorithm along with coordinate descent is introduced to compute the entire solution path, which permits the application to a wide variety of convex and nonconvex penalties. Monte Carlo simulations are conducted to investigate the performance of our modeling strategy. A real data example is also given to illustrate our procedure.

stat.ME

Estimation of oblique structure via penalized likelihood factor analysis

We consider the problem of sparse estimation via a lasso-type penalized likelihood procedure in a factor analysis model. Typically, the model estimation is done under the assumption that the common factors are orthogonal (uncorrelated). However, the lasso-type penalization method based on the orthogonal model can often estimate a completely different model from that with the true factor structure when the common factors are correlated. In order to overcome this problem, we propose to incorporate a factor correlation into the model, and estimate the factor correlation along with parameters included in the orthogonal model by maximum penalized likelihood procedure. An entire solution path is computed by the EM algorithm with coordinate descent, which permits the application to a wide variety of convex and nonconvex penalties. The proposed method can provide sufficiently sparse solutions, and be applied to the data where the number of variables is larger than the number of observations. Monte Carlo simulations are conducted to investigate the effectiveness of our modeling strategies. The results show that the lasso-type penalization based on the orthogonal model cannot often approximate the true factor structure, whereas our approach performs well in various situations. The usefulness of the proposed procedure is also illustrated through the analysis of real data.

stat.ME