Searcharxiv⌕ Search

arXiv subjects

Mohsen Pourahmadi

Publications and source records attributed to Mohsen Pourahmadi.

12 recordsLinked to original sources

An Alternative Graphical Lasso Algorithm for Precision Matrices

The Graphical Lasso (GLasso) algorithm is fast and widely used for estimating sparse precision matrices (Friedman et al., 2008). Its central role in the literature of high-dimensional covariance estimation rivals that of Lasso regression for sparse estimation of the mean vector. Some mysteries regarding its optimization target, convergence, positive-definiteness and performance have been unearthed, resolved and presented in Mazumder and Hastie (2011), leading to a new/improved (dual-primal) DP-GLasso. Using a new and slightly different reparametriztion of the last column of a precision matrix we show that the regularized normal log-likelihood naturally decouples into a sum of two easy to minimize convex functions one of which is a Lasso regression problem. This decomposition is the key in developing a transparent, simple iterative block coordinate descent algorithm for computing the GLasso updates with performance comparable to DP-GLasso. In particular, our algorithm has the precision matrix as its optimization target right at the outset, and retains all the favorable properties of the DP-GLasso algorithm.

stat.CO↗

Adaptive Bayesian Variable Clustering via Structural Learning of Breast Cancer Data

Clustering of proteins is of interest in cancer cell biology. This article proposes a hierarchical Bayesian model for protein (variable) clustering hinging on correlation structure. Starting from a multivariate normal likelihood, we enforce the clustering through prior modeling using angle based unconstrained reparameterization of correlations and assume a truncated Poisson distribution (to penalize the large number of clusters) as prior on the number of clusters. The posterior distributions of the parameters are not in explicit form and we use a reversible jump Markov chain Monte Carlo (RJMCMC) based technique is used to simulate the parameters from the posteriors. The end products of the proposed method are estimated cluster configuration of the proteins (variables) along with the number of clusters. The Bayesian method is flexible enough to cluster the proteins as well as the estimate the number of clusters. The performance of the proposed method has been substantiated with extensive simulation studies and one protein expression data with a hereditary disposition in breast cancer where the proteins are coming from different pathways.

stat.CO↗

Estimation of Constrained Mean-Covariance of Normal Distributions

Estimation of the mean vector and covariance matrix is of central importance in the analysis of multivariate data. In the framework of generalized linear models, usually the variances are certain functions of the means with the normal distribution being an exception. We study some implications of functional relationships between covariance and the mean by focusing on the maximum likelihood and Bayesian estimation of the mean-covariance under the joint constraint $\bmΣ\bmμ = \bmμ$ for a multivariate normal distribution. A novel structured covariance is proposed through reparameterization of the spectral decomposition of $\bmΣ$ involving its eigenvalues and $\bmμ$. This is designed to address the challenging issue of positive-definiteness and to reduce the number of covariance parameters from quadratic to linear function of the dimension. We propose a fast (noniterative) method for approximating the maximum likelihood estimator by maximizing a lower bound for the profile likelihood function, which is concave. We use normal and inverse gamma priors on the mean and eigenvalues, and approximate the maximum aposteriori estimators by both MH within Gibbs sampling and a faster iterative method. A simulation study shows good performance of our estimators.

stat.ME↗

Learning Bayesian Networks through Birkhoff Polytope: A Relaxation Method

We establish a novel framework for learning a directed acyclic graph (DAG) when data are generated from a Gaussian, linear structural equation model. It consists of two parts: (1) introduce a permutation matrix as a new parameter within a regularized Gaussian log-likelihood to represent variable ordering; and (2) given the ordering, estimate the DAG structure through sparse Cholesky factor of the inverse covariance matrix. For permutation matrix estimation, we propose a relaxation technique that avoids the NP-hard combinatorial problem of order estimation. Given an ordering, a sparse Cholesky factor is estimated using a cyclic coordinatewise descent algorithm which decouples row-wise. Our framework recovers DAGs without the need for an expensive verification of the acyclicity constraint or enumeration of possible parent sets. We establish numerical convergence of the algorithm, and consistency of the Cholesky factor estimator when the order of variables is known. Through several simulated and macro-economic datasets, we study the scope and performance of the proposed methodology.

stat.ML↗

Time Series Graphical Lasso and Sparse VAR Estimation

We improve upon the two-stage sparse vector autoregression (sVAR) method in Davis et al. (2016) by proposing an alternative two-stage modified sVAR method which relies on time series graphical lasso to estimate sparse inverse spectral density in the first stage, and the second stage refines non-zero entries of the AR coefficient matrices using a false discovery rate (FDR) procedure. Our method has the advantage of avoiding the inversion of the spectral density matrix but has to deal with optimization over Hermitian matrices with complex-valued entries. It significantly improves the computational time with a little loss in forecasting performance. We study the properties of our proposed method and compare the performance of the two methods using simulated and a real macro-economic dataset. Our simulation results show that the proposed modification or msVAR is a preferred choice when the goal is to learn the structure of the AR coefficient matrices while sVAR outperforms msVAR when the ultimate task is forecasting.

stat.CO↗

MLE of Jointly Constrained Mean-Covariance of Multivariate Normal Distributions

Estimating the unconstrained mean and covariance matrix is a popular topic in statistics. However, estimation of the parameters of $N_p(μ,Σ)$ under joint constraints such as $Σμ= μ$ has not received much attention. It can be viewed as a multivariate counterpart of the classical estimation problem in the $N(θ,θ^2)$ distribution. In addition to the usual inference challenges under such non-linear constraints among the parameters (curved exponential family), one has to deal with the basic requirements of symmetry and positive definiteness when estimating a covariance matrix. We derive the non-linear likelihood equations for the constrained maximum likelihood estimator of $(μ,Σ)$ and solve them using iterative methods. Generally, the MLE of covariance matrices computed using iterative methods do not satisfy the constraints. We propose a novel algorithm to modify such (infeasible) estimators or any other (reasonable) estimator. The key step is to re-align the mean vector along the eigenvectors of the covariance matrix using the idea of regression. In using the Lagrangian function for constrained MLE (Aitchison et al. 1958), the Lagrange multiplier entangles with the parameters of interest and presents another computational challenge. We handle this by either iterative or explicit calculation of the Lagrange multiplier. The existence and nature of location of the constrained MLE are explored within a data-dependent convex set using recent results from random matrix theory. A simulation study illustrates our methodology and shows that the modified estimators perform better than the initial estimators from the iterative methods.

stat.ME↗

Fused-Lasso Regularized Cholesky Factors of Large Nonstationary Covariance Matrices of Longitudinal Data

Smoothness of the subdiagonals of the Cholesky factor of large covariance matrices is closely related to the degrees of nonstationarity of autoregressive models for time series and longitudinal data. Heuristically, one expects for a nearly stationary covariance matrix the entries in each subdiagonal of the Cholesky factor of its inverse to be nearly the same in the sense that sum of absolute values of successive terms is small. Statistically such smoothness is achieved by regularizing each subdiagonal using fused-type lasso penalties. We rely on the standard Cholesky factor as the new parameters within a regularized normal likelihood setup which guarantees: (1) joint convexity of the likelihood function, (2) strict convexity of the likelihood function restricted to each subdiagonal even when $n<p$, and (3) positive-definiteness of the estimated covariance matrix. A block coordinate descent algorithm, where each block is a subdiagonal, is proposed and its convergence is established under mild conditions. Lack of decoupling of the penalized likelihood function into a sum of functions involving individual subdiagonals gives rise to some computational challenges and advantages relative to two recent algorithms for sparse estimation of the Cholesky factor which decouple row-wise. Simulation results and real data analysis show the scope and good performance of the proposed methodology.

stat.ML↗

Stationary subspace analysis of nonstationary covariance processes: eigenstructure description and testing

Stationary subspace analysis (SSA) searches for linear combinations of the components of nonstationary vector time series that are stationary. These linear combinations and their number defne an associated stationary subspace and its dimension. SSA is studied here for zero mean nonstationary covariance processes. We characterize stationary subspaces and their dimensions in terms of eigenvalues and eigenvectors of certain symmetric matrices. This characterization is then used to derive formal statistical tests for estimating dimensions of stationary subspaces. Eigenstructure-based techniques are also proposed to estimate stationary subspaces, without relying on previously used computationally intensive optimization-based methods. Finally, the introduced methodologies are examined on simulated and real data.

stat.ME↗

Baxter's inequality for finite predictor coefficients of multivariate long-memory stationary processes

For a multivariate stationary process, we develop explicit representations for the finite predictor coefficient matrices, the finite prediction error covariance matrices and the partial autocorrelation function (PACF) in terms of the Fourier coefficients of its phase function in the spectral domain. The derivation is based on a novel alternating projection technique and the use of the forward and backward innovations corresponding to the predictions based on the infinite past and future, respectively. We show that such representations are ideal for studying the rates of convergence of the finite predictor coefficients, prediction error covariances, and the PACF as well as for proving a multivariate version of Baxter's inequality for a multivariate FARIMA process with a common fractional differencing order for all components of the process.

math.PR↗

The intersection of past and future for multivariate stationary processes

We consider an intersection of past and future property of multivariate stationary processes which is the key to deriving various representation theorems for their linear predictor coefficient matrices. We extend useful spectral characterizations for this property from univariate processes to multivariate processes.

math.PR↗

Covariance Estimation: The GLM and Regularization Perspectives

Finding an unconstrained and statistically interpretable reparameterization of a covariance matrix is still an open problem in statistics. Its solution is of central importance in covariance estimation, particularly in the recent high-dimensional data environment where enforcing the positive-definiteness constraint could be computationally expensive. We provide a survey of the progress made in modeling covariance matrices from two relatively complementary perspectives: (1) generalized linear models (GLM) or parsimony and use of covariates in low dimensions, and (2) regularization or sparsity for high-dimensional data. An emerging, unifying and powerful trend in both perspectives is that of reducing a covariance estimation problem to that of estimating a sequence of regression problems. We point out several instances of the regression-based formulation. A notable case is in sparse estimation of a precision matrix or a Gaussian graphical model leading to the fast graphical LASSO algorithm. Some advantages and limitations of the regression-based Cholesky decomposition relative to the classical spectral (eigenvalue) and variance-correlation decompositions are highlighted. The former provides an unconstrained and statistically interpretable reparameterization, and guarantees the positive-definiteness of the estimated covariance matrix. It reduces the unintuitive task of covariance estimation to that of modeling a sequence of regressions at the cost of imposing an a priori order among the variables. Elementwise regularization of the sample covariance matrix such as banding, tapering and thresholding has desirable asymptotic properties and the sparse estimated covariance matrix is positive definite with probability tending to one for large samples and dimensions.

stat.ME↗

Applications of a finite-dimensional duality principle to some prediction problems

Some of the most important results in prediction theory and time series analysis when finitely many values are removed from or added to its infinite past have been obtained using difficult and diverse techniques ranging from duality in Hilbert spaces of analytic functions (Nakazi, 1984) to linear regression in statistics (Box and Tiao, 1975). We unify these results via a finite-dimensional duality lemma and elementary ideas from the linear algebra. The approach reveals the inherent finite-dimensional character of many difficult prediction problems, the role of duality and biorthogonality for a finite set of random variables. The lemma is particularly useful when the number of missing values is small, like one or two, as in the case of Kolmogorov and Nakazi prediction problems. The stationarity of the underlying process is not a requirement. It opens up the possibility of extending such results to nonstationary processes.

math.PR↗