Searcharxiv⌕ Search

arXiv subjects

Bala Rajaratnam

Publications and source records attributed to Bala Rajaratnam.

At least 19 recordsLinked to original sources

High dimensional Bayesian inference for Gaussian directed acyclic graph models

We study centered Gaussian models Markov with respect to a directed acyclic graph (DAG) whose vertices have a fixed parent ordering. We construct a conjugate family on the modified Cholesky parameters, with one shape parameter per vertex, and derive its induced distributions on incomplete covariance and precision coordinates. The distribution is proper exactly when \(α_i>|\mathrm{pa}(i)|+2\) for every vertex, and its nodewise conditional-variance and regression parameters are independent across vertices. This factorization gives a closed-form normalizing constant, conjugate updating, marginal likelihoods, and explicit full-matrix posterior means. We distinguish these posterior means from nonlinear completions of incomplete-coordinate means. We also distinguish the transformed Cholesky-coordinate mode from modes defined using covariance or precision coordinates. The model-selection procedure searches only over DAGs compatible with the specified ordering. Historical simulation and data examples illustrate the method; their evidentiary limitations and reproducibility requirements are stated explicitly.

math.ST↗

Large Scale Partial Correlation Screening with Uncertainty Quantification

Identifying multivariate dependencies in high-dimensional data is an important problem in large-scale inference. This problem has motivated recent advances in mining (partial) correlations, which focus on the challenging ultra-high dimensional setting where the sample size, n, is fixed, while the number of features, p, grows without bound. The state-of-the-art method for partial correlation screening can lead to undesirable results. This paper introduces a novel principled framework for partial correlation screening with error control (PARSEC), which leverages the connection between partial correlations and regression coefficients. We establish the inferential properties of PARSEC when n is fixed and p grows super-exponentially. First, we provide "fixed-n-large-p" asymptotic expressions for the familywise error rate (FWER) and k-FWER. Equally importantly, our analysis leads to a novel discovery which permits the calculation of exact marginal p-values for controlling the false discovery rate (FDR), and also the positive FDR (pFDR). To our knowledge, no other competing approach in the "fixed-n large-p" setting allows for error control across the spectrum of multiple hypothesis testing metrics. We establish the computational complexity of PARSEC and rigorously demonstrate its scalability to the large p setting. The theory and methods are successfully validated on simulated and real data, and PARSEC is shown to outperform the current state-of-the-art.

stat.ME↗

Hierarchical Relational Learning for Few-Shot Knowledge Graph Completion

Knowledge graphs (KGs) are powerful in terms of their inference abilities, but are also notorious for their incompleteness and long-tail distribution of relations. To address these challenges and expand the coverage of KGs, few-shot KG completion aims to make predictions for triplets involving novel relations when only a few training triplets are provided as reference. Previous methods have focused on designing local neighbor aggregators to learn entity-level information and/or imposing a potentially invalid sequential dependency assumption at the triplet level to learn meta relation information. However, pairwise triplet-level interactions and context-level relational information have been largely overlooked for learning meta representations of few-shot relations. In this paper, we propose a hierarchical relational learning method (HiRe) for few-shot KG completion. By jointly capturing three levels of relational information (entity-level, triplet-level and context-level), HiRe can effectively learn and refine meta representations of few-shot relations, and thus generalize well to new unseen relations. Extensive experiments on benchmark datasets validate the superiority of HiRe over state-of-the-art methods. The code can be found in https://github.com/alexhw15/HiRe.git.

cs.LG↗

Scalable and non-iterative graphical model estimation

Graphical models have found widespread applications in many areas of modern statistics and machine learning. Iterative Proportional Fitting (IPF) and its variants have become the default method for undirected graphical model estimation, and are thus ubiquitous in the field. As the IPF is an iterative approach, it is not always readily scalable to modern high-dimensional data regimes. In this paper we propose a novel and fast non-iterative method for positive definite graphical model estimation in high dimensions, one that directly addresses the shortcomings of IPF and its variants. In addition, the proposed method has a number of other attractive properties. First, we show formally that as the dimension p grows, the proportion of graphs for which the proposed method will outperform the state-of-the-art in terms of computational complexity and performance tends to 1, affirming its efficacy in modern settings. Second, the proposed approach can be readily combined with scalable non-iterative thresholding-based methods for high-dimensional sparsity selection. Third, the proposed method has high-dimensional statistical guarantees. Moreover, our numerical experiments also show that the proposed method achieves scalability without compromising on statistical precision. Fourth, unlike the IPF, which depends on the Gaussian likelihood, the proposed method is much more robust.

stat.ME↗

Preserving positivity for rank-constrained matrices

Entrywise functions preserving the cone of positive semidefinite matrices have been studied by many authors, most notably by Schoenberg [Duke Math. J. 9, 1942] and Rudin [Duke Math. J. 26, 1959]. Following their work, it is well-known that entrywise functions preserving Loewner positivity in all dimensions are precisely the absolutely monotonic functions. However, there are strong theoretical and practical motivations to study functions preserving positivity in a fixed dimension $n$. Such characterizations for a fixed value of $n$ are difficult to obtain, and in fact are only known in the $2 \times 2$ case. In this paper, using a novel and intuitive approach, we study entrywise functions preserving positivity on distinguished submanifolds inside the cone obtained by imposing rank constraints. These rank constraints are prevalent in applications, and provide a natural way to relax the elusive original problem of preserving positivity in a fixed dimension. In our main result, we characterize entrywise functions mapping $n \times n$ positive semidefinite matrices of rank at most $l$ into positive semidefinite matrices of rank at most $k$ for $1 \leq l \leq n$ and $1 \leq k < n$. We also demonstrate how an important necessary condition for preserving positivity by Horn and Loewner [Trans. Amer. Math. Soc. 136, 1969] can be significantly generalized by adding rank constraints. Finally, our techniques allow us to obtain an elementary proof of the classical characterization of functions preserving positivity in all dimensions obtained by Schoenberg and Rudin.

math.FA↗

Pseudo-likelihood Estimators for Graphical Models: Existence and Uniqueness

Graphical and sparse (inverse) covariance models have found widespread use in modern sample-starved high dimensional applications. A part of their wide appeal stems from the significantly low sample sizes required for the existence of estimators, especially in comparison with the classical full covariance model. For undirected Gaussian graphical models, the minimum sample size required for the existence of maximum likelihood estimators had been an open question for almost half a century, and has been recently settled. The very same question for pseudo-likelihood estimators has remained unsolved ever since their introduction in the '70s. Pseudo-likelihood estimators have recently received renewed attention as they impose fewer restrictive assumptions and have better computational tractability, improved statistical performance, and appropriateness in modern high dimensional applications, thus renewing interest in this longstanding problem. In this paper, we undertake a comprehensive study of this open problem within the context of the two classes of pseudo-likelihood methods proposed in the literature. We provide a precise answer to this question for both pseudo-likelihood approaches and relate the corresponding solutions to their Gaussian counterpart.

math.ST↗

The Khinchin-Kahane and Levy inequalities for abelian metric groups, and transfer from normed (abelian semi)groups to Banach spaces

The Khinchin-Kahane inequality is a fundamental result in the probability literature, with the most general version to date holding in Banach spaces. Motivated by modern settings and applications, we generalize this inequality to arbitrary metric groups which are abelian. If instead of abelian one assumes the group's metric to be a norm (i.e., $\mathbb{Z}_{>0}$-homogeneous), then we explain how the inequality improves to the same one as in Banach spaces. This occurs via a "transfer principle" that helps carry over questions involving normed metric groups and abelian normed semigroups into the Banach space framework. This principle also extends the notion of the expectation to random variables with values in arbitrary abelian normed metric semigroups $\mathscr{G}$. We provide additional applications, including studying weakly $\ell_p$ $\mathscr{G}$-valued sequences and related Rademacher series. On a related note, we also formulate a "general" Levy inequality, with two features: (i) It subsumes several known variants in the Banach space literature; and (ii) We show the inequality in the minimal framework required to state it: abelian metric groups.

math.PR↗

A unified framework for correlation mining in ultra-high dimension

Many applications benefit from theory relevant to the identification of variables having large correlations or partial correlations in high dimension. Recently there has been progress in the ultra-high dimensional setting when the sample size $n$ is fixed and the dimension $p$ tends to infinity. Despite these advances, the correlation screening framework suffers from practical, methodological and theoretical deficiencies. For instance, previous correlation screening theory requires that the population covariance matrix be sparse and block diagonal. This block sparsity assumption is however restrictive in practical applications. As a second example, correlation and partial correlation screening requires the estimation of dependence measures, which can be computationally prohibitive. In this paper, we propose a unifying approach to correlation and partial correlation mining that is not restricted to block diagonal correlation structure, thus yielding a methodology that is suitable for modern applications. By making connections to random geometric graphs, the number of highly correlated or partial correlated variables are shown to have compound Poisson finite-sample characterizations, which hold for both the finite $p$ case and when $p$ tends to infinity. The unifying framework also demonstrates a duality between correlation and partial correlation screening with theoretical and practical consequences.

math.ST↗

Differential calculus on the space of countable labelled graphs

The study of very large graphs is a prominent theme in modern-day mathematics. In this paper we develop a rigorous foundation for studying the space of finite labelled graphs and their limits. These limiting objects are naturally countable graphs, and the completed graph space $\mathscr{G}(V)$ is identified with the 2-adic integers as well as the Cantor set. The goal of this paper is to develop a model for differentiation on graph space in the spirit of the Newton-Leibnitz calculus. To this end, we first study the space of all finite labelled graphs and their limiting objects, and establish analogues of left-convergence, homomorphism densities, a Counting Lemma, and a large family of topologically equivalent metrics on labelled graph space. We then establish results akin to the First and Second Derivative Tests for real-valued functions on countable graphs, and completely classify the permutation automorphisms of graph space that preserve its topological and differential structures.

math.CO↗

Probability inequalities and tail estimates for metric semigroups

We study probability inequalities leading to tail estimates in a general semigroup $\mathscr{G}$ with a translation-invariant metric $d_{\mathscr{G}}$. (An important and central example of this in the functional analysis literature is that of $\mathscr{G}$ a Banach space.) Using our prior work [Ann. Prob. 2017] that extends the Hoffmann-Jorgensen inequality to all metric semigroups, we obtain tail estimates and approximate bounds for sums of independent semigroup-valued random variables, their moments, and decreasing rearrangements. In particular, we obtain the "correct" universal constants in several cases, extending results in the Banach space literature by Johnson-Schechtman-Zinn [Ann. Prob. 1985], Hitczenko [Ann. Prob. 1994], and Hitczenko and Montgomery-Smith [Ann. Prob. 2001]. Our results also hold more generally, in a very primitive mathematical framework required to state them: metric semigroups $\mathscr{G}$. This includes all compact, discrete, or (connected) abelian Lie groups.

math.PR↗

The critical exponent: a novel graph invariant

A surprising result of FitzGerald and Horn (1977) shows that $A^{\circ α} := (a_{ij}^α)$ is positive semidefinite (p.s.d.) for every entrywise nonnegative $n \times n$ p.s.d. matrix $A = (a_{ij})$ if and only if $α$ is a positive integer or $α\geq n-2$. Given a graph $G$, we consider the refined problem of characterizing the set $\mathcal{H}_G$ of entrywise powers preserving positivity for matrices with a zero pattern encoded by $G$. Using algebraic and combinatorial methods, we study how the geometry of $G$ influences the set $\mathcal{H}_G$. Our treatment provides new and exciting connections between combinatorics and analysis, and leads us to introduce and compute a new graph invariant called the critical exponent.

math.CO↗

The Hoffmann-Jorgensen inequality in metric semigroups

We prove a refinement of the inequality by Hoffmann-Jorgensen that is significant for three reasons. First, our result improves on the state-of-the-art even for real-valued random variables. Second, the result unifies several versions in the Banach space literature, including those by Johnson and Schechtman [Ann. Probab. 17 (1989)], Klass and Nowicki [Ann. Probab. 28 (2000)], and Hitczenko and Montgomery-Smith [Ann. Probab. 29 (2001)]. Finally, we show that the Hoffmann-Jorgensen inequality (including our generalized version) holds not only in Banach spaces but more generally, in a very primitive mathematical framework required to state the inequality: a metric semigroup $\mathscr{G}$. This includes normed linear spaces as well as all compact, discrete, or (connected) abelian Lie groups.

math.PR↗

From Random Walks to Random Leaps: Generalizing Classic Markov Chains for Big Data Applications

Simple random walks are a basic staple of the foundation of probability theory and form the building block of many useful and complex stochastic processes. In this paper we study a natural generalization of the random walk to a process in which the allowed step sizes take values in the set $\{\pm1,\pm2,\ldots,\pm k\}$, a process we call a random leap. The need to analyze such models arises naturally in modern-day data science and so-called "big data" applications. We provide closed-form expressions for quantities associated with first passage times and absorption events of random leaps. These expressions are formulated in terms of the roots of the characteristic polynomial of a certain recurrence relation associated with the transition probabilities. Our analysis shows that the expressions for absorption probabilities for the classical simple random walk are a special case of a universal result that is very elegant. We also consider an important variant of a random leap: the reflecting random leap. We demonstrate that the reflecting random leap exhibits more interesting behavior in regard to the existence of a stationary distribution and properties thereof. Questions relating to recurrence/transience are also addressed, as well as an application of the random leap.

math.PR↗

Scalable Bayesian shrinkage and uncertainty quantification for high-dimensional regression

Bayesian shrinkage methods have generated a lot of recent interest as tools for high-dimensional regression and model selection. These methods naturally facilitate tractable uncertainty quantification and incorporation of prior information. This benefit has led to extensive use of the Bayesian shrinkage methods across diverse applications. A common feature of these models is that the corresponding priors on the regression coefficients can be expressed as scale mixture of normals. While the three-step Gibbs sampler used to sample from the often intractable associated posterior density has been shown to be geometrically ergodic for several of these models, it has been demonstrated recently that convergence of this sampler can still be quite slow in modern high-dimensional settings despite this apparent theoretical safeguard. We propose a new method to draw from the same posterior via a tractable two-step blocked Gibbs sampler. We demonstrate that our proposed two-step blocked sampler exhibits vastly superior convergence behavior compared to the original three-step sampler in high-dimensional regimes on both real and simulated data. We also provide a detailed theoretical underpinning to the new method in the context of the Bayesian lasso. First, we prove that the proposed two-step sampler is geometrically ergodic, and derive explicit upper bounds for the (geometric) rate of convergence. Furthermore, we demonstrate theoretically that while the original Bayesian lasso chain is not Hilbert-Schmidt, the proposed chain is trace class (and hence Hilbert-Schmidt). The trace class property implies that the corresponding Markov operator is compact, and its (countably many) eigenvalues are summable. It also facilitates a rigorous comparison of the two-step blocked chain with "sandwich" algorithms which aim to improve performance of the two-step chain by inserting an inexpensive extra step.

stat.ME↗

Scalable Bayesian shrinkage and uncertainty quantification in high-dimensional regression

Bayesian shrinkage methods have generated a lot of recent interest as tools for high-dimensional regression and model selection. These methods naturally facilitate tractable uncertainty quantification and incorporation of prior information. A common feature of these models, including the Bayesian lasso, global-local shrinkage priors, and spike-and-slab priors is that the corresponding priors on the regression coefficients can be expressed as scale mixture of normals. While the three-step Gibbs sampler used to sample from the often intractable associated posterior density has been shown to be geometrically ergodic for several of these models (Khare and Hobert, 2013; Pal and Khare, 2014), it has been demonstrated recently that convergence of this sampler can still be quite slow in modern high-dimensional settings despite this apparent theoretical safeguard. We propose a new method to draw from the same posterior via a tractable two-step blocked Gibbs sampler. We demonstrate that our proposed two-step blocked sampler exhibits vastly superior convergence behavior compared to the original three- step sampler in high-dimensional regimes on both real and simulated data. We also provide a detailed theoretical underpinning to the new method in the context of the Bayesian lasso. First, we derive explicit upper bounds for the (geometric) rate of convergence. Furthermore, we demonstrate theoretically that while the original Bayesian lasso chain is not Hilbert-Schmidt, the proposed chain is trace class (and hence Hilbert-Schmidt). The trace class property has useful theoretical and practical implications. It implies that the corresponding Markov operator is compact, and its eigenvalues are summable. It also facilitates a rigorous comparison of the two-step blocked chain with "sandwich" algorithms which aim to improve performance of the two-step chain by inserting an inexpensive extra step.

stat.CO↗

Generalized Pseudolikelihood Methods for Inverse Covariance Estimation

We introduce PseudoNet, a new pseudolikelihood-based estimator of the inverse covariance matrix, that has a number of useful statistical and computational properties. We show, through detailed experiments with synthetic and also real-world finance as well as wind power data, that PseudoNet outperforms related methods in terms of estimation error and support recovery, making it well-suited for use in a downstream application, where obtaining low estimation error can be important. We also show, under regularity conditions, that PseudoNet is consistent. Our proof assumes the existence of accurate estimates of the diagonal entries of the underlying inverse covariance matrix; we additionally provide a two-step method to obtain these estimates, even in a high-dimensional setting, going beyond the proofs for related methods. Unlike other pseudolikelihood-based methods, we also show that PseudoNet does not saturate, i.e., in high dimensions, there is no hard limit on the number of nonzero entries in the PseudoNet estimate. We present a fast algorithm as well as screening rules that make computing the PseudoNet estimate over a range of tuning parameters tractable.

stat.ME↗

A convex framework for high-dimensional sparse Cholesky based covariance estimation

Covariance estimation for high-dimensional datasets is a fundamental problem in modern day statistics with numerous applications. In these high dimensional datasets, the number of variables p is typically larger than the sample size n. A popular way of tackling this challenge is to induce sparsity in the covariance matrix, its inverse or a relevant transformation. In particular, methods inducing sparsity in the Cholesky pa- rameter of the inverse covariance matrix can be useful as they are guaranteed to give a positive definite estimate of the covariance matrix. Also, the estimated sparsity pattern corresponds to a Directed Acyclic Graph (DAG) model for Gaussian data. In recent years, two useful penalized likelihood methods for sparse estimation of this Cholesky parameter (with no restrictions on the sparsity pattern) have been developed. How- ever, these methods either consider a non-convex optimization problem which can lead to convergence issues and singular estimates of the covariance matrix when p > n, or achieve a convex formulation by placing a strict constraint on the conditional variance parameters. In this paper, we propose a new penalized likelihood method for sparse estimation of the inverse covariance Cholesky parameter that aims to overcome some of the shortcomings of current methods, but retains their respective strengths. We ob- tain a jointly convex formulation for our objective function, which leads to convergence guarantees, even when p > n. The approach always leads to a positive definite and symmetric estimator of the covariance matrix. We establish high-dimensional estima- tion and graph selection consistency, and also demonstrate finite sample performance on simulated/real data.

stat.ME↗

Two-stage Sampling, Prediction and Adaptive Regression via Correlation Screening (SPARCS)

This paper proposes a general adaptive procedure for budget-limited predictor design in high dimensions called two-stage Sampling, Prediction and Adaptive Regression via Correlation Screening (SPARCS). SPARCS can be applied to high dimensional prediction problems in experimental science, medicine, finance, and engineering, as illustrated by the following. Suppose one wishes to run a sequence of experiments to learn a sparse multivariate predictor of a dependent variable $Y$ (disease prognosis for instance) based on a $p$ dimensional set of independent variables $\mathbf X=[X_1,\ldots, X_p]^T$ (assayed biomarkers). Assume that the cost of acquiring the full set of variables $\mathbf X$ increases linearly in its dimension. SPARCS breaks the data collection into two stages in order to achieve an optimal tradeoff between sampling cost and predictor performance. In the first stage we collect a few ($n$) expensive samples $\{y_i,\mathbf x_i\}_{i=1}^n$, at the full dimension $p\gg n$ of $\mathbf X$, winnowing the number of variables down to a smaller dimension $l < p$ using a type of cross-correlation or regression coefficient screening. In the second stage we collect a larger number $(t-n)$ of cheaper samples of the $l$ variables that passed the screening of the first stage. At the second stage, a low dimensional predictor is constructed by solving the standard regression problem using all $t$ samples of the selected variables. SPARCS is an adaptive online algorithm that implements false positive control on the selected variables, is well suited to small sample sizes, and is scalable to high dimensions. We establish asymptotic bounds for the Familywise Error Rate (FWER), specify high dimensional convergence rates for support recovery, and establish optimal sample allocation rules to the first and second stages.

stat.ML↗