SearcharxivSearch

arXiv subjects

Subhabrata Sen

Publications and source records attributed to Subhabrata Sen.

At least 19 recordsLinked to original sources

DAIF: A Data-Driven Intermediate Fusion Framework for Multimodal Supervised Learning via Approximate Message Passing

Multimodal supervised learning seeks to leverage multiple heterogeneous data sources to improve predictive performance. A central challenge is determining the fusion granularity across modalities: over-integration may amplify noise while under-integration fails to exploit cross-modal dependence. Existing approaches rely on pre-specified fusion architectures, from early to late fusion, that may not adapt to the underlying dependence structure among modalities. We propose DAIF, a data adaptive intermediate fusion framework that combines random matrix theory and non-parametric dependence measures to learn fusion structure directly from data. We operate under a Bayesian multimodal factor model where the prior on the latent factors determines the cross-modal dependence. Our method clusters modalities based on estimated intermodal dependence, then performs clusterwise empirical Bayes estimation of the priors. These estimated priors are used to construct denoisers within an approximate message passing (AMP) framework, yielding denoised low-dimensional features that borrow strength across related modalities while preserving modality-specific signal. The resulting embeddings are used for downstream supervised prediction. We evaluate the framework through simulations under varying dependence structures and signal regimes, comparing against several benchmark methods, and demonstrate its practical utility on two multimodal datasets, namely a trimodal TEA-seq dataset (Swanson et al., 2021) and TCGA-BRCA dataset (Goldman et al., 2020). In the first example, we predict the expression level of a T-cell differentiation marker protein and in the second case we analyze patient survival prediction based on multimodal information. Our method competes with or outperforms the state-of-the-art techniques in both prediction problems, demonstrating its versatility across diverse supervised learning tasks.

stat.ME

Sharp Spectral Thresholds for Multi-View Spiked Wigner Models

Motivated by multimodal estimation, we study a multi-view spiked Wigner model in which several noisy matrix observations contain correlated latent spikes. We derive a spectral estimator for the latent spikes by linearizing approximate message passing (AMP). Our main result is an explicit sharp transition formula for its spectrum: for $L \geq 2$ views, letting $λ$ be the $L$-dimensional vector of spike strengths and $B$ the $L\times L$ limiting Gram matrix of the spikes, the critical parameter is $\mathsf{SNR}(λ,B)=λ_{\max}[\mathrm{Diag}(\sqrtλ) (B \odot B) \mathrm{Diag}(\sqrtλ)]$. When $\mathsf{SNR}(λ,B)<1$, the linearized AMP matrix has no outlier beyond the right edge of its bulk spectrum. When $\mathsf{SNR}(λ,B)>1$, an informative outlier is pinned at the distinguished point $1$, and the associated eigenvector has explicit, nontrivial overlaps with the latent signals. Thus $\mathsf{SNR}(λ,B)=1$ gives the exact spectral weak-recovery threshold for the linearized AMP method. To establish our results, we analyze the correlated Gaussian noise matrix through a matrix Dyson equation and combine this deterministic description with finite-rank perturbation arguments adapted to the multi-view spike structure. We also show that, for a broad class of spike priors, the spectral threshold $\mathsf{SNR}(λ,B)=1$ coincides with the information-theoretic threshold for weak recovery, ruling out a statistical-computational gap for this class of priors.

math.PR

Multi-layer Cross-attention is Provably Optimal for Multi-modal In-context Learning

Recent progress has rapidly advanced our understanding of the mechanisms underlying in-context learning in modern attention-based neural networks. However, existing results focus exclusively on unimodal data; in contrast, the theoretical underpinnings of in-context learning for multi-modal data remain poorly understood. We introduce a mathematically tractable framework for studying multi-modal learning and explore when transformer-like architectures can recover Bayes-optimal performance in-context. To model multi-modal problems, we assume the observed data arises from a latent factor model. Our first result comprises a negative take on expressibility: we prove that single-layer, linear self-attention fails to recover the Bayes-optimal predictor uniformly over the task distribution. To address this limitation, we introduce a novel, linearized cross-attention mechanism, which we study in the regime where both the number of cross-attention layers and the context length are large. We show that this cross-attention mechanism is provably Bayes optimal when optimized using gradient flow. Our results underscore the benefits of depth for in-context learning and establish the provable utility of cross-attention for multi-modal distributions.

stat.ML

High-Dimensional Statistics: Reflections on Progress and Open Problems

Over the past two decades, the field of high-dimensional statistics has experienced substantial progress, driven largely by technological advances that have dramatically reduced the cost and effort for data collection and storage across a broad range of domains, including biology, medicine, astronomy, and the social and environmental sciences. Modern datasets are increasingly complex, often exhibiting rich dependency, heterogeneity, and other features that challenge traditional statistical methods. In response, high-dimensional statistics has evolved to address more sophisticated estimation and inference problems. This evolution has, in turn, fostered deep connections with and contributions to a wide range of research areas, including optimization, concentration of measure, random matrix theory, information theory, and theoretical computer science. Given the rapid pace of recent developments in high-dimensional statistics, our goal is to synthesize representative advances, highlight common themes and open problems, and point to important works that offer entry points into the field.

math.ST

Causal inference under interference: computational barriers and algorithmic solutions

We study causal effect estimation under interference from network data. We work under the chain-graph formulation pioneered in Tchetgen Tchetgen et. al (2021). Our first result shows that polynomial time evaluation of treatment effects is computationally hard in this framework without additional assumptions on the underlying chain graph. Subsequently, we assume that the interactions among the study units are governed either by (i) a dense graph or (ii) an i.i.d. Gaussian matrix. In each case, we show that the treatment effects have well-defined limits as the population size diverges to infinity. Additionally, we develop polynomial time algorithms to consistently evaluate the treatment effects in each case. Finally, we estimate the unknown parameters from the observed data using maximum pseudo-likelihood estimates, and establish the stability of our causal effect estimators under this perturbation. Our algorithms provably approximate the causal effects in polynomial time even in low-temperature regimes where the canonical MCMC samplers are slow mixing. For dense graphs, our results use the notion of regularity partitions; for Gaussian interactions, our approach uses ideas from spin glass theory and Approximate Message Passing.

math.ST

De-Preferential Attachment Random Graphs

In this work we consider a growing random graph sequence where a new vertex is less likely to join to an existing vertex with high degree and more likely to join to a vertex with low degree. In contrast to the well studied \emph{preferential attachment random graphs} \cite{BarAlb99}, we call such a sequence a \emph{de-preferential attachment random graph model}. We consider two types of models, namely, \emph{inverse de-preferential}, where the attachment probabilities are inversely proportional to the degree and \emph{linear de-preferential}, where the attachment probabilities are proportional to $c-$degree, where $c > 0$ is a constant. For the case when each new vertex comes with exactly one half-edge we show that the degree of a fixed vertex is asymptotically of the order $\sqrt{\log n}$ for the inverse de-preferential case and of the order $\log n$ for the linear case. These show that compared to preferential attachment, the degree of a fixed vertex grows to infinity at a much slower rate for these models. We also show that in both cases limiting degree distributions have exponential tails. In fact we show that for the inverse de-preferential model the tail of the limiting degree distribution is faster than exponential while that for the linear de-preferential model is exactly the $\mbox{Geometric}\left(\frac{1}{2}\right)$ distribution. For the case when each new vertex comes with $m > 1$ half-edges, we show that similar asymptotic results hold for fixed vertex degree in both inverse and linear de-preferential models. Our proofs make use of the martingale approach as well as embedding to certain continuous time age dependent branching processes.

math.PR

Provable Benefits of Unsupervised Pre-training and Transfer Learning via Single-Index Models

Unsupervised pre-training and transfer learning are commonly used techniques to initialize training algorithms for neural networks, particularly in settings with limited labeled data. In this paper, we study the effects of unsupervised pre-training and transfer learning on the sample complexity of high-dimensional supervised learning. Specifically, we consider the problem of training a single-layer neural network via online stochastic gradient descent. We establish that pre-training and transfer learning (under concept shift) reduce sample complexity by polynomial factors (in the dimension) under very general assumptions. We also uncover some surprising settings where pre-training grants exponential improvement over random initialization in terms of sample complexity.

stat.ML

A large deviation principle for block models

We initiate a study of large deviations for block model random graphs in the dense regime. Following Chatterjee-Varadhan(2011), we establish an LDP for dense block models, viewed as random graphons. As an application of our result, we study upper tail large deviations for homomorphism densities of regular graphs. We identify the existence of a "symmetric" phase, where the graph, conditioned on the rare event, looks like a block model with the same block sizes as the generating graphon. In specific examples, we also identify the existence of a "symmetry breaking" regime, where the conditional structure is not a block model with compatible dimensions. This identifies a "reentrant phase transition" phenomenon for this problem -- analogous to one established for Erdos-Renyi random graphs (Chatterjee-Dey(2010), Chatterjee-Varadhan(2011)). Finally, extending the analysis of Lubetzky-Zhao(2015), we identify the precise boundary between the symmetry and symmetry breaking regime for homomorphism densities of regular graphs and the operator norm on Erdos-Renyi bipartite graphs.

math.PR

Understanding Optimal Feature Transfer via a Fine-Grained Bias-Variance Analysis

In the transfer learning paradigm models learn useful representations (or features) during a data-rich pretraining stage, and then use the pretrained representation to improve model performance on data-scarce downstream tasks. In this work, we explore transfer learning with the goal of optimizing downstream performance. We introduce a simple linear model that takes as input an arbitrary pretrained feature transform. We derive exact asymptotics of the downstream risk and its \textit{fine-grained} bias-variance decomposition. We then identify the pretrained representation that optimizes the asymptotic downstream bias and variance averaged over an ensemble of downstream tasks. Our theoretical and empirical analysis uncovers the surprising phenomenon that the optimal featurization is naturally sparse, even in the absence of explicit sparsity-inducing priors or penalties. Additionally, we identify a phase transition where the optimal pretrained representation shifts from hard selection to soft selection of relevant features.

stat.ML

Bayes optimal learning in high-dimensional linear regression with network side information

Supervised learning problems with side information in the form of a network arise frequently in applications in genomics, proteomics and neuroscience. For example, in genetic applications, the network side information can accurately capture background biological information on the intricate relations among the relevant genes. In this paper, we initiate a study of Bayes optimal learning in high-dimensional linear regression with network side information. To this end, we first introduce a simple generative model (called the Reg-Graph model) which posits a joint distribution for the supervised data and the observed network through a common set of latent parameters. Next, we introduce an iterative algorithm based on Approximate Message Passing (AMP) which is provably Bayes optimal under very general conditions. In addition, we characterize the limiting mutual information between the latent signal and the data observed, and thus precisely quantify the statistical impact of the network side information. Finally, supporting numerical experiments suggest that the introduced algorithm has excellent performance in finite samples.

math.ST

Causal effect estimation under network interference with mean-field methods

We study causal effect estimation from observational data under interference. The interference pattern is captured by an observed network. We adopt the chain graph framework of Tchetgen Tchetgen et. al. (2021), which allows (i) interaction among the outcomes of distinct study units connected along the graph and (ii) long range interference, whereby the outcome of an unit may depend on the treatments assigned to distant units connected along the interference network. For ``mean-field" interaction networks, we develop a new scalable iterative algorithm to estimate the causal effects. For gaussian weighted networks, we introduce a novel causal effect estimation algorithm based on Approximate Message Passing (AMP). Our algorithms are provably consistent under a ``high-temperature" condition on the underlying model. We estimate the (unknown) parameters of the model from data using maximum pseudo-likelihood and establish $\sqrt{n}$-consistency of this estimator in all parameter regimes. Finally, we prove that the downstream estimators obtained by plugging in estimated parameters into the aforementioned algorithms are consistent at high-temperature. Our methods can accommodate dense interactions among the study units -- a setting beyond reach using existing techniques. Our algorithms originate from the study of variational inference approaches in high-dimensional statistics; overall, we demonstrate the usefulness of these ideas in the context of causal effect estimation under interference.

math.ST

On Naive Mean-Field Approximation for high-dimensional canonical GLMs

We study the validity of the Naive Mean Field (NMF) approximation for canonical GLMs with product priors. This setting is challenging due to the non-conjugacy of the likelihood and the prior. Using the theory of non-linear large deviations (Austin 2019, Chatterjee, Dembo 2016, Eldan 2018), we derive sufficient conditions for the tightness of the NMF approximation to the log-normalizing constant of the posterior distribution. As a second contribution, we establish that under minor conditions on the design, any NMF optimizer is a product distribution where each component is a quadratic tilt of the prior. In turn, this suggests novel iterative algorithms for fitting the NMF optimizer to the target posterior. Finally, we establish that if the NMF optimization problem has a "well-separated maximizer", then this optimizer governs the probabilistic properties of the posterior. Specifically, we derive credible intervals with average coverage guarantees, and characterize the prediction performance on an out-of-sample datapoint in terms of this dominant optimizer.

math.ST

A Friendly Tutorial on Mean-Field Spin Glass Techniques for Non-Physicists

This tutorial is based on lecture notes written for a class taught in the Statistics Department at Stanford in the Winter Quarter of 2017. The objective was to provide a working knowledge of some of the techniques developed over the last 40 years by theoretical physicists and mathematicians to study mean field spin glasses and their applications to high-dimenensional statistics and statistical learning.

math.ST

Fundamental limits of community detection from multi-view data: multi-layer, dynamic and partially labeled block models

Multi-view data arises frequently in modern network analysis e.g. relations of multiple types among individuals in social network analysis, longitudinal measurements of interactions among observational units, annotated networks with noisy partial labeling of vertices etc. We study community detection in these disparate settings via a unified theoretical framework, and investigate the fundamental thresholds for community recovery. We characterize the mutual information between the data and the latent parameters, provided the degrees are sufficiently large. Based on this general result, (i) we derive a sharp threshold for community detection in an inhomogeneous multilayer block model \citep{chen2022global}, (ii) characterize a sharp threshold for weak recovery in a dynamic stochastic block model \citep{matias2017statistical}, and (iii) identify the limiting mutual information in an unbalanced partially labeled block model. Our first two results are derived modulo coordinate-wise convexity assumptions on specific functions -- we provide extensive numerical evidence for their correctness. Finally, we introduce iterative algorithms based on Approximate Message Passing for community detection in these problems.

math.ST

A Mean Field Approach to Empirical Bayes Estimation in High-dimensional Linear Regression

We study empirical Bayes estimation in high-dimensional linear regression. To facilitate computationally efficient estimation of the underlying prior, we adopt a variational empirical Bayes approach, introduced originally in Carbonetto and Stephens (2012) and Kim et al. (2022). We establish asymptotic consistency of the nonparametric maximum likelihood estimator (NPMLE) and its (computable) naive mean field variational surrogate under mild assumptions on the design and the prior. Assuming, in addition, that the naive mean field approximation has a dominant optimizer, we develop a computationally efficient approximation to the oracle posterior distribution, and establish its accuracy under the 1-Wasserstein metric. This enables computationally feasible Bayesian inference; e.g., construction of posterior credible intervals with an average coverage guarantee, Bayes optimal estimation for the regression coefficients, estimation of the proportion of non-nulls, etc. Our analysis covers both deterministic and random designs, and accommodates correlations among the features. To the best of our knowledge, this provides the first rigorous nonparametric empirical Bayes method in a high-dimensional regression setting without sparsity.

math.ST

Long ties accelerate noisy threshold-based contagions

Network structure can affect when and how widely new ideas, products, and behaviors are adopted. In widely-used models of biological contagion, interventions that randomly rewire edges (on average making them "longer") accelerate spread. However, there are other models relevant to social contagion, such as those motivated by myopic best-response in games with strategic complements, in which an individual's behavior is described by a threshold number of adopting neighbors above which adoption occurs (i.e., complex contagions). Recent work has argued that highly clustered, rather than random, networks facilitate spread of these complex contagions. Here we show that minor modifications to this model, which make it more realistic, reverse this result, thereby harmonizing qualitative facts about how network structure affects contagion. To model the trade-off between long and short edges we analyze the rate of spread over networks that are the union of circular lattices and random graphs on $n$ nodes. Allowing for noise in adoption decisions (i.e., adoptions below threshold) to occur with order ${n}^{-1/θ}$ probability along at least some "short" cycle edges is enough to ensure that random rewiring accelerates the spread of a noisy threshold-$θ$ contagion. This conclusion also holds under partial but frequent enough rewiring and when adoption decisions are reversible but infrequently so, as well as in high-dimensional lattice structures that facilitate faster-expanding contagions. Simulations illustrate the robustness of these results to several variations on this noisy best-response behavior. Hypothetical interventions that randomly rewire existing edges or add random edges (versus adding "short", triad-closing edges) in hundreds of empirical social networks reduce time to spread.

cs.SI

Sparse Signal Detection in Heteroscedastic Gaussian Sequence Models: Sharp Minimax Rates

Given a heterogeneous Gaussian sequence model with unknown mean $θ\in \mathbb R^d$ and known covariance matrix $Σ= \operatorname{diag}(σ_1^2,\dots, σ_d^2)$, we study the signal detection problem against sparse alternatives, for known sparsity $s$. Namely, we characterize how large $ε^*>0$ should be, in order to distinguish with high probability the null hypothesis $θ=0$ from the alternative composed of $s$-sparse vectors in $\mathbb R^d$, separated from $0$ in $L^t$ norm ($t \in [1,\infty]$) by at least $ε^*$. We find minimax upper and lower bounds over the minimax separation radius $ε^*$ and prove that they are always matching. We also derive the corresponding minimax tests achieving these bounds. Our results reveal new phase transitions regarding the behavior of $ε^*$ with respect to the level of sparsity, to the $L^t$ metric, and to the heteroscedasticity profile of $Σ$. In the case of the Euclidean (i.e. $L^2$) separation, we bridge the remaining gaps in the literature.

math.ST

Spectral Universality of Regularized Linear Regression with Nearly Deterministic Sensing Matrices

It has been observed that the performances of many high-dimensional estimation problems are universal with respect to underlying sensing (or design) matrices. Specifically, matrices with markedly different constructions seem to achieve identical performance if they share the same spectral distribution and have ``generic'' singular vectors. We prove this universality phenomenon for the case of convex regularized least squares (RLS) estimators under a linear regression model with additive Gaussian noise. Our main contributions are two-fold: (1) We introduce a notion of universality classes for sensing matrices, defined through a set of deterministic conditions that fix the spectrum of the sensing matrix and precisely capture the previously heuristic notion of generic singular vectors; (2) We show that for all sensing matrices that lie in the same universality class, the dynamics of the proximal gradient descent algorithm for solving the regression problem, as well as the performance of RLS estimators themselves (under additional strong convexity conditions) are asymptotically identical. In addition to including i.i.d. Gaussian and rotational invariant matrices as special cases, our universality class also contains highly structured, strongly correlated, or even (nearly) deterministic matrices. Examples of the latter include randomly signed versions of incoherent tight frames and randomly subsampled Hadamard transforms. As a consequence of this universality principle, the asymptotic performance of regularized linear regression on many structured matrices constructed with limited randomness can be characterized by using the rotationally invariant ensemble as an equivalent yet mathematically more tractable surrogate.

cs.IT