SearcharxivSearch

arXiv subjects

Paromita Dubey

Publications and source records attributed to Paromita Dubey.

14 recordsLinked to original sources

Iterative Exploration-Driven Sparse SDP Clustering via Thompson Sampling

High-dimensional sparse clustering is a challenging combinatorial problem due to the tight coupling between cluster assignment and variable selection. We demonstrate that semidefinite programming (SDP) relaxation of K-means is robust to variable over-selection by establishing minimax separation bounds. Leveraging this robustness, we offer a twist on the recent trend of using exogenous randomness to evaluate feature importance and control false discovery by partially shifting the objective toward active feature exploration. We propose a randomized alternating block-coordinate framework that alternates between SDP clustering and feature selection. Here, feature selection is formulated as a multi-armed bandit problem, solved via Thompson Sampling with stochastic rewards generated by clustering and robust permutation tests. This approach introduces adaptive memory by aggregating historical variable-selection outcomes into posterior distributions, and selects features via posterior sampling, enabling stochastic exploration that promotes the inclusion of underexplored features and facilitates escape from local maxima. We establish uniform exact-recovery guarantees for adaptively selected feature sets. We extend the method to settings with unknown covariance through a scalable estimation procedure. Synthetic experiments and a real-data application in document clustering show that the proposed memory-driven randomized approach outperforms state-of-the-art sparse clustering methods across the studied settings.

stat.ME

DFNN: A Deep Fréchet Neural Network Framework for Learning Metric-Space-Valued Responses

Regression with non-Euclidean responses---e.g., probability distributions, networks, symmetric positive-definite matrices, and compositions---has become increasingly important in modern applications. In this paper, we propose deep Fréchet neural networks (DFNNs), an end-to-end deep learning framework for predicting non-Euclidean responses---which are considered as random objects in a metric space---from Euclidean predictors. Our method utilizes the representation-learning power of deep neural networks (DNNs) to the task of approximating conditional Fréchet means of the response given the predictors, the metric-space analogue of conditional expectations, by minimizing a Fréchet risk. The framework is highly flexible, accommodating diverse metrics and high-dimensional predictors. We establish a universal approximation theorem for DFNNs, advancing the state-of-the-art of neural network approximation theory to general metric-space-valued responses, without making model assumptions or relying on local smoothing. We further establish rigorous generalization guarantees for DFNNs and derive corresponding risk bounds, providing, to the best of our knowledge, the first such theoretical results for deep learning regression with metric-space-valued responses. Empirical studies on synthetic distributional and network-valued responses, as well as real-world applications to predicting compositional responses in an Aitchison simplex and spherical responses, demonstrate that DFNNs consistently outperform all existing methods.

stat.ML

DiPMInd: Distance Profile based Mutual Independence testing for random objects

This paper develops a novel unified framework for testing mutual independence among random objects residing in possibly different metric spaces. The framework generalizes existing methodologies and introduces new measures of mutual independence, and proposes associated tests that achieve minimax rate optimality and exhibit strong empirical power. The foundation of the proposed tests is the new concept of joint distance profiles, which uniquely characterize the joint law of random objects under a mild condition on either the joint law or the metric spaces. Our test statistics quantify the difference of the joint distance profiles of each data point with respect to the joint law and the product of marginal laws of the vector of random objects. To enhance power, we consider integrating this difference with respect to different measures and incorporate flexible data-adaptive weight profiles in the test statistics. We derive the limiting distribution of the test statistics under the null hypothesis of mutual independence and show that the proposed tests with certain weight profiles are asymptotically distribution-free if the marginal distance profiles are continuous. Furthermore, we establish the consistency of the tests under sequences of alternative hypotheses converging to the null. For practical implementations, we employ a permutation scheme to approximate the $p$-values and provide theoretical guarantees that the permutation-based tests maintain type I error control under the null and achieve consistency under alternatives. We demonstrate the power of the proposed tests across various types of data objects through simulations and real data applications, where the new tests exhibit better performance compared with popular existing approaches.

stat.ME

Inference for Fréchet Regression

Linear regression is widely used to model relationships between responses and predictors. In modern applications, one encounters data where the responses are non-Euclidean random objects situated in a metric space, paired with Euclidean predictors. Global Fréchet regression generalizes linear regression to such general settings, however statistical inference has remained largely unexplored. We develop a significance test for the null hypothesis that the Fréchet regression function does not depend on the predictors, addressing the challenge of an absence of linear operations in metric spaces. We also develop a test for the partial effect of a subset of the predictors in analogy to, but quite different from, the partial F-tests commonly used in classical linear regression under Gaussian assumptions. Key ideas are to employ random multipliers to obtain non-degenerate null distributions for the proposed test statistics and the Cauchy combination method. We obtain consistency and convergence results under the null hypothesis and contiguous alternatives and demonstrate the finite sample performance of the proposed tests through simulations on network data represented by graph Laplacians and spherical data with geodesic distances. We further illustrate our method using transport networks arising from New York City taxi trip data and U.S. energy source compositional data.

stat.ME

Change Point Inference for Non-Euclidean Data Sequences using Distance Profiles

We introduce a powerful scan statistic and the corresponding test for detecting the presence and pinpointing the location of a change point within the distribution of a data sequence with the data elements residing in a separable metric space $(Ω, d)$. These change points mark abrupt shifts in the distribution of the data sequence as characterized using distance profiles, where the distance profile of an element $ω\in Ω$ is the distribution of distances from $ω$ as dictated by the data. This approach is tuning parameter free, fully non-parametric and universally applicable to diverse data types, including distributional and network data, as long as distances between the data objects are available. We obtain an explicit characterization of the asymptotic distribution of the test statistic under the null hypothesis of no change points, rigorous guarantees on the consistency of the test in the presence of change points under fixed and local alternatives and near-optimal convergence of the estimated change point location, all under practicable settings. To compare with state-of-the-art methods we conduct simulations covering multivariate data, bivariate distributional data and sequences of graph Laplacians, and illustrate our method on real data sequences of the U.S. electricity generation compositions and Bluetooth proximity networks.

stat.ME

LLmFPCA-detect: LLM-powered Multivariate Functional PCA for Anomaly Detection in Sparse Longitudinal Texts

Sparse longitudinal (SL) textual data arises when individuals generate text repeatedly over time (e.g., customer reviews, occasional social media posts, electronic medical records across visits), but the frequency and timing of observations vary across individuals. These complex textual data sets have immense potential to inform future policy and targeted recommendations. However, because SL text data lack dedicated methods and are noisy, heterogeneous, and prone to anomalies, detecting and inferring key patterns is challenging. We introduce LLmFPCA-detect, a flexible framework that pairs LLM-based text embeddings with functional data analysis to detect clusters and infer anomalies in large SL text datasets. First, LLmFPCA-detect embeds each piece of text into an application-specific numeric space using LLM prompts. Sparse multivariate functional principal component analysis (mFPCA) conducted in the numeric space forms the workhorse to recover primary population characteristics, and produces subject-level scores which, together with baseline static covariates, facilitate data segmentation, unsupervised anomaly detection and inference, and enable other downstream tasks. In particular, we leverage LLMs to perform dynamic keyword profiling guided by the data segments and anomalies discovered by LLmFPCA-detect, and we show that cluster-specific functional PC scores from LLmFPCA-detect, used as features in existing pipelines, help boost prediction performance. We support the stability of LLmFPCA-detect with experiments and evaluate it on two different applications using public datasets, Amazon customer-review trajectories, and Wikipedia talk-page comment streams, demonstrating utility across domains and outperforming state-of-the-art baselines.

stat.ML

Online network change point detection with missing values and temporal dependence

In this paper we study online change point detection in dynamic networks with time heterogeneous missing pattern within networks and dependence across the time course. The missingness probabilities, the entrywise sparsity of networks, the rank of networks and the jump size in terms of the Frobenius norm, are all allowed to vary as functions of the pre-change sample size. On top of a thorough handling of all the model parameters, we notably allow the edges and missingness to be dependent. To the best of our knowledge, such general framework has not been rigorously nor systematically studied before in the literature. We propose a polynomial time change point detection algorithm, with a version of soft-impute algorithm (e.g. Mazumder et al., 2010; Klopp, 2015) as the imputation sub-routine. Piecing up these standard sub-routines algorithms, we are able to solve a brand new problem with sharp detection delay subject to an overall Type-I error control. Extensive numerical experiments are conducted demonstrating the outstanding performances of our proposed method in practice.

stat.ME

Modeling Time-Varying Random Objects and Dynamic Networks

Samples of dynamic or time-varying networks and other random object data such as time-varying probability distributions are increasingly encountered in modern data analysis. Common methods for time-varying data such as functional data analysis are infeasible when observations are time courses of networks or other complex non-Euclidean random objects that are elements of general metric spaces. In such spaces, only pairwise distances between the data objects are available and a strong limitation is that one cannot carry out arithmetic operations due to the lack of an algebraic structure. We combat this complexity by a generalized notion of mean trajectory taking values in the object space. For this, we adopt pointwise Fréchet means and then construct pointwise distance trajectories between the individual time courses and the estimated Fréchet mean trajectory, thus representing the time-varying objects and networks by functional data. Functional principal component analysis of these distance trajectories can reveal interesting features of dynamic networks and object time courses and is useful for downstream analysis. Our approach also makes it possible to study the empirical dynamics of time-varying objects, including dynamic regression to the mean or explosive behavior over time. We demonstrate desirable asymptotic properties of sample based estimators for suitable population targets under mild assumptions. The utility of the proposed methodology is illustrated with dynamic networks, time-varying distribution data and longitudinal growth data.

stat.ME

Metric Statistics: Exploration and Inference for Random Objects With Distance Profiles

This article provides an overview on the statistical modeling of complex data as increasingly encountered in modern data analysis. It is argued that such data can often be described as elements of a metric space that satisfies certain structural conditions and features a probability measure. We refer to the random elements of such spaces as random objects and to the emerging field that deals with their statistical analysis as metric statistics. Metric statistics provides methodology, theory and visualization tools for the statistical description, quantification of variation, centrality and quantiles, regression and inference for populations of random objects, inferring these quantities from available data and samples. In addition to a brief review of current concepts, we focus on distance profiles as a major tool for object data in conjunction with the pairwise Wasserstein transports of the underlying one-dimensional distance distributions. These pairwise transports lead to the definition of intuitive and interpretable notions of transport ranks and transport quantiles as well as two-sample inference. An associated profile metric complements the original metric of the object space and may reveal important features of the object data in data analysis. We demonstrate these tools for the analysis of complex data through various examples and visualizations.

stat.ME

Learning delay dynamics for multivariate stochastic processes, with application to the prediction of the growth rate of COVID-19 cases in the United States

Delay differential equations form the underpinning of many complex dynamical systems. The forward problem of solving random differential equations with delay has received increasing attention in recent years. Motivated by the challenge to predict the COVID-19 caseload trajectories for individual states in the U.S., we target here the inverse problem. Given a sample of observed random trajectories obeying an unknown random differential equation model with delay, we use a functional data analysis framework to learn the model parameters that govern the underlying dynamics from the data. We show existence and uniqueness of the analytical solutions of the population delay random differential equation model when one has discrete time delays in the functional concurrent regression model and also for a second scenario where one has a delay continuum or distributed delay. The latter involves a functional linear regression model with history index. The derivative of the process of interest is modeled using the process itself as predictor and also other functional predictors with predictor-specific delayed impacts. This dynamics learning approach is shown to be well suited to model the growth rate of COVID-19 for the states that are part of the U.S., by pooling information from the individual states, using the case process and concurrently observed economic and mobility data as predictors.

math.AP

Fréchet Change Point Detection

We propose a method to infer the presence and location of change-points in the distribution of a sequence of independent data taking values in a general metric space, where change-points are viewed as locations at which the distribution of the data sequence changes abruptly in terms of either its Fréchet mean or Fréchet variance or both. The proposed method is based on comparisons of Fréchet variances before and after putative change-point locations and does not require a tuning parameter except for the specification of cut-off intervals near the endpoints where change-points are assumed not to occur. Our results include theoretical guarantees for consistency of the test under contiguous alternatives when a change-point exists and also for consistency of the estimated location of the change-point if it exists, where under the null hypothesis of no change-point the limit distribution of the proposed scan function is the square of a standardized Brownian Bridge. These consistency results are applicable for a broad class of metric spaces under mild entropy conditions. Examples include the space of univariate probability distributions and the space of graph Laplacians for networks. Simulation studies demonstrate the effectiveness of the proposed methods, both for inferring the presence of a change-point and estimating its location. We also develop theory that justifies bootstrap-based inference and illustrate the new approach with sequences of maternal fertility distributions and communication networks.

stat.ME

Functional Models for Time-Varying Random Objects

In recent years, samples of time-varying object data such as time-varying networks that are not in a vector space have been increasingly collected. These data can be viewed as elements of a general metric space that lacks local or global linear structure and therefore common approaches that have been used with great success for the analysis of functional data, such as functional principal component analysis, cannot be applied. In this paper we propose metric covariance, a novel association measure for paired object data lying in a metric space $(Ω,d)$ that we use to define a metric auto-covariance function for a sample of random $Ω$-valued curves, where $Ω$ generally will not have a vector space or manifold structure. The proposed metric auto-covariance function is non-negative definite when the squared semimetric $d^2$ is of negative type. Then the eigenfunctions of the linear operator with the auto-covariance function as kernel can be used as building blocks for an object functional principal component analysis for $Ω$-valued functional data, including time-varying probability distributions, covariance matrices and time-dynamic networks. Analogues of functional principal components for time-varying objects are obtained by applying Fréchet means and projections of distance functions of the random object trajectories in the directions of the eigenfunctions, leading to real-valued Fréchet scores. Using the notion of generalized Fréchet integrals, we construct object functional principal components that lie in the metric space $Ω$.

stat.ME

Fréchet Analysis Of Variance For Random Objects

Fréchet mean and variance provide a way of obtaining mean and variance for general metric space valued random variables and can be used for statistical analysis of data objects that lie in abstract spaces devoid of algebraic structure and operations. Examples of such spaces include covariance matrices, graph Laplacians of networks and univariate probability distribution functions. We derive a central limit theorem for Fréchet variance under mild regularity conditions, utilizing empirical process theory, and also provide a consistent estimator of the asymptotic variance. These results lead to a test to compare k populations based on Fréchet variance for general metric space valued data objects, with emphasis on comparing means and variances. We examine the finite sample performance of this inference procedure through simulation studies for several special cases that include probability distributions and graph Laplacians, which leads to tests to compare populations of networks. The proposed methodology has good finite sample performance in simulations for different kinds of random objects. We illustrate the proposed methods with data on mortality profiles of various countries and resting state Functional Magnetic Resonance Imaging data.

math.ST

Modeling Bimodal Discrete Data Using Conway-Maxwell-Poisson Mixture Models

Bimodal truncated count distributions are frequently observed in aggregate survey data and in user ratings when respondents are mixed in their opinion. They also arise in censored count data, where the highest category might create an additional mode. Modeling bimodal behavior in discrete data is useful for various purposes, from comparing shapes of different samples (or survey questions) to predicting future ratings by new raters. The Poisson distribution is the most common distribution for fitting count data and can be modified to achieve mixtures of truncated Poisson distributions. However, it is suitable only for modeling equi-dispersed distributions and is limited in its ability to capture bimodality. The Conway-Maxwell-Poisson (CMP) distribution is a two-parameter generalization of the Poisson distribution that allows for over- and under-dispersion. In this work, we propose a mixture of CMPs for capturing a wide range of truncated discrete data, which can exhibit unimodal and bimodal behavior. We present methods for estimating the parameters of a mixture of two CMP distributions using an EM approach. Our approach introduces a special two-step optimization within the M step to estimate multiple parameters. We examine computational and theoretical issues. The methods are illustrated for modeling ordered rating data as well as truncated count data, using simulated and real examples.

stat.ME