SearcharxivSearch

arXiv subjects

Debangan Dey

Publications and source records attributed to Debangan Dey.

10 recordsLinked to original sources

Bigraphical Mat\'ern-Whittle (BMW) Processes for Fast Inference of Big Multivariate Spatial Data on General Domains

Large spatial data sets now record many correlated variables at many thousands of locations, often on domains where Euclidean distance misrepresents proximity. The central difficulty is modelling the cross-variable dependence jointly while retaining variable-level interpretation. We introduce the bigraphical Mat'ern-Whittle process, a multivariate Gaussian process that resolves this with two graphs. A spatial graph generates the Mat'ern structure of each variable through a fractional power of a graph Laplacian, so the process is valid on any topology, with per-variable range, smoothness and amplitude. A directed acyclic variable graph encodes the scientific structure: we prove that each absent edge yields an exact conditional independence between the corresponding fields. We further prove that the operator determinant does not involve the cross-dependence coefficients, which keeps matrix-free likelihood evaluation and Bayesian learning of the variable graph tractable at scale. Estimation requires only sparse matrix-vector products and scales to tens of millions of space-variable pairs. In simulations the method recovered parameters and graphs accurately, remained robust under misspecification, and halved held-out prediction error on a non-convex domain. In a spatial transcriptomics section with 19,809 cells and 1,122 genes, fitted in 75 minutes on a laptop, borrowing across the learned gene graph reduced held-out prediction error by 50 to 91 percent. Theoretical challenges, such as the achievable efficiency of estimating the variance of the nugget, are also explored.

stat.ME

Do In-Match Hydration Breaks Alter Match Momentum? A Within-Match Case-Crossover Analysis of the 2026 FIFA World Cup

The 2026 FIFA World Cup was the first to mandate hydration breaks, brief stoppages near the twenty-second minute of each half, in every match regardless of weather. A widespread claim holds that these breaks blunt the momentum of whichever side is dominant. We test the claim with a within-match case-crossover design, using non-break minutes of the same match as self-matched controls. This differences out every match-level characteristic and adjusts only for game time, scoreline, and pre-break momentum state. The apparent post-break fade is not caused by the break: the dominant side's momentum decays at essentially the same rate whether or not a break occurs, reverting through the stoppage as if play had continued. What the break removes instead is the continuation of that side's pressure during the stopped minutes, and this leaves no trace in net expected goals. The average effect is only weakly identified, resting on a trend extrapolation into a window with no untreated minutes, whereas break-by-covariate interactions are stable; that average is small and near zero. With only 99 matches the design lacks power to reject a small effect, so we read the results by effect size and confidence interval. The estimates are not uniform across match states. The clearest modifier is team strength: the break effect shifts by about +1.6 momentum points per 100 Elo points of rating advantage, favouring a stronger in-form side over a weaker side that is momentarily dominant, and it is the only statistically significant interaction. A weaker pattern appears by scoreline, where a comfortably-leading side of average strength shows an implied loss near -4 points on a pre-break level of about +20. We report these state-dependent signals as exploratory and in need of replication.

stat.AP

Doubly-Unlinked Regression for Dependent Data

Shuffled regression concerns settings in which covariates and responses are observed without their correct pairing. In dependent-data problems, a second form of missing correspondence can arise when responses are also detached from the latent temporal, spatial, or geometric domain that induces their dependence structure. We study regression under this joint loss of correspondence and, to our knowledge, provide the first systematic treatment of this setting. Specifically, we consider a doubly-unlinked regression model in which both the covariate-response link and the response-domain link are unknown, represented by two latent permutation matrices, while dependence is induced by an unobserved stochastic process. This framework unifies shuffled regression and latent-domain permutation models within a common dependent-data setting. We characterize signal-to-noise regimes governing recovery of the regression parameter and the latent permutations, and show that consistent estimation of the regression coefficient can be achieved under strictly weaker conditions than exact permutation recovery. To address the combinatorial difficulty of inference, we develop REPAIR, a variational Bayes method based on a block-structured permutation model that captures localized scrambling while substantially reducing computational complexity. Simulations and an applied example illustrate the empirical behavior of REPAIR and support the theoretical results.

math.ST

Multivariate Functional Principal Component Analysis for Mixed-Type mHealth Data: An Application to Mood Disorders

Modern mobile health (mHealth) assessment combines self-reported measures of participants' health experiences with passively collected health behavior data throughout the day. These data are collected across multiple measurement scales, including continuous (physical activity), truncated (pain), ordinal (mood), and binary (daily life events). When indexed by time of day and stacked across assessment domains, these data structures can be treated as multivariate functional data comprising continuous, truncated, ordinal, and binary variables. Motivated by these applications, we propose a multivariate functional principal component analysis for mixed-type data ($M^2$FPCA). The approach is based on a semiparametric Gaussian copula model and assumes that the observed data arise from an underlying multivariate generalized latent nonparanormal functional process. Latent temporal and inter-variable dependence are estimated semiparametrically through Kendall's tau bridging method. Two covariance estimation procedures are developed: a fully multivariate block-wise estimator and a computationally efficient alternative based on partial separability that assumes shared principal components across domains. The proposed method yields interpretable latent functional principal component scores that can serve as participant-specific digital biomarkers. Simulation studies demonstrate the method's competitive performance under various complex dependence structures. The method is applied to mHealth data from 307 participants in the National Institute of Mental Health Family Study of Mood and Affective Spectrum Disorders. Our approach identifies time-of-day patterns shared across mood, anxiety, energy, and physical activity that meaningfully stratify mood disorder subtypes.

stat.ME

Functional Principal Component Analysis for Continuous non-Gaussian, Truncated, and Discrete Functional Data

Mobile health studies often collect multiple within-day self-reported assessments of participants' behavior and well-being on different scales such as physical activity (continuous), pain levels (truncated), mood states (ordinal), and life events (binary). These assessments, when indexed by time of day, can be treated as functional data of different types - continuous, truncated, ordinal, and binary. We develop a functional principal component analysis that deals with all four types of functional data in a unified manner. It employs a semiparametric Gaussian copula model, assuming a generalized latent non-paranormal process as the underlying mechanism for these four types of functional data. We specify latent temporal dependence using a covariance estimated through Kendall's tau bridging method, incorporating smoothness during the bridging process. Simulation studies demonstrate the method's competitive performance under both dense and sparse sampling conditions. We then apply this approach to data from 497 participants in the National Institute of Mental Health Family Study of the Mood Disorder Spectrum to characterize within-day temporal patterns of mood differences among individuals with major mood disorder subtypes, including Major Depressive Disorder, Type 1, and Type 2 Bipolar Disorder.

stat.ME

Graph-constrained Analysis for Multivariate Functional Data

Functional Gaussian graphical models (GGM) used for analyzing multivariate functional data customarily estimate an unknown graphical model representing the conditional relationships between the functional variables. However, in many applications of multivariate functional data, the graph is known and existing functional GGM methods cannot preserve a given graphical constraint. In this manuscript, we demonstrate how to conduct multivariate functional analysis that exactly conforms to a given inter-variable graph. We first show the equivalence between partially separable functional GGM and graphical Gaussian processes (GP), proposed originally for constructing optimal covariance functions for multivariate spatial data that retain the conditional independence relations in a given graphical model. The theoretical connection help design a new algorithm that leverages Dempster's covariance selection to calculate the maximum likelihood estimate of the covariance function for multivariate functional data under graphical constraints. We also show that the finite term truncation of functional GGM basis expansion used in practice is equivalent to a low-rank graphical GP, which is known to oversmooth marginal distributions. To remedy this, we extend our algorithm to better preserve marginal distributions while still respecting the graph and retaining computational scalability. The insights obtained from the new results presented in this manuscript will help practitioners better understand the relationship between these graphical models and in deciding on the appropriate method for their specific multivariate data analysis task. The benefits of the proposed algorithms are illustrated using empirical experiments and an application to functional modeling of neuroimaging data using the connectivity graph among regions of the brain.

stat.ME

Regression and Dimension Reduction for Multivariate Mixed-Type Data via Semiparametric Gaussian Copula

Clinical and epidemiological studies encode participant information in multivariate vectors with mixed type variables on continuous, truncated, ordinal, and binary scales. Semiparametric Gaussian Copula (SGC) assumes that observed data is generated by latent multivariate normal random variables which marginals are monotonically transformed and then truncated/ordinalized/binarized. In SGC, the latent correlation matrix fully determines the dependence structure and it is estimated through an inversion of ``bridges'' between Kendall's Tau rank correlations of observed variables and latent correlations. By employing SGC, we develop regression (SGC-Reg), principal component analysis (SGC-PCA), and principal component regression (SGC-PCR) for latent representations of observed data. To build our framework, we make several key contributions: i) establishing novel bridging results for general ordinal type variables, ii) developing regression estimation on the latent space and deriving asymptotic normality of estimators, iii) developing a computationally efficient algorithm that reduces calculation complexity of all steps including calculation of asymptotic covariance matrix from $O(n^4)$ to $O(n\log n)$, iv) developing methods to predict latent representations of observed data and perform imputation of missing data, and v) developing principal component analysis and principal component regression on the latent space. We apply our framework to study the association between a 5-year mortality and 61 frailty-related measures composed of 29 continuous, 17 ordinal, and 15 binary variables in 9478 participants of 1999-2010 waves of National Health and Nutrition Examination Survey (NHANES).

stat.ME

Graphical Gaussian Process Models for Highly Multivariate Spatial Data

For multivariate spatial Gaussian process (GP) models, customary specifications of cross-covariance functions do not exploit relational inter-variable graphs to ensure process-level conditional independence among the variables. This is undesirable, especially for highly multivariate settings, where popular cross-covariance functions such as the multivariate Matérn suffer from a "curse of dimensionality" as the number of parameters and floating point operations scale up in quadratic and cubic order, respectively, in the number of variables. We propose a class of multivariate "Graphical Gaussian Processes" using a general construction called "stitching" that crafts cross-covariance functions from graphs and ensures process-level conditional independence among variables. For the Matérn family of functions, stitching yields a multivariate GP whose univariate components are Matérn GPs, and conforms to process-level conditional independence as specified by the graphical model. For highly multivariate settings and decomposable graphical models, stitching offers massive computational gains and parameter dimension reduction. We demonstrate the utility of the graphical Matérn GP to jointly model highly multivariate spatial data using simulation examples and an application to air-pollution modelling.

stat.ME

Connecting population-level AUC and latent scale-invariant $R^2$ via Semiparametric Gaussian Copula and rank correlations

Area Under the Curve (AUC) is arguably the most popular measure of classification accuracy. We use a semiparametric framework to introduce a latent scale-invariant $R^2$, a novel measure of variation explained for an observed binary outcome and an observed continuous predictor, and then directly link the latent $R^2$ to AUC. This enables a mutually consistent simultaneous use of AUC as a measure of classification accuracy and the latent $R^2$ as a scale-invariant measure of explained variation. Specifically, we employ Semiparametric Gaussian Copula (SGC) to model a joint dependence between observed binary outcome and observed continuous predictor via the correlation of latent standard normal random variables. Under SGC, we show how, both population-level AUC and latent scale-invariant $R^2$, defined as a squared latent correlation, can be estimated using any of the four rank statistics calculated on binary-continuous pairs: Wilcoxon rank-sum, Kendall's Tau, Spearman's Rho, and Quadrant rank correlations. We then focus on three implications and applications: i) we explicitly show that under SGC, the population-level AUC and the population-level latent $R^2$ are related via a monotone function that depends on the population-level prevalence rate, ii) we propose Quadrant rank correlation as a robust semiparametric version of AUC; iii) we demonstrate how, under complex-survey designs, Wilcoxon rank sum statistics and Spearman and Quadrant rank correlations provide asymptotically consistent estimators of the population-level AUC using only single-participant survey weights. We illustrate these applications using binary outcome of five-year mortality and continuous predictors including Albumin, Systolic Blood Pressure, and accelerometry-derived measures of total volume of physical activity collected in 2003-2006 National Health and Nutrition Examination Survey (NHANES) cohorts.

stat.ME

Spatial modeling of shot conversion in soccer to single out goalscoring ability

Goals are results of pin-point shots and it is a pivotal decision in soccer when, how and where to shoot. The main contribution of this study is two-fold. At first, after showing that there exists high spatial correlation in the data of shots across games, we introduce a spatial process in the error structure to model the probability of conversion from a shot depending on positional and situational covariates. The model is developed using a full Bayesian framework. Secondly, based on the proposed model, we define two new measures that can appropriately quantify the impact of an individual in soccer, by evaluating the positioning senses and shooting abilities of the players. As a practical application, the method is implemented on Major League Soccer data from 2016/17 season.

stat.AP