SearcharxivSearch

arXiv subjects

Pramita Bagchi

Publications and source records attributed to Pramita Bagchi.

8 recordsLinked to original sources

COMPLEX: A Closed-Form Certified Embedding of Multiparameter Persistence Modules

Every multiparameter persistence vectorization we know of carries a one-sided Lipschitz upper bound and nothing below it: without a lower gauge there is no sense in which the features are faithful, and no per-prediction guarantee can be built on them. This paper supplies the missing side. COMPLEX is a closed-form, training-free embedding of multiparameter modules -- slice the module along a fixed near-diagonal net, embed each slice barcode by the certified PLACE/PALACE landmark map, concatenate. Under a checkable witnessing-slice coherence condition, holding on 100% of audited pairs on Orbit5k, a single slice carries a closed-form lower gauge: separated modules stay separated in the embedding. With the standard upper bound this gives, to our knowledge, the first two-sided distortion bound for a multiparameter feature map, making faithfulness measurable. Measuring it, we find the floor tight within a small factor of realized distances yet operationally local: an RBF-SVM reaches 91% where 1-NN reaches 78% on the same features. Local per-prediction certification therefore fails for a structural reason common to every landmark embedding whose lower gauge is witnessed by one coordinate. With no learned embedding and no held-out calibration -- only a cross-validated SVM head -- COMPLEX sets the state of the art on both Orbit benchmarks (91.95% on Orbit5k, 92.98% on Orbit100k), level with or above Euler-characteristic surfaces and above transformers and graphcode. On graphs it exceeds GRIL on all four shared molecular benchmarks with one fixed configuration, including the only multiparameter method to clear COX2's majority baseline by more than three points. Closed-form selection -- of the landmark radius, the kernel (certificate-preserving), and the bifiltration set -- buys further accuracy; gradient-shaped adaptation buys none.

cs.LG

Statistical Inference for Persistence Diagrams via Landmark Embeddings: Minimax Theory and Finite Approximation

Hilbert-space embeddings enable inference for populations of persistence diagrams, but separation between individual diagrams need not survive population averaging. We develop a framework for inference on population mean embeddings, with particular attention to the additive landmark representations PLACE and PALACE. Treating each diagram as one independent observation, we apply Hilbert-space limit theory to obtain covariance estimators, two-sample tests, and confidence balls under suitable moment conditions, without requiring a lower-distortion bound. For additive embeddings, we identify the population mean as an embedding of the mean counting measure and show that geometric separation of these measures alone cannot guarantee uniform testing power. We then introduce a model with latent template diagrams, missing features, and location perturbations. Under common or feature-specific prevalence conditions, a diagram-level lower-distortion certificate yields explicit lower bounds on population mean separation. These margins provide finite-sample uniform power guarantees, and an additional information-divergence comparison gives matching sample-complexity bounds over restricted scale ranges. Confidence sets yield lower bounds on transport separation of population mean measures and exclusion guarantees for specified structured alternatives. We also quantify how orthogonal truncation changes the certified signal and the approximation allowance needed for confidence sets targeting the full embedding, relating sample size, retained coordinates, and template separation. Simulations examine calibration, power, and coverage, and an analysis of resting-state connectivity from the Autism Brain Imaging Data Exchange illustrates the procedures.

math.ST

A Closed-Form Persistence-Landmark Pipeline for Certified Point-Cloud and Graph Classification

We introduce PLACE (Persistence-Landmark Analytic Classification Engine), a closed-form pipeline for classifying point clouds and graphs through their persistent-homology signatures. Three quantitative guarantees -- a margin-based excess-risk rate, a closed-form descriptor-selection rule, and a per-prediction certificate -- are derived from training labels alone, with no learned weights or held-out calibration. The embedding sums Mitra-Virk single-point coordinate functions over a sparse landmark grid; the closed-form weight rule $w_k^2 \propto (d_{k+1}^2 - d_k^2)/R_k^2$ maximizes the distortion slope in Mitra-Virk's affine certificate under $ν$-coherence. (i) An $O(kR/(Δ\sqrt{m_{\min}}))$ margin bound, driven by class-mean separation $Δ$ and embedding radius $R$, matched in the sample-starved regime $m \lesssim R/Δ$ by a Le Cam minimax lower bound. (ii) The Mahalanobis margin under Ledoit-Wolf-shrunk covariance is the strongest closed-form ranker on a 64-descriptor chemical-graph pool (mean Spearman $ρ= +0.56$ across 11 benchmarks, positive on 10 of 11); the isotropic surrogate $Δ/\sqrt{\ell}$ admits a closed-form selection-consistency rate on the homogeneous protein/social pools. (iii) A training-time-decided certificate, with no per-prediction overhead, in three concrete radii (Pinelis, Gaussian plug-in, and variance-aware Pinelis-Bernstein). Empirically, PLACE is the strongest diagram-based method on Orbit5k and matches the strongest topology-based baseline within statistical noise on MUTAG and COX2; remaining gaps fall into two diagnosable regimes (descriptor blindness on NCI1/NCI109; pool-coverage limits elsewhere). The Pinelis-Bernstein radius fires on 8 of the 12 benchmarks; on MUTAG the empirical and population nearest-centroid rules agree on every one of 940 held-out test predictions, validating the certificate's mechanism.

cs.LG

A Closed-Form Adaptive-Landmark Kernel for Certified Point-Cloud and Graph Classification

We introduce PALACE (Persistence Adaptive-Landmark Analytic Classification Engine), the data-adaptive companion to PLACE, paying a small cross-validation tier on three knobs (budget, radii, bandwidth; $\leq 5$ choices each). A cover-theoretic core (Lebesgue-number criterion on the landmark cover) yields four closed-form guarantees. (i) A structural lower distortion bound $λ(τ;ν)$ on $\mathcal{D}_n$ under cross-diagram non-interference, with a $(D/L)^2$ budget reduction over the uniform grid when diagrams concentrate. (ii) Equal weights $w_k = K^{-1/2}$ maximizing $λ$, and farthest-point-sampling positions $2$-approximating the optimal $k$-center covering radius; both derived from training labels alone, no gradient training. (iii) A kernel-RKHS classification rate $O((k-1)\sqrt{K}/(γ\sqrt{m_{\min}}))$ with binary necessity threshold $m = Ω(\sqrt K/γ)$ from a matching Le Cam lower bound, and a closed-form filtration-selection rule. The kernel-Mahalanobis margin $\hatρ_{\mathrm{Mah}}$ is the strongest closed-form ranker across the chemical-graph pool (mean Spearman $ρ\approx +0.60$); the isotropic surrogate $\hatγ/\sqrt{K}$ admits a selection-consistency rate, and $\widehatλ$ from (i) provides an independent data-level signal (positive on COX2 and PTC). (iv) A per-prediction certificate, in non-asymptotic Pinelis and asymptotic Gaussian forms, with no calibration split. Empirically, PALACE is the strongest closed-form diagram-based method on Orbit5k ($91.3 \pm 1.0\%$, matching Persformer), leads every diagram-based competitor on COX2 and MUTAG, and is competitive on DHFR (within 1 pp of ECP). At $8\times$ domain inflation, adaptive placement maintains $94\%$ while the uniform grid collapses to chance ($25\%$ on 4-class data).

cs.LG

Adaptive Frequency Band Analysis for Functional Time Series

The frequency-domain properties of nonstationary functional time series often contain valuable information. These properties are characterized through its time-varying power spectrum. Practitioners seeking low-dimensional summary measures of the power spectrum often partition frequencies into bands and create collapsed measures of power within bands. However, standard frequency bands have largely been developed through manual inspection of time series data and may not adequately summarize power spectra. In this article, we propose a framework for adaptive frequency band estimation of nonstationary functional time series that optimally summarizes the time-varying dynamics of the series. We develop a scan statistic and search algorithm to detect changes in the frequency domain. We establish theoretical properties of this framework and develop a computationally-efficient implementation. The validity of our method is also justified through numerous simulation studies and an application to analyzing electroencephalogram data in participants alternating between eyes open and eyes closed conditions.

stat.ME

A Test for Separability in Covariance Operators of Random Surfaces

The assumption of separability is a simplifying and very popular assumption in the analysis of spatio-temporal or hypersurface data structures. It is often made in situations where the covariance structure cannot be easily estimated, for example because of a small sample size or because of computational storage problems. In this paper we propose a new and very simple test to validate this assumption. Our approach is based on a measure of separability which is zero in the case of separability and positive otherwise. The measure can be estimated without calculating the full non-separable covariance operator. We prove asymptotic normality of the corresponding statistic with a limiting variance, which can easily be estimated from the available data. As a consequence quantiles of the standard normal distribution can be used to obtain critical values and the new test of separability is very easy to implement. In particular, our approach does neither require projections on subspaces generated by the eigenfunctions of the covariance operator, nor resampling procedures to obtain critical values nor distributional assumptions as used by other available methods of constructing tests for separability. We investigate the finite sample performance by means of a simulation study and also provide a comparison with the currently available methodology. Finally, the new procedure is illustrated analyzing wind speed and temperature data.

stat.ME

A simple test for white noise in functional time series

We propose a new procedure for white noise testing of a functional time series. Our approach is based on an explicit representation of the $L^2$-distance between the spectral density operator and its best ($L^2$-)approximation by a spectral density operator corresponding to a white noise process. The estimation of this distance can be easily accomplished by sums of periodogram kernels and it is shown that an appropriately standardized version of the estimator is asymptotically normal distributed under the null hypothesis (of functional white noise) and under the alternative. As a consequence we obtain a very simple test (using the quantiles of the normal distribution) for the hypothesis of a white noise functional process. In particular the test does neither require the estimation of a long run variance (including a fourth order cumulant) nor resampling procedures to calculate critical values. Moreover, in contrast to all other methods proposed in the literature our approach also allows to test for "relevant" deviations from white noise and to construct confidence intervals for a measure which measures the discrepancy of the underlying process from a functional white noise process.

math.ST

Inference for Monotone Trends Under Dependence

We focus on the problem estimating a monotone trend function under additive and dependent noise. New point-wise confidence interval estimators under both short- and long-range dependent errors are introduced and studied. These intervals are obtained via the method of inversion of certain discrepancy statistics arising in hypothesis testing problems. The advantage of this approach is that it avoids the estimation of nuisance parameters such as the derivative of the unknown function, which existing methods are forced to deal with. While the methodology is motivated by earlier work in the independent context, the dependence of the errors, especially longrange dependence leads to new challenges, such as the study of convex minorants of drifted fractional Brownian motion that may be of independent interest. We also unravel a new family of universal limit distributions (and tabulate selected quantiles) that can henceforth be used for inference in monotone function problems involving dependence.

math.ST