SearcharxivSearch

arXiv subjects

Radhendushka Srivastava

Publications and source records attributed to Radhendushka Srivastava.

15 recordsLinked to original sources

A structural equation formulation for general quasi-periodic Gaussian processes

This paper introduces a structural equation formulation that gives rise to a new family of quasi-periodic Gaussian processes, useful to process a broad class of natural and physiological signals. The proposed formulation simplifies generation and forecasting, and provides hyperparameter estimates, which we exploit in a convergent and consistent iterative estimation algorithm. A bootstrap approach for standard error estimation and confidence intervals is also provided. We demonstrate the computational and scaling benefits of the proposed approach on a broad class of problems, including water level tidal analysis, CO\textsubscript{2} emission data, and sunspot numbers data. By leveraging the structural equations, our method reduces the cost of likelihood evaluations and predictions from $\mathcal{O}(k^2 p^2)$ to $\mathcal{O}(p^2)$, significantly improving scalability.

stat.ME

Quasi-Periodic Gaussian Process Predictive Iterative Learning Control

Repetitive motion tasks are common in robotics, but performance can degrade over time due to environmental changes and robot wear and tear. Iterative learning control (ILC) improves performance by using information from previous iterations to compensate for expected errors in future iterations. This work incorporates the use of Quasi-Periodic Gaussian Processes (QPGPs) into a predictive ILC framework to model and forecast disturbances and drift across iterations. Using a recent structural equation formulation of QPGPs, the proposed approach enables efficient inference with complexity $\mathcal{O}(p^3)$ instead of $\mathcal{O}(i^2p^3)$, where $p$ denotes the number of points within an iteration and $i$ represents the total number of iterations, specially for larger $i$. This formulation also enables parameter estimation without loss of information, making continual GP learning computationally feasible within the control loop. By predicting next-iteration error profiles rather than relying only on past errors, the controller achieves faster convergence and maintains this under time-varying disturbances. We benchmark the method against both standard ILC and conventional Gaussian Process (GP)-based predictive ILC on three tasks, autonomous vehicle trajectory tracking, a three-link robotic manipulator, and a real-world Stretch robot experiment. Across all cases, the proposed approach converges faster and remains robust under injected and natural disturbances while reducing computational cost. This highlights its practicality across a range of repetitive dynamical systems.

cs.RO

Correction of Pooling Matrix Mis-specifications in Compressed Sensing Based Group Testing

Compressed sensing, which involves the reconstruction of sparse signals from an under-determined linear system, has been recently used to solve problems in group testing. In a public health context, group testing aims to determine the health status values of p subjects from n<<p pooled tests, where a pool is defined as a mixture of small, equal-volume portions of the samples of a subset of subjects. This approach saves on the number of tests administered in pandemics or other resource-constrained scenarios. In practical group testing in time-constrained situations, a technician can inadvertently make a small number of errors during pool preparation, which leads to errors in the pooling matrix, which we term `model mismatch errors' (MMEs). This poses difficulties while determining health status values of the participating subjects from the results on n<<p pooled tests. In this paper, we present an algorithm to correct the MMEs in the pooled tests directly from the pooled results and the available (inaccurate) pooling matrix. Our approach then reconstructs the signal vector from the corrected pooling matrix, in order to determine the health status of the subjects. We further provide theoretical guarantees for the correction of the MMEs and the reconstruction error from the corrected pooling matrix. We also provide several supporting numerical results.

stat.AP

A semi-parametric model for assessing the effect of temperature on ice accumulation rate from Antarctic ice core data

In this paper, we present a semiparametric model for describing the effect of temperature on Antarctic ice accumulation on a paleoclimatic time scale. The model is motivated by sharp ups and downs in the rate of ice accumulation apparent from ice core data records, which are synchronous with movements of temperature. We prove strong consistency of the estimators under reasonable conditions. We conduct extensive simulations to assess the performance of the estimators and bootstrap based standard errors and confidence limits for the requisite range of sample sizes. Analysis of ice core data from two Antarctic locations over several hundred thousand years shows a reasonable fit. The apparent accumulation rate exhibits a thinning pattern that should facilitate the understanding of ice condensation, transformation and flow over the ages. There is a very strong linear relationship between temperature and the apparent accumulation rate adjusted for thinning.

stat.ME

Robust Non-adaptive Group Testing under Errors in Group Membership Specifications

Given $p$ samples, each of which may or may not be defective, group testing (GT) aims to determine their defect status by performing tests on $n < p$ `groups', where a group is formed by mixing a subset of the $p$ samples. Assuming that the number of defective samples is very small compared to $p$, GT algorithms have provided excellent recovery of the status of all $p$ samples with even a small number of groups. Most existing methods, however, assume that the group memberships are accurately specified. This assumption may not always be true in all applications, due to various resource constraints. Such errors could occur, eg, when a technician, preparing the groups in a laboratory, unknowingly mixes together an incorrect subset of samples as compared to what was specified. We develop a new GT method, the Debiased Robust Lasso Test Method (DRLT), that handles such group membership specification errors. The proposed DRLT method is based on an approach to debias, or reduce the inherent bias in, estimates produced by Lasso, a popular and effective sparse regression technique. We also provide theoretical upper bounds on the reconstruction error produced by our estimator. Our approach is then combined with two carefully designed hypothesis tests respectively for (i) the identification of defective samples in the presence of errors in group membership specifications, and (ii) the identification of groups with erroneous membership specifications. The DRLT approach extends the literature on bias mitigation of statistical estimators such as the LASSO, to handle the important case when some of the measurements contain outliers, due to factors such as group membership specification errors. We present numerical results which show that our approach outperforms several baselines and robust regression techniques for identification of defective samples as well as erroneously specified groups.

stat.ML

Fast Debiasing of the LASSO Estimator

In high-dimensional sparse regression, the \textsc{Lasso} estimator offers excellent theoretical guarantees but is well-known to produce biased estimates. To address this, \cite{Javanmard2014} introduced a method to ``debias" the \textsc{Lasso} estimates for a random sub-Gaussian sensing matrix $\boldsymbol{A}$. Their approach relies on computing an ``approximate inverse" $\boldsymbol{M}$ of the matrix $\boldsymbol{A}^\top \boldsymbol{A}/n$ by solving a convex optimization problem. This matrix $\boldsymbol{M}$ plays a critical role in mitigating bias and allowing for construction of confidence intervals using the debiased \textsc{Lasso} estimates. However the computation of $\boldsymbol{M}$ is expensive in practice as it requires iterative optimization. In the presented work, we re-parameterize the optimization problem to compute a ``debiasing matrix" $\boldsymbol{W} := \boldsymbol{AM}^{\top}$ directly, rather than the approximate inverse $\boldsymbol{M}$. This reformulation retains the theoretical guarantees of the debiased \textsc{Lasso} estimates, as they depend on the \emph{product} $\boldsymbol{AM}^{\top}$ rather than on $\boldsymbol{M}$ alone. Notably, we provide a simple, computationally efficient, closed-form solution for $\boldsymbol{W}$ under similar conditions for the sensing matrix $\boldsymbol{A}$ used in the original debiasing formulation, with an additional condition that the elements of every row of $\boldsymbol{A}$ have uncorrelated entries. Also, the optimization problem based on $\boldsymbol{W}$ guarantees a unique optimal solution, unlike the original formulation based on $\boldsymbol{M}$. We verify our main result with numerical simulations.

stat.ML

Apparent ice accumulation rate in East Antarctica: Relation with temperature and thinning pattern

We present here formal evidence of a strong linkage between temperature and East Antarctic ice accumulation over the past eight hundred kiloyears, after accounting for thinning. The conclusions are based on statistical analysis of a proposed empirical model based on ice core data from multiple locations with ground topography ranging from local peaks to local valleys. The method permits adjustment of the apparent accumulation rate for a very general thinning process of ice sheet over the ages, is robust to any misspecification of the age scale, and does not require delineation of the accumulation rate from thinning. Records show 5% to 8% increase in the accumulation rate for every 1${}^\circ$C rise in temperature. This is consistent with the theoretical expectation on the average rate of increase in moisture absorption capacity of the atmosphere with rise in temperature, as inferred from the Clausius-Clapeyron equation. This finding reinforces indications of the resilience of the Antarctic Ice Sheet to the effects of warming induced by climate change, which have been documented in other studies based on recent data. Analysis of the thinning pattern of ice revealed an exponential rate of thinning over several glacial cycles and eventual attainment of a saturation level.

physics.ao-ph

Feature Sensitive Curve Registration by Kernel Matching

In this paper, we argue that the problem of registering two sets of functional data, where the underlying mean function has sharp features, is not properly addressed by methods designed to align a bunch of growth curves data. We provide a new method, which is able to pool local information without smoothing and to match sharp landmarks without manual identification. This method, which we refer to as kernel-matched registration, is based on maximizing a kernel-based measure of alignment. We prove that the proposed method is consistent under fairly general conditions. Simulation results show superiority of the performance of the proposed method over two existing methods. The proposed method is illustrated through the analysis of three sets of paleoclimatic data.

stat.ME

AR(1) sequence with random coefficients: Regenerative properties and its application

Let $\{X_n\}_{n\ge0}$ be a sequence of real valued random variables such that $X_n=ρ_n X_{n-1}+ε_n,~n=1,2,\ldots$, where $\{(ρ_n,ε_n)\}_{n\ge1}$ are i.i.d. and independent of initial value (possibly random) $X_0$. In this paper it is shown that, under some natural conditions on the distribution of $(ρ_1,ε_1)$, the sequence $\{X_n\}_{n\ge0}$ is regenerative in the sense that it could be broken up into i.i.d. components. Further, when $ρ_1$ and $ε_1$ are independent, we construct a non-parametric strongly consistent estimator of the characteristic functions of $ρ_1$ and $ε_1$.

math.PR

Feature Sensitive and Automated Curve Registration

Given two sets of functional data having a common underlying mean function but different degrees of distortion in time measurements, we provide a method of estimating the time transformation necessary to align (or `register') them. We prove that the proposed method is consistent under fairly general conditions. Simulation results show superiority of the performance of the proposed method over two existing methods. The proposed method is illustrated through the analysis of three paleoclimatic data sets.

stat.ME

RAPTT: An Exact Two-Sample Test in High Dimensions Using Random Projections

In high dimensions, the classical Hotelling's $T^2$ test tends to have low power or becomes undefined due to singularity of the sample covariance matrix. In this paper, this problem is overcome by projecting the data matrix onto lower dimensional subspaces through multiplication by random matrices. We propose RAPTT (RAndom Projection T-Test), an exact test for equality of means of two normal populations based on projected lower dimensional data. RAPTT does not require any constraints on the dimension of the data or the sample size. A simulation study indicates that in high dimensions the power of this test is often greater than that of competing tests. The advantage of RAPTT is illustrated on high-dimensional gene expression data involving the discrimination of tumor and normal colon tissues.

stat.ME

Effect of sampling on the estimation of drift parameter of continuous time AR(1) processes

We study the effect of stochastic sampling on the estimation of the drift parameter of continuous time AR(1) process. A natural distribution free moment estimator is considered for the drift based on stochastically observed time points. The effect of the constraint of the minimum separation between successive samples on the estimation of the drift is studied.

math.ST

Asymptotic distribution of a consistent cross-spectrum estimator based on uniformly spaced samples of a non-bandlimited process

It is well known that if the power spectral density of a continuous time stationary stochastic process does not have a compact support, data sampled from that process at any uniform sampling rate leads to biased and inconsistent spectrum estimators. In a recent paper, the authors showed that the smoothed periodogram estimator can be consistent, if the sampling interval is allowed to shrink to zero at a suitable rate as the sample size goes to infinity. In this paper, this `shrinking asymptotics' approach is used to obtain the limiting distribution of the smoothed periodogram estimator of spectra and cross-spectra. It is shown that, under suitable conditions, the scaling that ensures weak convergence of the estimator to a limiting normal random vector can range from cube-root of the sample size to square-root of the sample size, depending on the strength of the assumption made. The results are used to construct asymptotic confidence intervals for spectra and cross spectra. It is shown through a Monte-Carlo simulation study that these intervals have appropriate empirical coverage probabilities at moderate sample sizes.

math.ST

Consistent estimation of non-bandlimited spectral density from uniformly spaced samples

In the matter of selection of sample time points for the estimation of the power spectral density of a continuous time stationary stochastic process, irregular sampling schemes such as Poisson sampling are often preferred over regular (uniform) sampling. A major reason for this preference is the well-known problem of inconsistency of estimators based on regular sampling, when the underlying power spectral density is not bandlimited. It is argued in this paper that, in consideration of a large sample property like consistency, it is natural to allow the sampling rate to go to infinity as the sample size goes to infinity. Through appropriate asymptotic calculations under this scenario, it is shown that the smoothed periodogram based on regularly spaced data is a consistent estimator of the spectral density, even when the latter is not band-limited. It transpires that, under similar assumptions, the estimators based on uniformly sampled and Poisson-sampled data have about the same rate of convergence. Apart from providing this reassuring message, the paper also gives a guideline for practitioners regarding appropriate choice of the sampling rate. Theoretical calculations for large samples and Monte-Carlo simulations for small samples indicate that the smoothed periodogram based on uniformly sampled data have less variance and more bias than its counterpart based on Poisson sampled data.

math.ST

Effect of inter-sample spacing constraint on spectrum estimation with irregular sampling

A practical constraint that comes in the way of spectrum estimation of a continuous time stationary stochastic process is the minimum separation between successively observed samples of the process. When the underlying process is not band-limited, sampling at any uniform rate leads to aliasing, while certain stochastic sampling schemes, including Poisson process sampling, are rendered infeasible by the constraint of minimum separation. It is shown in this paper that, subject to this constraint, no point process sampling scheme is alias-free for the class of all spectra. It turns out that point process sampling under this constraint can be alias-free for band-limited spectra. However, the usual construction of a consistent spectrum estimator does not work in such a case. Simulations indicate that a commonly used estimator, which is consistent in the absence of this constraint, performs poorly when the constraint is present. These results should help practitioners in rationalizing their expectations from point process sampling as far as spectrum estimation is concerned, and motivate researchers to look for appropriate estimators of bandlimited spectra.

math.ST