SearcharxivSearch

arXiv subjects

Abhishek Kaul

Publications and source records attributed to Abhishek Kaul.

12 recordsLinked to original sources

High Dimensional Change Point Models for Two-Directional Data

We develop methodology for recovery of change points for data observed on more than one temporal index where changes may occur simultaneous in both indices, where the spatial component may be high dimensional. The work is motivated by climate monitoring problems where long series of data are available, e.g., daily observations (index 1) over several years (index 2). Such data may be evolving over the annual time scale, along with dynamic seasonal changes in the shorter time scale. We model this as a high dimensional mean process observed on a two dimensional grid with change points. Asymptotic estimation and inference results are developed under a single change point setup, including rates of convergence of the proposed method as well the resulting limiting distributions. The method is extended to the case of multiple changes. Theoretical results are supported numerically with monte-carlo simulations. We implement our work on a large scale climate data for the Pacific Northwest region of the United States.

stat.ME

Inference for Change Points in High Dimensional Mean Shift Models

We consider the problem of constructing confidence intervals for the locations of change points in a high-dimensional mean shift model. To that end, we develop a locally refitted least squares estimator and obtain component-wise and simultaneous rates of estimation of the underlying change points. The simultaneous rate is the sharpest available in the literature by at least a factor of $\log p,$ while the component-wise one is optimal. These results enable existence of limiting distributions. Component-wise distributions are characterized under both vanishing and non-vanishing jump size regimes, while joint distributions for any finite subset of change point estimates are characterized under the latter regime, which also yields asymptotic independence of these estimates. The combined results are used to construct asymptotically valid component-wise and simultaneous confidence intervals for the change point parameters. The results are established under a high dimensional scaling, allowing for diminishing jump sizes, in the presence of diverging number of change points and under subexponential errors. They are illustrated on synthetic data and on sensor measurements from smartphones for activity recognition.

stat.ME

Segmentation of high dimensional means over multi-dimensional change points and connections to regression trees

This article is motivated by the objective of providing a new analytically tractable and fully frequentist framework to characterize and implement regression trees while also allowing a multivariate (potentially high dimensional) response. The connection to regression trees is made by a high dimensional model with dynamic mean vectors over multi-dimensional change axes. Our theoretical analysis is carried out under a single two dimensional change point setting. An optimal rate of convergence of the proposed estimator is obtained, which in turn allows existence of limiting distributions. Distributional behavior of change point estimates are split into two distinct regimes, the limiting distributions under each regime is then characterized, in turn allowing construction of asymptotically valid confidence intervals for $2d$-location of change. All results are obtained under a high dimensional scaling $s\log^2 p=o(T_wT_h),$ where $p$ is the response dimension, $s$ is a sparsity parameter, and $T_w,T_h$ are sampling periods along change axes. We characterize full regression trees by defining a multiple multi-dimensional change point model. Natural extensions of the single $2d$-change point estimation methodology are provided. Two applications, first on segmentation of {\it Infra-red astronomy satellite (IRAS)} data and second to segmentation of digital images are provided. Methodology and theoretical results are supported with monte-carlo simulations.

stat.ME

Inference on the Change Point for High Dimensional Dynamic Graphical Models

We develop an estimator for the change point parameter for a dynamically evolving graphical model, and also obtain its asymptotic distribution under high dimensional scaling. To procure the latter result, we establish that the proposed estimator exhibits an $O_p(ψ^{-2})$ rate of convergence, wherein $ψ$ represents the jump size between the graphical model parameters before and after the change point. Further, it retains sufficient adaptivity against plug-in estimates of the graphical model parameters. We characterize the forms of the asymptotic distribution under the both a vanishing and a non-vanishing regime of the magnitude of the jump size. Specifically, in the former case it corresponds to the argmax of a negative drift asymmetric two sided Brownian motion, while in the latter case to the argmax of a negative drift asymmetric two sided random walk, whose increments depend on the distribution of the graphical model. Easy to implement algorithms are provided for estimating the change point and their performance assessed on synthetic data. The proposed methodology is further illustrated on RNA-sequenced microbiome data and their changes between young and older individuals.

math.ST

Inference on the change point in high dimensional time series models via plug in least squares

We study a plug in least squares estimator for the change point parameter where change is in the mean of a high dimensional random vector under subgaussian or subexponential distributions. We obtain sufficient conditions under which this estimator possesses sufficient adaptivity against plug in estimates of mean parameters in order to yield an optimal rate of convergence $O_p(ξ^{-2})$ in the integer scale. This rate is preserved while allowing high dimensionality as well as a potentially diminishing jump size $ξ,$ provided $s\log (p\vee T)=o(\surd(Tl_T))$ or $s\log^{3/2}(p\vee T)=o(\surd(Tl_T))$ in the subgaussian and subexponential cases, respectively. Here $s,p,T$ and $l_T$ represent a sparsity parameter, model dimension, sampling period and the separation of the change point from its parametric boundary. Moreover, since the rate of convergence is free of $s,p$ and logarithmic terms of $T,$ it allows the existence of limiting distributions. These distributions are then derived as the {\it argmax} of a two sided negative drift Brownian motion or a two sided negative drift random walk under vanishing and non-vanishing jump size regimes, respectively. Thereby allowing inference of the change point parameter in the high dimensional setting. Feasible algorithms for implementation of the proposed methodology are provided. Theoretical results are supported with monte-carlo simulations.

stat.ME

Inference on the change point with the jump size near the boundary of the region of detectability in high dimensional time series models

We develop a projected least squares estimator for the change point parameter in a high dimensional time series model with a potential change point. Importantly we work under the setup where the jump size may be near the boundary of the region of detectability. The proposed methodology yields an optimal rate of convergence despite high dimensionality of the assumed model and a potentially diminishing jump size. The limiting distribution of this estimate is derived, thereby allowing construction of a confidence interval for the location of the change point. A secondary near optimal estimate is proposed which is required for the implementation of the optimal projected least squares estimate. The prestep estimation procedure is designed to also agnostically detect the case where no change point exists, thereby removing the need to pretest for the existence of a change point for the implementation of the inference methodology. Our results are presented under a general positive definite spatial dependence setup, assuming no special structure on this dependence. The proposed methodology is designed to be highly scalable, and applicable to very large data. Theoretical results regarding detection and estimation consistency and the limiting distribution are numerically supported via monte carlo simulations.

math.ST

Pivotal Estimation via Self-Normalization for High-Dimensional Linear Models with Error in Variables

We propose a new estimator for the high-dimensional linear regression model with observation error in the design where the number of coefficients is potentially larger than the sample size. The main novelty of our procedure is that the choice of penalty parameters is pivotal. The estimator is based on applying a self-normalization to the constraints that characterize the estimator. Importantly, we show how to cast the computation of the estimator as the solution of a convex program with second order cone constraints. This allows the use of algorithms with theoretical guarantees and reliable implementation. Under sparsity assumptions, we derive $\ell_q$-rates of convergence and show that consistency can be achieved even if the number of regressors exceeds the sample size. We further provide a simple to implement rule to threshold the estimator that yields a provably sparse estimator with similar $\ell_2$ and $\ell_1$-rates of convergence. The thresholds are data-driven and component dependents. Finally, we also study the rates of convergence of estimators that refit the data based on a selected support with possible model selection mistakes. In addition to our finite sample theoretical results that allow for non-i.i.d. data, we also present simulations to compare the performance of the proposed estimators.

stat.ME

Detection and estimation of parameters in high dimensional multiple change point regression models via $\ell_1/\ell_0$ regularization and discrete optimization

Binary segmentation, which is sequential in nature is thus far the most widely used method for identifying multiple change points in statistical models. Here we propose a top down methodology called arbitrary segmentation that proceeds in a conceptually reverse manner. We begin with an arbitrary superset of the parametric space of the change points, and locate unknown change points by suitably filtering this space down. Critically, we reframe the problem as that of variable selection in the change point parameters, this enables the filtering down process to be achieved in a single step with the aid of an $\ell_0$ regularization, thus avoiding the sequentiality of binary segmentation. We study this method under a high dimensional multiple change point linear regression model and show that rates convergence of the error in the regression and change point estimates are near optimal. We propose a simulated annealing (SA) approach to implement a key finite state space discrete optimization that arises in our method. Theoretical results are numerically supported via simulations. The proposed method is shown to possess the ability to agnostically detect the `no change' scenario. Furthermore, its computational complexity is of order $O(Np^2)$+SA, where SA is the cost of a SA optimization on a $N$(no. of change points) dimensional grid. Thus, the proposed methodology is significantly more computationally efficient than existing approaches. Finally, our theoretical results are obtained under weaker model conditions than those assumed in the current literature.

math.ST

An efficient two step algorithm for high dimensional change point regression models without grid search

We propose a two step algorithm based on $\ell_1/\ell_0$ regularization for the detection and estimation of parameters of a high dimensional change point regression model and provide the corresponding rates of convergence for the change point as well as the regression parameter estimates. Importantly, the computational cost of our estimator is only $2\cdotp$Lasso$(n,p)$, where Lasso$(n,p)$ represents the computational burden of one Lasso optimization in a model of size $(n,p)$. In comparison, existing grid search based approaches to this problem require a computational cost of at least $n\cdot {\rm Lasso}(n,p)$ optimizations. Additionally, the proposed method is shown to be able to consistently detect the case of `no change', i.e., where no finite change point exists in the model. We work under a subgaussian random design where the underlying assumptions in our study are milder than those currently assumed in the high dimensional change point regression literature. We allow the true change point parameter $τ_0$ to possibly move to the boundaries of its parametric space, and the jump size $\|β_0-γ_0\|_2$ to possibly diverge as $n$ increases. We then characterize the corresponding effects on the rates of convergence of the change point and regression estimates. In particular, we show that, while an increasing jump size may have a beneficial effect on the change point estimate, however the optimal rate of regression parameter estimates are preserved only upto a certain rate of the increasing jump size. This behavior in the rate of regression parameter estimates is unique to high dimensional change point regression models only. Simulations are performed to empirically evaluate performance of the proposed estimators. The methodology is applied to community level socio-economic data of the U.S., collected from the 1990 U.S. census and other sources.

stat.ME

Confidence Bands for Coefficients in High Dimensional Linear Models with Error-in-variables

We study high-dimensional linear models with error-in-variables. Such models are motivated by various applications in econometrics, finance and genetics. These models are challenging because of the need to account for measurement errors to avoid non-vanishing biases in addition to handle the high dimensionality of the parameters. A recent growing literature has proposed various estimators that achieve good rates of convergence. Our main contribution complements this literature with the construction of simultaneous confidence regions for the parameters of interest in such high-dimensional linear models with error-in-variables. These confidence regions are based on the construction of moment conditions that have an additional orthogonal property with respect to nuisance parameters. We provide a construction that requires us to estimate an additional high-dimensional linear model with error-in-variables for each component of interest. We use a multiplier bootstrap to compute critical values for simultaneous confidence intervals for a subset $S$ of the components. We show its validity despite of possible model selection mistakes, and allowing for the cardinality of $S$ to be larger than the sample size. We apply and discuss the implications of our results to two examples and conduct Monte Carlo simulations to illustrate the performance of the proposed procedure.

math.ST

Analysis of High Dimensional Compositional Data Containing Structural Zeros with Applications to Microbiome Data

This paper is motivated by the recent interest in the analysis of high dimen- sional microbiome data. A key feature of this data is the presence of `structural zeros' which are microbes missing from an observation vector due to an underlying biological process and not due to error in measurement. Typical notions of missingness are insufficient to model these structural zeros. We define a general framework which allows for structural zeros in the model and propose methods of estimating sparse high dimensional covariance and precision matrices under this setup. We establish error bounds in the spectral and frobenius norms for the proposed esti- mators and empirically support them with a simulation study. We also apply the proposed methodology to the global human gut microbiome data of Yatsunenko (2012).

stat.AP

Two Stage Non-penalized Corrected Least Squares for High Dimensional Linear Models with Measurement error or Missing Covariates

This paper provides an alternative to penalized estimators for estimation and vari- able selection in high dimensional linear regression models with measurement error or missing covariates. We propose estimation via bias corrected least squares after model selection. We show that by separating model selection and estimation, it is possible to achieve an improved rate of convergence of the L2 estimation error compared to the rate sqrt{s log p/n} achieved by simultaneous estimation and variable selection methods such as L1 penalized corrected least squares. If the correct model is selected with high probability then the L2 rate of convergence for the proposed method is indeed the oracle rate of sqrt{s/n}. Here s, p are the number of non zero parameters and the model dimension, respectively, and n is the sample size. Under very general model selection criteria, the proposed method is computationally simpler and statistically at least as efficient as the L1 penalized corrected least squares method, performs model selection without the availability of the bias correction matrix, and is able to provide estimates with only a small sub-block of the bias correction covariance matrix of order s x s in comparison to the p x p correction matrix required for computation of the L1 penalized version. Furthermore we show that the model selection requirements are met by a correlation screening type method and the L1 penalized corrected least squares method. Also, the proposed methodology when applied to the estimation of precision matrices with missing observations, is seen to perform at least as well as existing L1 penalty based methods. All results are supported empirically by a simulation study.

stat.ME