Searcharxiv⌕ Search

arXiv subjects

Debashis Paul

Publications and source records attributed to Debashis Paul.

36 records · Page 2Linked to original sources

Spectral analysis of linear time series in moderately high dimensions

This article is concerned with the spectral behavior of $p$-dimensional linear processes in the moderately high-dimensional case when both dimensionality $p$ and sample size $n$ tend to infinity so that $p/n\to0$. It is shown that, under an appropriate set of assumptions, the empirical spectral distributions of the renormalized and symmetrized sample autocovariance matrices converge almost surely to a nonrandom limit distribution supported on the real line. The key assumption is that the linear process is driven by a sequence of $p$-dimensional real or complex random vectors with i.i.d. entries possessing zero mean, unit variance and finite fourth moments, and that the $p\times p$ linear process coefficient matrices are Hermitian and simultaneously diagonalizable. Several relaxations of these assumptions are discussed. The results put forth in this paper can help facilitate inference on model parameters, model diagnostics and prediction of future values of the linear process.

math.ST↗

On the Marčenko-Pastur law for linear time series

This paper is concerned with extensions of the classical Marčenko-Pastur law to time series. Specifically, $p$-dimensional linear processes are considered which are built from innovation vectors with independent, identically distributed (real- or complex-valued) entries possessing zero mean, unit variance and finite fourth moments. The coefficient matrices of the linear process are assumed to be simultaneously diagonalizable. In this setting, the limiting behavior of the empirical spectral distribution of both sample covariance and symmetrized sample autocovariance matrices is determined in the high-dimensional setting $p/n\to c\in (0,\infty)$ for which dimension $p$ and sample size $n$ diverge to infinity at the same rate. The results extend existing contributions available in the literature for the covariance case and are one of the first of their kind for the autocovariance case.

math.ST↗

Time-warped growth processes, with applications to the modeling of boom-bust cycles in house prices

House price increases have been steady over much of the last 40 years, but there have been occasional declines, most notably in the recent housing bust that started around 2007, on the heels of the preceding housing bubble. We introduce a novel growth model that is motivated by time-warping models in functional data analysis and includes a nonmonotone time-warping component that allows the inclusion and description of boom-bust cycles and facilitates insights into the dynamics of asset bubbles. The underlying idea is to model longitudinal growth trajectories for house prices and other phenomena, where temporal setbacks and deflation may be encountered, by decomposing such trajectories into two components. A first component corresponds to underlying steady growth driven by inflation that anchors the observed trajectories on a simple first order linear differential equation, while a second boom-bust component is implemented as time warping. Time warping is a commonly encountered phenomenon and reflects random variation along the time axis. Our approach to time warping is more general than previous approaches by admitting the inclusion of nonmonotone warping functions. The anchoring of the trajectories on an underlying linear dynamic system also makes the time-warping component identifiable and enables straightforward estimation procedures for all model components. The application to the dynamics of housing prices as observed for 19 metropolitan areas in the U.S. from December 1998 to July 2013 reveals that the time setbacks corresponding to nonmonotone time warping vary substantially across markets and we find indications that they are related to market-specific growth rates.

stat.AP↗

Adaptation in a class of linear inverse problems

We consider the linear inverse problem of estimating an unknown signal $f$ from noisy measurements on $Kf$ where the linear operator $K$ admits a wavelet-vaguelette decomposition (WVD). We formulate the problem in the Gaussian sequence model and propose estimation based on complexity penalized regression on a level-by-level basis. We adopt squared error loss and show that the estimator achieves exact rate-adaptive optimality as $f$ varies over a wide range of Besov function classes.

math.ST↗

Nonparametric estimation of dynamics of monotone trajectories

We study a class of nonlinear nonparametric inverse problems. Specifically, we propose a nonparametric estimator of the dynamics of a monotonically increasing trajectory defined on a finite time interval. Under suitable regularity conditions, we prove consistency of the proposed estimator and show that in terms of $L^2$-loss, the optimal rate of convergence for the proposed estimator is the same as that for the estimation of the derivative of a trajectory. This is a new contribution to the area of nonlinear nonparametric inverse problems. We conduct a simulation study to examine the finite sample behavior of the proposed estimator and apply it to the Berkeley growth data.

math.ST↗

High-dimensional genome-wide association study and misspecified mixed model analysis

We study behavior of the restricted maximum likelihood (REML) estimator under a misspecified linear mixed model (LMM) that has received much attention in recent gnome-wide association studies. The asymptotic analysis establishes consistency of the REML estimator of the variance of the errors in the LMM, and convergence in probability of the REML estimator of the variance of the random effects in the LMM to a certain limit, which is equal to the true variance of the random effects multiplied by the limiting proportion of the nonzero random effects present in the LMM. The aymptotic results also establish convergence rate (in probability) of the REML estimators as well as a result regarding convergence of the asymptotic conditional variance of the REML estimator. The asymptotic results are fully supported by the results of empirical studies, which include extensive simulation studies that compare the performance of the REML estimator (under the misspecified LMM) with other existing methods.

math.ST↗

Limiting spectral distribution of renormalized separable sample covariance matrices when $p/n\to 0$

We are concerned with the behavior of the eigenvalues of renormalized sample covariance matrices of the form C_n=\sqrt{\frac{n}{p}}\left(\frac{1}{n}A_{p}^{1/2}X_{n}B_{n}X_{n}^{*}A_{p}^{1/2}-\frac{1}{n}\tr(B_{n})A_{p}\right) as $p,n\to \infty$ and $p/n\to 0$, where $X_{n}$ is a $p\times n$ matrix with i.i.d. real or complex valued entries $X_{ij}$ satisfying $E(X_{ij})=0$, $E|X_{ij}|^2=1$ and having finite fourth moment. $A_{p}^{1/2}$ is a square-root of the nonnegative definite Hermitian matrix $A_{p}$, and $B_{n}$ is an $n\times n$ nonnegative definite Hermitian matrix. We show that the empirical spectral distribution (ESD) of $C_n$ converges a.s. to a nonrandom limiting distribution under some assumptions. The probability density function of the LSD of $C_{n}$ is derived and it is shown that it depends on the LSD of $A_{p}$ and the limiting value of $n^{-1}\tr(B_{n}^2)$. We propose a computational algorithm for evaluating this limiting density when the LSD of $A_{p}$ is a mixture of point masses. In addition, when the entries of $X_{n}$ are sub-Gaussian, we derive the limiting empirical distribution of $\{\sqrt{n/p}(λ_j(S_n) - n^{-1}\tr(B_n) λ_j(A_{p}))\}_{j=1}^p$ where $S_n := n^{-1} A_{p}^{1/2}X_{n}B_{n}X_{n}^{*}A_{p}^{1/2}$ is the sample covariance matrix and $λ_j$ denotes the $j$-th largest eigenvalue, when $F^A$ is a finite mixture of point masses. These results are utilized to propose a test for the covariance structure of the data where the null hypothesis is that the joint covariance matrix is of the form $A_{p} \otimes B_n$ for $\otimes$ denoting the Kronecker product, as well as $A_{p}$ and the first two spectral moments of $B_n$ are specified. The performance of this test is illustrated through a simulation study.

math.ST↗

Minimax bounds for sparse PCA with noisy high-dimensional data

We study the problem of estimating the leading eigenvectors of a high-dimensional population covariance matrix based on independent Gaussian observations. We establish a lower bound on the minimax risk of estimators under the $l_2$ loss, in the joint limit as dimension and sample size increase to infinity, under various models of sparsity for the population eigenvectors. The lower bound on the risk points to the existence of different regimes of sparsity of the eigenvectors. We also propose a new method for estimating the eigenvectors by a two-stage coordinate selection scheme.

math.ST↗

Augmented sparse principal component analysis for high dimensional data

We study the problem of estimating the leading eigenvectors of a high-dimensional population covariance matrix based on independent Gaussian observations. We establish lower bounds on the rates of convergence of the estimators of the leading eigenvectors under $l^q$-sparsity constraints when an $l^2$ loss function is used. We also propose an estimator of the leading eigenvectors based on a coordinate selection scheme combined with PCA and show that the proposed estimator achieves the optimal rate of convergence under a sparsity regime. Moreover, we establish that under certain scenarios, the usual PCA achieves the minimax convergence rate.

math.ST↗

Semiparametric modeling of autonomous nonlinear dynamical systems with application to plant growth

We propose a semiparametric model for autonomous nonlinear dynamical systems and devise an estimation procedure for model fitting. This model incorporates subject-specific effects and can be viewed as a nonlinear semiparametric mixed effects model. We also propose a computationally efficient model selection procedure. We show by simulation studies that the proposed estimation as well as model selection procedures can efficiently handle sparse and noisy measurements. Finally, we apply the proposed method to a plant growth data used to study growth displacement rates within meristems of maize roots under two different experimental conditions.

stat.AP↗

Shrinking the Quadratic Estimator

We study a regression characterization for the quadratic estimator of weak lensing, developed by Hu and Okamoto (2001,2002), for cosmic microwave background observations. This characterization motivates a modification of the quadratic estimator by an adaptive Wiener filter which uses the robust Bayesian techniques described in Strawderman (1971) and Berger (1980). This technique requires the user to propose a fiducial model for the spectral density of the unknown lensing potential but the resulting estimator is developed to be robust to misspecification of this model. The role of the fiducial spectral density is to give the estimator superior statistical performance in a "neighborhood of the fiducial model" while controlling the statistical errors when the fiducial spectral density is drastically wrong. Our estimate also highlights some advantages provided by a Bayesian analysis of the quadratic estimator.

astro-ph.IM↗

Geometric kernel smoothing of tensor fields

In this paper, we study a kernel smoothing approach for denoising a tensor field. Particularly, both simulation studies and theoretical analysis are conducted to understand the effects of the noise structure and the structure of the tensor field on the performance of different smoothers arising from using different metrics, viz., Euclidean, log-Euclidean and affine invariant metrics. We also study the Rician noise model and compare two regression estimators of diffusion tensors based on raw diffusion weighted imaging data at each voxel.

stat.ME↗

Semiparametric modeling of autonomous nonlinear dynamical systems with applications

In this paper, we propose a semi-parametric model for autonomous nonlinear dynamical systems and devise an estimation procedure for model fitting. This model incorporates subject-specific effects and can be viewed as a nonlinear semi-parametric mixed effects model. We also propose a computationally efficient model selection procedure. We prove consistency of the proposed estimator under suitable regularity conditions. We show by simulation studies that the proposed estimation as well as model selection procedures can efficiently handle sparse and noisy measurements. Finally, we apply the proposed method to a plant growth data used to study growth displacement rates within meristems of maize roots under two different experimental conditions.

stat.ME↗

Tie-respecting bootstrap methods for estimating distributions of sets and functions of eigenvalues

Bootstrap methods are widely used for distribution estimation, although in some problems they are applicable only with difficulty. A case in point is that of estimating the distributions of eigenvalue estimators, or of functions of those estimators, when one or more of the true eigenvalues are tied. The $m$-out-of-$n$ bootstrap can be used to deal with problems of this general type, but it is very sensitive to the choice of $m$. In this paper we propose a new approach, where a tie diagnostic is used to determine the locations of ties, and parameter estimates are adjusted accordingly. Our tie diagnostic is governed by a probability level, $β$, which in principle is an analogue of $m$ in the $m$-out-of-$n$ bootstrap. However, the tie-respecting bootstrap (TRB) is remarkably robust against the choice of $β$. This makes the TRB significantly more attractive than the $m$-out-of-$n$ bootstrap, where the value of $m$ has substantial influence on the final result. The TRB can be used very generally; for example, to test hypotheses about, or construct confidence regions for, the proportion of variability explained by a set of principal components. It is suitable for both finite-dimensional data and functional data.

math.ST↗

Principal components analysis for sparsely observed correlated functional data using a kernel smoothing approach

In this paper, we consider the problem of estimating the covariance kernel and its eigenvalues and eigenfunctions from sparse, irregularly observed, noise corrupted and (possibly) correlated functional data. We present a method based on pre-smoothing of individual sample curves through an appropriate kernel. We show that the naive empirical covariance of the pre-smoothed sample curves gives highly biased estimator of the covariance kernel along its diagonal. We attend to this problem by estimating the diagonal and off-diagonal parts of the covariance kernel separately. We then present a practical and efficient method for choosing the bandwidth for the kernel by using an approximation to the leave-one-curve-out cross validation score. We prove that under standard regularity conditions on the covariance kernel and assuming i.i.d. samples, the risk of our estimator, under $L^2$ loss, achieves the optimal nonparametric rate when the number of measurements per curve is bounded. We also show that even when the sample curves are correlated in such a way that the noiseless data has a separable covariance structure, the proposed method is still consistent and we quantify the role of this correlation in the risk of the estimator.

stat.ME↗

Consistency of restricted maximum likelihood estimators of principal components

In this paper we consider two closely related problems : estimation of eigenvalues and eigenfunctions of the covariance kernel of functional data based on (possibly) irregular measurements, and the problem of estimating the eigenvalues and eigenvectors of the covariance matrix for high-dimensional Gaussian vectors. In Peng and Paul (2007), a restricted maximum likelihood (REML) approach has been developed to deal with the first problem. In this paper, we establish consistency and derive rate of convergence of the REML estimator for the functional data case, under appropriate smoothness conditions. Moreover, we prove that when the number of measurements per sample curve is bounded, under squared-error loss, the rate of convergence of the REML estimators of eigenfunctions is near-optimal. In the case of Gaussian vectors, asymptotic consistency and an efficient score representation of the estimators are obtained under the assumption that the effective dimension grows at a rate slower than the sample size. These results are derived through an explicit utilization of the intrinsic geometry of the parameter space, which is non-Euclidean. Moreover, the results derived in this paper suggest an asymptotic equivalence between the inference on functional data with dense measurements and that of the high dimensional Gaussian vectors.

math.ST↗

A geometric approach to maximum likelihood estimation of the functional principal components from sparse longitudinal data

In this paper, we consider the problem of estimating the eigenvalues and eigenfunctions of the covariance kernel (i.e., the functional principal components) from sparse and irregularly observed longitudinal data. We approach this problem through a maximum likelihood method assuming that the covariance kernel is smooth and finite dimensional. We exploit the smoothness of the eigenfunctions to reduce dimensionality by restricting them to a lower dimensional space of smooth functions. The estimation scheme is developed based on a Newton-Raphson procedure using the fact that the basis coefficients representing the eigenfunctions lie on a Stiefel manifold. We also address the selection of the right number of basis functions, as well as that of the dimension of the covariance kernel by a second order approximation to the leave-one-curve-out cross-validation score that is computationally very efficient. The effectiveness of our procedure is demonstrated by simulation studies and an application to a CD4 counts data set. In the simulation studies, our method performs well on both estimation and model selection. It also outperforms two existing approaches: one based on a local polynomial smoothing of the empirical covariances, and another using an EM algorithm.

stat.ME↗

"Pre-conditioning" for feature selection and regression in high-dimensional problems

We consider regression problems where the number of predictors greatly exceeds the number of observations. We propose a method for variable selection that first estimates the regression function, yielding a "pre-conditioned" response variable. The primary method used for this initial regression is supervised principal components. Then we apply a standard procedure such as forward stepwise selection or the LASSO to the pre-conditioned response variable. In a number of simulated and real data examples, this two-step procedure outperforms forward stepwise selection or the usual LASSO (applied directly to the raw outcome). We also show that under a certain Gaussian latent variable model, application of the LASSO to the pre-conditioned response variable is consistent as the number of predictors and observations increases. Moreover, when the observational noise is rather large, the suggested procedure can give a more accurate estimate than LASSO. We illustrate our method on some real problems, including survival analysis with microarray data.

math.ST↗