SearcharxivSearch

arXiv subjects

Anindya Roy

Publications and source records attributed to Anindya Roy.

At least 19 recordsLinked to original sources

A Maximum Entropy Implementation of Differential Privacy Under Linear Invariants

Differential privacy is the standard for ensuring data privacy and is widely used in major data publications, including reporting results from the U.S. decennial census. Common implementation of differential privacy uses independent Gaussian or Laplace noise addition to the database. However, there could be aggregate (linear) queries to the database that are excluded from the privacy budget, for example, state totals that can not be perturbed due to constitutional mandates. Any implementation of a differential privacy is required to honor these constraints, also referred to as invariants. Under aggregation constraints, the noise vector is no longer independent and the traditional differential privacy guarantees have to be re-evaluated. We propose a high entropy differential privacy implementation that maintains the aggregation invariants with probability one or exponentially close to one and derive the privacy guarantees for the implementation under the invariants. The theoretical proof covers a partial solution to an open question about the null space of correlation matrices. Moreover, the methodology has general use in the context of sampling from normal mixture models under linear equality constraints.

cs.CR

Bayesian Graphical High-Dimensional Time Series Models for Detecting Structural Changes

We study the structural changes in multivariate time-series by estimating and comparing stationary graphs for macroeconomic time series before and after an economic crisis such as the Great Recession. Building on a latent time series framework called Orthogonally-rotated Univariate Time-series (OUT), we propose a shared-parameter framework-the spOUT autoregressive model (spOUTAR)-that jointly models two related multivariate time series and enables coherent Bayesian estimation of their corresponding stationary precision matrices. This framework provides a principled mechanism to detect and quantify which conditional relationships among the variables changed, or formed following the crisis. Specifically, we study the impact of the Great Recession (December 2007-June 2009) that substantially disrupted global and national economies, prompting long-lasting shifts in macroeconomic indicators and their interrelationships. While many studies document its economic consequences, far less is known about how the underlying conditional dependency structure among economic variables changed as economies moved from pre-crisis stability through the shock and back to normalcy. Using the proposed approach to analyze U.S. and OECD macroeconomic data, we demonstrate that spOUTAR effectively captures recession-induced changes in stationary graphical structure, offering a flexible and interpretable tool for studying structural shifts in economic systems.

stat.ME

Bayesian Inference for High-dimensional Time Series with a Stationary Directed Acyclic Graphical Structure

In multivariate time series analysis, understanding the underlying causal relationships among variables is often of interest for various applications. Directed acyclic graphs (DAGs) provide a powerful framework for representing causal dependencies. This paper proposes a novel Bayesian approach for modeling multivariate time series where conditional independencies and causal structure are encoded by a DAG. The proposed model allows structural properties such as stationarity to be easily accommodated. Given the application, we further extend the model for matrix-variate time series. We take a Bayesian approach to inference, and a ``immersion-posterior'' based efficient computational algorithm is developed. The posterior convergence properties of the proposed method are established along with two identifiability results for the unrestricted structural equation models. The utility of the proposed method is demonstrated through simulation studies and real data analysis.

stat.ME

Achieving Privacy Utility Balance for Multivariate Time Series Data

Utility-preserving data privatization is of utmost importance for data-producing agencies. The popular noise-addition privacy mechanism distorts autocorrelation patterns in time series data, thereby marring utility; in response, McElroy et al. (2023) introduced all-pass filtering (FLIP) as a utility-preserving time series data privatization method. Adapting this concept to multivariate data is more complex, and in this paper we propose a multivariate all-pass (MAP) filtering method, employing an optimization algorithm to achieve the best balance between data utility and privacy protection. To test the effectiveness of our approach, we apply MAP filtering to both simulated and real data, sourced from the U.S. Census Bureau's Quarterly Workforce Indicator (QWI) dataset.

stat.ME

Relational Graph in Vector Autoregression: A Case Study on the Effect of the Great Recession on Connectivity of Economic Indicators

Under a high-dimensional vector autoregressive (VAR) model, we propose a way of efficiently estimating both the stationary graph structure between the nodal time series and their temporal dynamics. The framework is then used to make inferences on the change in interdependencies between several economic indicators due to the impact of the Great Recession, the financial crisis that lasted from 2007 through 2009. There are several key advantages of the proposed framework; (1) it develops a reparametrized VAR likelihood that can be used in general high-dimensional VAR problems, (2) it strictly maintains causality of the estimated process, making inference on stationary features more meaningful and (3) it is computationally efficient due to the reduced rank structure of the parameterization. We apply the methodology to the seasonally adjusted quarterly economic indicators available in the FRED-QD database of the Federal Reserve. The analysis essentially confirms much of the prevailing knowledge about the impact of the Great Recession on different economic indicators. At the same time, it provides deeper insight into the nature and extent of the impact on the interplay of the different indicators. We also contribute to the theory of Bayesian VAR by showing the consistency of the posterior under sparse priors for the parameters of the reduced rank formulation of the VAR process.

stat.ME

Bayesian Learning of Relational Graph in Semiparametric High-dimensional Time Series

Time series data arising in many applications nowadays are high-dimensional. A large number of parameters describe features of these time series. We propose a novel approach to modeling a high-dimensional time series through several independent univariate time series, which are then orthogonally rotated and sparsely linearly transformed. With this approach, any specified intrinsic relations among component time series given by a graphical structure can be maintained at all time snapshots. We call the resulting process an Orthogonally-rotated Univariate Time series (OUT). Key structural properties of time series such as stationarity and causality can be easily accommodated in the OUT model. For Bayesian inference, we put suitable prior distributions on the spectral densities of the independent latent times series, the orthogonal rotation matrix, and the common precision matrix of the component times series at every time point. A likelihood is constructed using the Whittle approximation for univariate latent time series. An efficient Markov Chain Monte Carlo (MCMC) algorithm is developed for posterior computation. We study the convergence of the pseudo-posterior distribution based on the Whittle likelihood for the model's parameters upon developing a new general posterior convergence theorem for pseudo-posteriors. We find that the posterior contraction rate for independent observations essentially prevails in the OUT model under very mild conditions on the temporal dependence described in terms of the smoothness of the corresponding spectral densities. Through a simulation study, we compare the accuracy of estimating the parameters and identifying the graphical structure with other approaches. We apply the proposed methodology to analyze a dataset on different industrial components of the US gross domestic product between 2010 and 2019 and predict future observations.

stat.ME

Conic Sparsity: Estimation of Regression Parameters in Closed Convex Polyhedral Cones

Statistical problems often involve linear equality and inequality constraints on model parameters. Direct estimation of parameters restricted to general polyhedral cones, particularly when one is interested in estimating low dimensional features, may be challenging. We use a dual form parameterization to characterize parameter vectors restricted to lower dimensional faces of polyhedral cones and use the characterization to define a notion of 'sparsity' on such cones. We show that the proposed notion agrees with the usual notion of sparsity in the unrestricted case and prove the validity of the proposed definition as a measure of sparsity. The identifiable parameterization of the lower dimensional faces allows a generalization of popular spike-and-slab priors to a closed convex polyhedral cone. The prior measure utilizes the geometry of the cone by defining a Markov random field over the adjacency graph of the extreme rays of the cone. We describe an efficient way of computing the posterior of the parameters in the restricted case. We illustrate the usefulness of the proposed methodology for imposing linear equality and inequality constraints by using wearables data from the National Health and Nutrition Examination Survey (NHANES) actigraph study where the daily average activity profiles of participants exhibit patterns that seem to obey such constraints.

stat.ME

FLIP: A Utility Preserving Privacy Mechanism for Time Series

Guaranteeing privacy in released data is an important goal for data-producing agencies. There has been extensive research on developing suitable privacy mechanisms in recent years. Particularly notable is the idea of noise addition with the guarantee of differential privacy. There are, however, concerns about compromising data utility when very stringent privacy mechanisms are applied. Such compromises can be quite stark in correlated data, such as time series data. Adding white noise to a stochastic process may significantly change the correlation structure, a facet of the process that is essential to optimal prediction. We propose the use of all-pass filtering as a privacy mechanism for regularly sampled time series data, showing that this procedure preserves utility while also providing sufficient privacy guarantees to entity-level time series.

cs.CR

Efficient Integration of Aggregate Data and Individual Patient Data in One-Way Mixed Models

Often both Aggregate Data (AD) studies and Individual Patient Data (IPD) studies are available for specific treatments. Combining these two sources of data could improve the overall meta-analytic estimates of treatment effects. Moreover, often for some studies with AD, the associated IPD maybe available, albeit at some extra effort or cost to the analyst. We propose a method for combining treatment effects across trials when the response is from the exponential family of distribution and hence a generalized linear model structure can be used. We consider the case when treatment effects are fixed and common across studies. Using the proposed combination method, we evaluate the wisdom of choosing AD when IPD is available by studying the relative efficiency of analyzing all IPD studies versus combining various percentages of AD and IPD studies. For many different models design constraints under which the AD estimators are the IPD estimators, and hence fully efficient, are known. For such models we advocate a selection procedure that chooses AD studies over IPD studies in a manner that force least departure from design constraints and hence ensures a fully efficient combined AD and IPD estimator.

stat.ME

Constrained Parameterization of Reduced Rank and Co-integrated Vector Autoregression

The paper provides a parametrization of Vector Autoregression (VAR) that enables one to look at the parameters associated with unit root dynamics and those associated with stable dynamics separately. The task is achieved via a novel factorization of the VAR polynomial that partitions the polynomial spectrum into unit root and stable and zero roots via polynomial factors. The proposed factorization adds to the literature of spectral factorization of matrix polynomials. The main benefit is that using the parameterization, actions could be taken to model the dynamics due to a particular class of roots, e.g. unit roots or zero roots, without changing the properties of the dynamics due to other roots. For example, using the parameterization one is able to estimate cointegrating space with appropriate rank that maintains the root structure of the original VAR processes or one can estimate a reduced rank causal VAR process maintaining the constraints of causality. In essence, this parameterization provides the practitioner an option to perform estimation of VAR processes with constrained root structure (e.g., conintegrated VAR or reduced rank VAR) such that the estimated model maintains the assumed root structure.

stat.ME

Stochastic Approximation Algorithm for Estimating Mixing Distribution for Dependent Observations

Estimating the mixing density of a mixture distribution remains an interesting problem in statistics literature. Using a stochastic approximation method, Newton and Zhang (1999) introduced a fast recursive algorithm for estimating the mixing density of a mixture. Under suitably chosen weights the stochastic approximation estimator converges to the true solution. In Tokdar et. al. (2009) the consistency of this recursive estimation method was established. However, the proof of consistency of the resulting estimator used independence among observations as an assumption. Here, we extend the investigation of performance of Newton's algorithm to several dependent scenarios. We prove that the original algorithm under certain conditions remains consistent even when the observations are arising from a weakly dependent stationary process with the target mixture as the marginal density. We show consistency under a decay condition on the dependence among observations when the dependence is characterized by a quantity similar to mutual information between the observations.

math.ST

Note on Mean Vector Testing for High-Dimensional Dependent Observations

For the mean vector test in high dimension, Ayyala et al.(2017,153:136-155) proposed new test statistics when the observational vectors are M dependent. Under certain conditions, the test statistics for one-same and two-sample cases were shown to be asymptotically normal. While the test statistics and the asymptotic results are valid, some parts of the proof of asymptotic normality need to be corrected. In this work, we provide corrections to the proofs of their main theorems. We also note a few minor discrepancies in calculations in the publication.

math.ST

Fixed support positive-definite modification of covariance matrix estimators via linear shrinkage

In this work, we study the positive definiteness (PDness) problem in covariance matrix estimation. For high dimensional data, many regularized estimators are proposed under structural assumptions on the true covariance matrix including sparsity. They are shown to be asymptotically consistent and rate-optimal in estimating the true covariance matrix and its structure. However, many of them do not take into account the PDness of the estimator and produce a non-PD estimate. To achieve the PDness, researchers consider additional regularizations (or constraints) on eigenvalues, which make both the asymptotic analysis and computation much harder. In this paper, we propose a simple modification of the regularized covariance matrix estimator to make it PD while preserving the support. We revisit the idea of linear shrinkage and propose to take a convex combination between the first-stage estimator (the regularized covariance matrix without PDness) and a given form of diagonal matrix. The proposed modification, which we denote as FSPD (Fixed Support and Positive Definiteness) estimator, is shown to preserve the asymptotic properties of the first-stage estimator, if the shrinkage parameters are carefully selected. It has a closed form expression and its computation is optimization-free, unlike existing PD sparse estimators. In addition, the FSPD is generic in the sense that it can be applied to any non-PD matrix including the precision matrix. The FSPD estimator is numerically compared with other sparse PD estimators to understand its finite sample properties as well as its computational gain. It is also applied to two multivariate procedures relying on the covariance matrix estimator -- the linear minimax classification problem and the Markowitz portfolio optimization problem -- and is shown to substantially improve the performance of both procedures.

stat.ME

Tetrahedral Bonding in Twisted Bilayer Graphene by Carbon Intercalation

Based on ab initio calculations, we study the effect of intercalating twisted bilayer graphene with carbon. Surprisingly, we find that the intercalant pulls the atoms in the two layers closer together locally when placed in certain regions in between the layers, and the process is energetically favorable as well. This arises because in these regions of the supercell, the local environment allows the intercalant to form tetrahedral bonding with nearest atoms in the layers. Intercalating AB- or AA-bilayer graphene with carbon does not produce this effect; therefore, the nontrivial effect owes its origin to both using carbon as an intercalant and using twisted bilayer graphene as the host. This opens new routes to manipulating bilayer and multilayer van der Waals heterostructures and tuning their properties in an unconventional way.

cond-mat.mtrl-sci

Multi-physics simulations of lithiation-induced stress in Li$_{\rm 1+x}$Ti$_2$O$_4$ electrode particles

Cubic spinel Li$_{\rm 1+x}$Ti$_2$O$_4$ is a promising electrode material as it exhibits a high lithium diffusivity and undergoes minimal changes in lattice parameters during lithiation and delithiation, thereby ensuring favorable cycleability. The present work is a multi-physics and multi-scale study of Li$_{\rm 1+x}$Ti$_2$O$_4$ that combines first principles computations of thermodynamic and kinetic properties with continuum scale modeling of lithiation-delithiation kinetics. Density functional theory calculations and statistical mechanics methods are used to calculate lattice parameters, elastic coefficients, thermodynamic potentials, migration barriers and Li diffusion coefficients. These quantities then inform a phase field framework to model the coupled chemo-mechanical evolution of electrode particles. Several case studies accounting for either homogeneous or heterogeneous nucleation are considered to explore the temporal evolution of maximum principle stress values, which serve to indicate stress localization and the potential for crack initiation, during lithiation and delithiation.

cond-mat.mtrl-sci

Estimates of the thermal conductivity and the thermoelectric properties of PbTiO$_3$ from first principles

The lattice thermal conductivity ($κ_{\rm L}$) of PbTiO$_3$ (PTO) is estimated using a combination of {\em ab initio} calculations and semiclassical Boltzmann transport equation. The computed $κ_{\rm L}$ is remarkably low, nearly comparable with the $κ_{\rm L}$ of good thermoelectric materials such as PbTe. In addition, a semiclassical analysis of the electronic transport quantities is presented, which suggests excellent thermoelectric properties, with a figure of merit $zT$ well over 1 for a wide range of temperature. For thermoelectric applications, the $κ_{\rm L}$ could be further reduced by utilizing different morphologies and compositions.

cond-mat.mtrl-sci

Mean vector testing for high dimensional dependent observations

When testing for the mean vector in a high dimensional setting, it is generally assumed that the observations are independently and identically distributed. However if the data are dependent, the existing test procedures fail to preserve type I error at a given nominal significance level. We propose a new test for the mean vector when the dimension increases linearly with sample size and the data is a realization of an M -dependent stationary process. The order M is also allowed to increase with the sample size. Asymptotic normality of the test statistic is derived by extending the central limit theorem result for M -dependent processes using two dimensional triangular arrays. Finite sample simulation results indicate the cost of ignoring dependence amongst observations.

math.ST

Estimation of Causal Invertible VARMA Models

We present a re-parameterization of vector autoregressive moving average (VARMA) models that allows estimation of parameters under the constraints of causality and invertibility. The parameter constraints associated with a causal invertible VARMA model are highly complex. Currently there are no procedures that can maintain the constraints in the estimated VARMA process, except in the special case of a vector autoregression (VAR), where some moment based causal estimators are available. Even in the VAR case, the available likelihood based estimators are not causal. The maximum likelihood estimator based on the full likelihood that does not condition on the initial observations by definition satisfies the causal invertible constraints but optimization of the likelihood under the complex constraints is an intractable problem. The commonly used Bayesian procedure for VAR often has posterior mass outside the causal set because the priors are not constrained to the causal set of parameters. We provide an exact mathematical solution to this problem. An $m$-variate VARMA$(p, q)$ process contains $(p+ q) m^2 + \binom{m+1}{2}$ parameters, which must be constrained to a subset of Euclidean space in order to guarantee causality and invertibility. This space is implicitly described in this paper, through the device of parameterizing the entire space of block Toeplitz matrices in terms of positive definite matrices and orthogonal matrices. The parameterization has connection to Schur- stability of polynomials and the associated Stein transformation that are often used in dynamical systems literature. As an important by-product of our investigation, we generalize a classical result in dynamical systems to provide a characterization of Schur stable matrix polynomials.

math.ST