SearcharxivSearch

arXiv subjects

Sanjay Chaudhuri

Publications and source records attributed to Sanjay Chaudhuri.

17 recordsLinked to original sources

Data-driven techniques for translational neuroscience and personalized neuro-health

Neurodegenexrative diseases such as Alzheimer's disease and Parkinson's disease are diagnosed most reliably only after substantial, often irreversible, neuronal loss has already occurred, creating an urgent need for quantitative tools that can detect subtle, early, and individual-specific brain changes from neuroimaging data. This review surveys a broad and rapidly evolving toolkit of data-driven techniques for translational neuroscience and personalized neuro-health, organized around four complementary methodological pillars. Throughout, we emphasize how these methodologically diverse approaches converge on a common translational goal: personalized, mechanistically grounded, and clinically actionable models of individual brain health, and we close by discussing the principal open statistical, computational, and clinical challenges that remain.

q-bio.NC

Uncertainty Quantification Via the Posterior Predictive Variance

We use the law of total variance to generate multiple expansions for the posterior predictive variance. These expansions are sums of terms involving conditional expectations and conditional variances and provide a quantification of the sources of predictive uncertainty. Since the posterior predictive variance is fixed given the model, it represents a constant quantity that is conserved over these expansions. The terms in the expansions can be assessed in absolute or relative sense to understand the main contributors to the length of prediction intervals. We quantify the term-wise uncertainty across expansions varying in the number of terms and the order of conditionates. In particular, given that a specific term in one expansion is small or zero, we identify the other terms in other expansions that must also be small or zero. We illustrate this approach to predictive model assessment in several well-known models.

math.ST

On an Empirical Likelihood based Solution to the Approximate Bayesian Computation Problem

Approximate Bayesian Computation (ABC) methods are applicable to statistical models specified by generative processes with analytically intractable likelihoods. These methods try to approximate the posterior density of a model parameter by comparing the observed data with additional process-generated simulated datasets. For computational benefit, only the values of certain well-chosen summary statistics are usually compared, instead of the whole dataset. Most ABC procedures are computationally expensive, justified only heuristically, and have poor asymptotic properties. In this article, we introduce a new empirical likelihood-based approach to the ABC paradigm called ABCel. The proposed procedure is computationally tractable and approximates the target log posterior of the parameter as a sum of two functions of the data -- namely, the mean of the optimal log-empirical likelihood weights and the estimated differential entropy of the summary functions. We rigorously justify the procedure via direct and reverse information projections onto appropriate classes of probability densities. Past applications of empirical likelihood in ABC demanded constraints based on analytically tractable estimating functions that involve both the data and the parameter; although by the nature of the ABC problem such functions may not be available in general. In contrast, we use constraints that are functions of the summary statistics only. Equally importantly, we show that our construction directly connects to the reverse information projection. We show that ABCel is posterior consistent and has highly favourable asymptotic properties. Its construction justifies the use of simple summary statistics like moments, quantiles, etc, which in practice produce an accurate approximation of the posterior density. We illustrate the performance of the proposed procedure in a range of applications.

stat.ME

Population level information combined parameter estimation from complex survey datasets

We consider an empirical likelihood framework for inference for a statistical model based on an informative sampling design and population-level information. The population-level information is summarized in the form of estimating equations and incorporated into the inference through additional constraints. Covariate information is incorporated both through the weights and the estimating equations. The estimator is based on conditional weights. We show that under usual conditions, with population size increasing unbounded, the estimates are strongly consistent, asymptotically unbiased, and normally distributed. Moreover, they are more efficient than other probability-weighted analogs. Our framework provides additional justification for inverse probability weighted score estimators in terms of conditional empirical likelihood. We give an application to demographic hazard modeling by combining birth registration data with panel survey data to estimate annual first birth probabilities.

stat.ME

A Two-step Metropolis Hastings Method for Bayesian Empirical Likelihood Computation with Application to Quantile Regression and Bayesian Model Selection

Empirical likelihood-based methods have been used under the Bayesian framework (BayesEL) in recent times. For statistical inference, these methods require efficient Markov chain Monte Carlo (MCMC) samplers for drawing observations from the parameter posterior distributions. However, the complex, especially non-convex, nature of the empirical likelihood support makes such MCMC algorithms harder to design. Such difficulties have restricted the use of BayesEL methods in many applications. In this article, we propose a two-step Metropolis-Hastings algorithm to sample from the BayesEL posteriors. Our proposal uses the current values of suitable subsets of the parameters and the estimating equations determining the underlying empirical likelihood to propose values of the remaining parameters. The proposed method is thus suitable for sampling from BayesEL posteriors in many complex problems, especially those with discontinuous estimating equations, e.g., simultaneous quantile regression. Furthermore, the proposed method easily extends to BayesEL model selection through a reversible jump Markov chain Monte Carlo procedure. Several illustrative, real-life applications of our proposed methods are presented.

stat.ME

elhmc: An R Package for Hamiltonian Monte Carlo Sampling in Bayesian Empirical Likelihood

In this article, we describe a {\tt R} package for sampling from an empirical likelihood-based posterior using a Hamiltonian Monte Carlo method. Empirical likelihood-based methodologies have been used in Bayesian modeling of many problems of interest in recent times. This semiparametric procedure can easily combine the flexibility of a non-parametric distribution estimator together with the interpretability of a parametric model. The model is specified by estimating equations-based constraints. Drawing an inference from a Bayesian empirical likelihood (BayesEL) posterior is challenging. The likelihood is computed numerically, so no closed expression of the posterior exists. Moreover, for any sample of finite size, the support of the likelihood is non-convex, which hinders the fast mixing of many Markov Chain Monte Carlo (MCMC) procedures. It has been recently shown that using the properties of the gradient of log empirical likelihood, one can devise an efficient Hamiltonian Monte Carlo (HMC) algorithm to sample from a BayesEL posterior. The package requires the user to specify only the estimating equations, the prior, and their respective gradients. An MCMC sample drawn from the BayesEL posterior of the parameters, with various details required by the user is obtained.

stat.OT

A Unified Statistical Procedure to Analyse Irreversible Thermal Curves

DNA hybridisation experiments are crucial for studying the thermodynamic and kinetic profiles of various systems in nucleic acid chemistry. The phenomenon of hysteresis is commonly observed in many such UV thermal experiments involving unmodified or modified nucleic acids. In the presence of hysteresis, the thermal curves are irreversible and demand a significant effort to produce the reaction-specific kinetic and thermodynamic parameters. In this article, we describe a unified statistical procedure to analyse such thermal curves. More specifically, the proposed method allows one to handle the thermal curves for the formation of duplexes, triplexes, and various quadruplexes in exactly the same way. The proposed method uses a local polynomial regression to find the smoothed thermal curves and calculate their slopes. This method is more flexible and easier to implement than the least squares polynomial smoothing, which is currently almost universally used for such purposes. Full analyses of the curves, including computation of kinetic and thermodynamic parameters, can be done using freely available statistical software. The proposed procedure has been implemented in a web-based free software called anhysnuc, which can be found at https://sanjaychaudhuri.shinyapps.io/anhysnuc/. Finally, we illustrate our method by analysing irreversible curves encountered in the formation of a G-quadruplex and an LNA-modified parallel duplex.

physics.chem-ph

General Unbiased Estimating Equations for Variance Components in Linear Mixed Models

This paper introduces a general framework for estimating variance components in the linear mixed models via general unbiased estimating equations, which include some well-used estimators such as the restricted maximum likelihood estimator. We derive the asymptotic covariance matrices and second-order biases under general estimating equations without assuming the normality of the underlying distributions and identify a class of second-order unbiased estimators of variance components. It is also shown that the asymptotic covariance matrices and second-order biases do not depend on whether the regression coefficients are estimated by the generalized or ordinary least squares methods. We carry out numerical studies to check the performance of the proposed method based on typical linear mixed models.

stat.ME

On a Variational Approximation based Empirical Likelihood ABC Method

Many scientifically well-motivated statistical models in natural, engineering, and environmental sciences are specified through a generative process. However, in some cases, it may not be possible to write down the likelihood for these models analytically. Approximate Bayesian computation (ABC) methods allow Bayesian inference in such situations. The procedures are nonetheless typically computationally intensive. Recently, computationally attractive empirical likelihood-based ABC methods have been suggested in the literature. All of these methods rely on the availability of several suitable analytically tractable estimating equations, and this is sometimes problematic. We propose an easy-to-use empirical likelihood ABC method in this article. First, by using a variational approximation argument as a motivation, we show that the target log-posterior can be approximated as a sum of an expected joint log-likelihood and the differential entropy of the data generating density. The expected log-likelihood is then estimated by an empirical likelihood where the only inputs required are a choice of summary statistic, it's observed value, and the ability to simulate the chosen summary statistics for any parameter value under the model. The differential entropy is estimated from the simulated summaries using traditional methods. Posterior consistency is established for the method, and we discuss the bounds for the required number of simulated summaries in detail. The performance of the proposed method is explored in various examples.

stat.ME

Maximum Likelihood under constraints: Degeneracies and Random Critical Points

We investigate the problem of semi-parametric maximum likelihood under constraints on summary statistics. Such a procedure results in a discrete probability distribution that maximises the likelihood among all such distributions under the specified constraints (called estimating equations), and is an approximation to the underlying population distribution. The study of such empirical likelihood originates from the seminal work of Owen. We investigate this procedure in the setting of mis-specified (or biased) estimating equations, i.e. when the null hypothesis is not true. We establish that the behaviour of the optimal distribution under such mis-specification differ markedly from their properties under the null, i.e. when the estimating equations are unbiased and correctly specified. This is manifested by certain degeneracies in the optimal distribution which define the likelihood. Such degeneracies are not observed under the null. Furthermore, we establish an anomalous behaviour of the log-likelihood based Wilks statistic, which, unlike under the null, does not exhibit a chi-squared limit. In the Bayesian setting, we rigorously establish the posterior consistency of procedures based on these ideas, where instead of a parametric likelihood, an empirical likelihood is used to define the posterior distribution. In particular, we show that this posterior, as a random probability measure, rapidly converges to the delta measure at the true parameter value. A novel feature of our approach is the investigation of critical points of random functions in the context of such empirical likelihood. In particular, we obtain the location and the mass of the degenerate optimal weights as the leading and sub-leading terms in a canonical expansion of a particular critical point of a random function that is naturally associated with the model.

math.ST

A Conditional Empirical Likelihood Based Method for Model Parameter Estimation from Complex survey Datasets

We consider an empirical likelihood framework for inference for a statistical model based on an informative sampling design. Covariate information is incorporated both through the weights and the estimating equations. The estimator is based on conditional weights. We show that under usual conditions, with population size increasing unbounded, the estimates are strongly consistent, asymptotically unbiased and normally distributed. Our framework provides additional justification for inverse probability weighted score estimators in terms of conditional empirical likelihood. In doing so, it bridges the gap between design-based and model-based modes of inference in survey sampling settings. We illustrate these ideas with an application to an electoral survey.

stat.ME

An easy-to-use empirical likelihood ABC method

Many scientifically well-motivated statistical models in natural, engineering and environmental sciences are specified through a generative process, but in some cases it may not be possible to write down a likelihood for these models analytically. Approximate Bayesian computation (ABC) methods, which allow Bayesian inference in these situations, are typically computationally intensive. Recently, computationally attractive empirical likelihood based ABC methods have been suggested in the literature. These methods heavily rely on the availability of a set of suitable analytically tractable estimating equations. We propose an easy-to-use empirical likelihood ABC method, where the only inputs required are a choice of summary statistic, it's observed value, and the ability to simulate summary statistics for any parameter value under the model. It is shown that the posterior obtained using the proposed method is consistent, and its performance is explored using various examples.

stat.CO

Qualitative inequalities for squared partial correlations of a Gaussian random vector

We describe various sets of conditional independence relationships, sufficient for qualitatively comparing non-vanishing squared partial correlations of a Gaussian random vector. These sufficient conditions are satisfied by several graphical Markov models. Rules for comparing degree of association among the vertices of such Gaussian graphical models are also developed. We apply these rules to compare conditional dependencies on Gaussian trees. In particular for trees, we show that such dependence can be completely characterized by the length of the paths joining the dependent vertices to each other and to the vertices conditioned on. We also apply our results to postulate rules for model selection for polytree models. Our rules apply to mutual information of Gaussian random vectors as well.

math.ST

Variance Estimation for Tree Order Restricted Models

In this article we discuss estimation of the common variance of several normal populations with tree order restricted means. We discuss the asymptotic properties of the maximum likelihood estimator of the variance as the number of populations tends to infinity. We consider several cases of various orders of the sample sizes and show that the maximum likelihood estimator of the variance may or may not be consistent or be asymptotically normal.

math.ST

Reversing the Stein Effect

The Reverse Stein Effect is identified and illustrated: A statistician who shrinks his/her data toward a point chosen without reliable knowledge about the underlying value of the parameter to be estimated but based instead upon the observed data will not be protected by the minimax property of shrinkage estimators such as that of James and Stein, but instead will likely incur a greater error than if shrinkage were not used.

stat.ME

Estimation of a Covariance Matrix with Zeros

We consider estimation of the covariance matrix of a multivariate random vector under the constraint that certain covariances are zero. We first present an algorithm, which we call Iterative Conditional Fitting, for computing the maximum likelihood estimator of the constrained covariance matrix, under the assumption of multivariate normality. In contrast to previous approaches, this algorithm has guaranteed convergence properties. Dropping the assumption of multivariate normality, we show how to estimate the covariance matrix in an empirical likelihood approach. These approaches are then compared via simulation and on an example of gene expression.

math.ST

Quasiperiodic waves at the onset of zero Prandtl number convection with rotation

We show the possibility of quasiperiodic waves at the onset of thermal convection in a thin horizontal layer of slowly rotating zero-Prandtl number Boussinesq fluid confined between stress-free conducting boundaries. Two independent frequencies emerge due to an interaction between a stationary instability and a self-tuned wavy instability in presence of coriolis force, if Taylor number is raised above a critical value. Constructing a dynamical system for the hydrodynamical problem, the competition between the interacting instabilities is analyzed. The forward bifurcation from the conductive state is self-tuned.

physics.flu-dyn