Searcharxiv⌕ Search

arXiv subjects

Arkaprava Roy

Publications and source records attributed to Arkaprava Roy.

34 records · Page 2Linked to original sources

Nonparametric Modeling of Diffusion MRI Signal in Q-space

This paper describes a novel nonparametric model for modeling diffusion MRI signals in q-space. In q-space, diffusion MRI signal is measured for a sequence of magnetic strengths (b-values) and magnetic gradient directions (b-vectors). We propose a Poly-RBF model, which employs a bidirectional framework with polynomial bases to model the signal along the b-value direction and Gaussian radial bases across the b-vectors. The model can accommodate sparse data on b-values and moderately dense data on b-vectors. The utility of Poly-RBF is inspected for two applications: 1) prediction of the dMRI signal, and 2) harmonization of dMRI data collected under different acquisition protocols with different scanners. Our results indicate that the proposed Poly-RBF model can more accurately predict the unmeasured diffusion signal than its competitors such as the Gaussian process model in {\tt Eddy} of FSL. Applying it to harmonizing the diffusion signal can significantly improve the reproducibility of derived white matter microstructure measures.

stat.AP↗

Linking stability with molecular geometries of perovskites and lanthanide richness using machine learning methods

Oxide perovskite materials of type ABO3 have a wide range of technological applications, such as catalysts in solid oxide fuel cells and as light-absorbing materials in solar photovoltaics. These materials often exhibit differential structural and electrostatic properties through lanthanide or non-lanthanide derived A- and B- sites. Although, experimental and/or computational verification of these differences are often difficult. In this paper, we thus take a data-driven approach. Specifically, we run three analysis using the dataset Li, Jacobs, and Morgan [2018a] applying advanced machine learning tools to perform nonparametric regressions and also to produce data visualizations using latent factor analysis (LFA) and principal component analysis (PCA). We also implement a nonparametric feature screening step while performing our high dimensional regression analysis, ensuring robustness in our results

cond-mat.mtrl-sci↗

Spectral Clustering, Bayesian Spanning Forest, and Forest Process

Spectral clustering views the similarity matrix as a weighted graph, and partitions the data by minimizing a graph-cut loss. Since it minimizes the across-cluster similarity, there is no need to model the distribution within each cluster. As a result, one reduces the chance of model misspecification, which is often a risk in mixture model-based clustering. Nevertheless, compared to the latter, spectral clustering has no direct ways of quantifying the clustering uncertainty (such as the assignment probability), or allowing easy model extensions for complicated data applications. To fill this gap, we propose the Bayesian forest model as a generative graphical model for spectral clustering. This is motivated by our discovery that the posterior connecting matrix in a forest model has almost the same leading eigenvectors, as the ones used by normalized spectral clustering. To induce a distribution for the forest, we develop a ``forest process'' as a graph extension to the urn process, while we carefully characterize the differences in the partition probability. We derive a simple Markov chain Monte Carlo algorithm for posterior estimation, and demonstrate superior performance compared to existing algorithms. We illustrate several model-based extensions useful for data applications, including high-dimensional and multi-view clustering for images.

stat.ME↗

Nonparametric Group Variable Selection with Multivariate Response for Connectome-Based Modeling of Cognitive Scores

In this article, we study association between the structural connectome and cognitive profiles using a multi-response nonparametric regression model.The cognitive profiles are measured in terms of seven age-adjusted cognitive test scores. The structural connectomes are represented by undirected graphs. The connectivity properties of these graphs are available in terms of the nodal attributes. A collection of nodal centralities together can encode different patterns of connections in the brain network. In this article, we consider nine such attributes for each brain region.These nodal graph metrics may naturally be grouped together for each node, motivating us to introduce group sparsity for feature selection. We propose Gaussian RBF-nets with a novel group sparsity inducing prior to model the unknown mean functions. The covariance structure of the multivariate response is characterized in terms of a linear factor modeling framework. For posterior computation, we develop an efficient Markov chain Monte Carlo sampling algorithm. We show that the proposed method performs much better than all its competitors. Applying our proposed method to a Human Connectome Project (HCP) dataset, we identify the important brain regions and nodal attributes for cognitive functioning, as well as identify interesting low-dimensional dependency structures among the cognition related test scores. Keywords: Factor model; Group variable selection; High-dimension; Human Connectome Project (HCP); Markov chain Monte Carlo (MCMC); Neural network; Nonparametric inference; Radial basis network; Spike-and-slab prior; Variable selection.

stat.ME↗

Bayesian Semiparametric Multivariate Density Deconvolution via Stochastic Rotation of Replicates

We consider the problem of multivariate density deconvolution where the distribution of a random vector needs to be estimated from replicates contaminated with conditionally heteroscedastic measurement errors. We propose a conceptually straightforward yet fundamentally novel and highly robust approach to multivariate density deconvolution by stochastically rotating the replicates toward the corresponding true latent values. We also address the additionally significantly challenging problem of accommodating conditionally heteroscedastic measurement errors in this newly introduced framework. We take a Bayesian route to estimation and inference, implemented via an efficient Markov chain Monte Carlo algorithm, appropriately accommodating uncertainty in all aspects of our analysis. Asymptotic convergence guarantees for the method are also established. We illustrate the method's empirical efficacy through simulation experiments and its practical utility in estimating the long-term joint average intakes of different dietary components from their measurement error contaminated 24-hour dietary recalls.

stat.ME↗

Optimal Bayesian Smoothing of Functional Observations over a Large Graph

In modern contexts, some types of data are observed in high-resolution, essentially continuously in time. Such data units are best described as taking values in a space of functions. Subject units carrying the observations may have intrinsic relations among themselves, and are best described by the nodes of a large graph. It is often sensible to think that the underlying signals in these functional observations vary smoothly over the graph, in that neighboring nodes have similar underlying signals. This qualitative information allows borrowing of strength over neighboring nodes and consequently leads to more accurate inference. In this paper, we consider a model with Gaussian functional observations and adopt a Bayesian approach to smoothing over the nodes of the graph. We characterize the minimax rate of estimation in terms of the regularity of the signals and their variation across nodes quantified in terms of the graph Laplacian. We show that an appropriate prior constructed from the graph Laplacian can attain the minimax bound, while using a mixture prior, the minimax rate up to a logarithmic factor can be attained simultaneously for all possible values of functional and graphical smoothness. We also show that in the fixed smoothness setting, an optimal sized credible region has arbitrarily high frequentist coverage. A simulation experiment demonstrates that the method performs better than potential competing methods like the random forest. The method is also applied to a dataset on daily temperatures measured at several weather stations in the US state of North Carolina.

stat.ME↗

Time-varying auto-regressive models for count time-series

Count-valued time series data are routinely collected in many application areas. We are particularly motivated to study the count time series of daily new cases, arising from COVID-19 spread. We propose two Bayesian models, a time-varying semiparametric AR(p) model for count and then a time-varying INGARCH model considering the rapid changes in the spread. We calculate posterior contraction rates of the proposed Bayesian methods with respect to average Hellinger metric. Our proposed structures of the models are amenable to Hamiltonian Monte Carlo (HMC) sampling for efficient computation. We substantiate our methods by simulations that show superiority compared to some of the close existing methods. Finally we analyze the daily time series data of newly confirmed cases to study its spread through different government interventions.

stat.ME↗

Bayesian modelling of time-varying conditional heteroscedasticity

Conditional heteroscedastic (CH) models are routinely used to analyze financial datasets. The classical models such as ARCH-GARCH with time-invariant coefficients are often inadequate to describe frequent changes over time due to market variability. However we can achieve significantly better insight by considering the time-varying analogues of these models. In this paper, we propose a Bayesian approach to the estimation of such models and develop computationally efficient MCMC algorithm based on Hamiltonian Monte Carlo (HMC) sampling. We also established posterior contraction rates with increasing sample size in terms of the average Hellinger metric. The performance of our method is compared with frequentist estimates and estimates from the time constant analogues. To conclude the paper we obtain time-varying parameter estimates for some popular Forex (currency conversion rate) and stock market datasets.

math.ST↗

EmotionGIF-IITP-AINLPML: Ensemble-based Automated Deep Neural System for predicting category(ies) of a GIF response

In this paper, we describe the systems submitted by our IITP-AINLPML team in the shared task of SocialNLP 2020, EmotionGIF 2020, on predicting the category(ies) of a GIF response for a given unlabelled tweet. For the round 1 phase of the task, we propose an attention-based Bi-directional GRU network trained on both the tweet (text) and their replies (text wherever available) and the given category(ies) for its GIF response. In the round 2 phase, we build several deep neural-based classifiers for the task and report the final predictions through a majority voting based ensemble technique. Our proposed models attain the best Mean Recall (MR) scores of 52.92% and 53.80% in round 1 and round 2, respectively.

cs.CL↗

Perturbed factor analysis: Accounting for group differences in exposure profiles

In this article, we investigate group differences in phthalate exposure profiles using NHANES data. Phthalates are a family of industrial chemicals used in plastics and as solvents. There is increasing evidence of adverse health effects of exposure to phthalates on reproduction and neuro-development, and concern about racial disparities in exposure. We would like to identify a single set of low-dimensional factors summarizing exposure to different chemicals, while allowing differences across groups. Improving on current multi-group additive factor models, we propose a class of Perturbed Factor Analysis (PFA) models that assume a common factor structure after perturbing the data via multiplication by a group-specific matrix. Bayesian inference algorithms are defined using a matrix normal hierarchical model for the perturbation matrices. The resulting model is just as flexible as current approaches in allowing arbitrarily large differences across groups but has substantial advantages that we illustrate in simulation studies. Applying PFA to NHANES data, we learn common factors summarizing exposures to phthalates, while showing clear differences across groups.

stat.ME↗

Analyzing initial stage of COVID-19 transmission through Bayesian time-varying model

Recent outbreak of the novel coronavirus COVID-19 has affected all of our lives in one way or the other. While medical researchers are working hard to find a cure and doctors/nurses to attend the affected individuals, measures such as `lockdown', `stay-at-home', `social distancing' are being implemented in different parts of the world to curb its further spread. To model the non-stationary spread, we propose a novel time-varying semiparametric AR$(p)$ model for the count valued time-series of newly affected cases, collected every day and also extend it to propose a novel time-varying INGARCH model. Our proposed structures of the models are amenable to Hamiltonian Monte Carlo (HMC) sampling for efficient computation. We substantiate our methods by simulations that show superiority compared to some of the close existing methods. Finally we analyze the daily time series data of newly confirmed cases to study its spread through different government interventions.

stat.ME↗

Spatial shrinkage via the product independent Gaussian process prior

We study the problem of sparse signal detection on a spatial domain. We propose a novel approach to model continuous signals that are sparse and piecewise smooth as product of independent Gaussian processes (PING) with a smooth covariance kernel. The smoothness of the PING process is ensured by the smoothness of the covariance kernels of Gaussian components in the product, and sparsity is controlled by the number of components. The bivariate kurtosis of the PING process shows more components in the product results in thicker tail and sharper peak at zero. The simulation results demonstrate the improvement in estimation using the PING prior over Gaussian process (GP) prior for different image regressions. We apply our method to a longitudinal MRI dataset to detect the regions that are affected by multiple sclerosis (MS) in the greatest magnitude through an image-on-scalar regression model. Due to huge dimensionality of these images, we transform the data into the spectral domain and develop methods to conduct computation in this domain. In our MS imaging study, the estimates from the PING model are more informative than those from the GP model.

stat.ME↗

Nonparametric graphical model for counts

Although multivariate count data are routinely collected in many application areas, there is surprisingly little work developing flexible models for characterizing their dependence structure. This is particularly true when interest focuses on inferring the conditional independence graph. In this article, we propose a new class of pairwise Markov random field-type models for the joint distribution of a multivariate count vector. By employing a novel type of transformation, we avoid restricting to non-negative dependence structures or inducing other restrictions through truncations. Taking a Bayesian approach to inference, we choose a Dirichlet process prior for the distribution of a random effect to induce great flexibility in the specification. An efficient Markov chain Monte Carlo (MCMC) algorithm is developed for posterior computation. We prove various theoretical properties, including posterior consistency, and show that our COunt Nonparametric Graphical Analysis (CONGA) approach has good performance relative to competitors in simulation studies. The methods are motivated by an application to neuron spike count data in mice.

stat.ME↗

Bayesian time-aligned factor analysis of paired multivariate time series

Many modern data sets require inference methods that can estimate the shared and individual-specific components of variability in collections of matrices that change over time. Promising methods have been developed to analyze these types of data in static cases, but very few approaches are available for dynamic settings. To address this gap, we consider novel models and inference methods for pairs of matrices in which the columns correspond to multivariate observations at different time points. In order to characterize common and individual features, we propose a Bayesian dynamic factor modeling framework called Time Aligned Common and Individual Factor Analysis (TACIFA) that includes uncertainty in time alignment through an unknown warping function. We provide theoretical support for the proposed model, showing identifiability and posterior concentration. The structure enables efficient computation through a Hamiltonian Monte Carlo (HMC) algorithm. We show excellent performance in simulations, and illustrate the method through application to a social synchrony experiment.

stat.ME↗

Bayesian Modeling of the Structural Connectome for Studying Alzheimer Disease

We study possible relations between the structure of the connectome, white matter connecting different regions of brain, and Alzheimer disease. Regression models in covariates including age, gender and disease status for the extent of white matter connecting each pair of regions of brain are proposed. Subject We study possible relations between the Alzheimer's disease progression and the structure of the connectome, white matter connecting different regions of brain. Regression models in covariates including age, gender and disease status for the extent of white matter connecting each pair of regions of brain are proposed. Subject inhomogeneity is also incorporated in the model through random effects with an unknown distribution. As there are large number of pairs of regions, we also adopt a dimension reduction technique through graphon (Lovasz and Szegedy (2006)) functions, which reduces functions of pairs of regions to functions of regions. The connecting graphon functions are considered unknown but assumed smoothness allows putting priors of low complexity on them. We pursue a nonparametric Bayesian approach by assigning a Dirichlet process scale mixture of zero mean normal prior on the distributions of the random effects and finite random series of tensor products of B-splines priors on the underlying graphon functions. Markov chain Monte Carlo techniques, for drawing samples for the posterior distributions are developed. The proposed Bayesian method overwhelmingly outperforms similar ANCOVA models in the simulation setup. The proposed Bayesian approach is applied on a dataset of 100 subjects and 83 brain regions and key regions implicated in the changing connectome are identified.

stat.ME↗

High-dimensional single-index Bayesian modeling of brain atrophy

We propose a model of brain atrophy as a function of high-dimensional genetic information and low dimensional covariates such as gender, age, APOE gene, and disease status. A nonparametric single-index Bayesian model of high dimension is proposed to model the relationship with B-spline series prior on the unknown functions and Dirichlet process scale mixture of centered normal prior on the distributions of the random effects. The posterior rate of contraction without the random effect is established for a fixed number of regions and time points with increasing sample size. We implement an efficient computation algorithm through a Hamiltonian Monte Carlo (HMC) algorithm. The performance of the proposed Bayesian method is compared with the corresponding least square estimator in the linear model with horseshoe prior, LASSO and SCAD penalization on the high-dimensional covariates. The proposed Bayesian method is applied to a dataset on volumes of brain regions recorded over multiple visits of 748 individuals using 620,901 SNPs and 6 other covariates for each individual, to identify factors associated with brain atrophy.

stat.ME↗