SearcharxivSearch

arXiv subjects

Dipankar Bandyopadhyay

Publications and source records attributed to Dipankar Bandyopadhyay.

At least 19 recordsLinked to original sources

A Monotone Single-Index Modal Regression Powered by Deep Neural Networks for Non-Gaussian Periodontal Data

Pocket depth (PD) is a widely used biomarker for diagnosing risk of periodontal disease (PrD). However, PD typically exhibits skewness and heavy-tailedness, and its relationship with clinical risk factors is often nonlinear. Motivated by PrD studies, this paper develops a robust single-index modal regression framework for analyzing skewed and heavy-tailed data. Our method has the following novel features: (a) a flexible two-piece scale Student-$t$ error distribution that generalizes both normal and two-piece scale normal distributions; (b) a neural network with guaranteed monotonicity constraints to estimate the unknown single-index function; and (c) theoretical \revone{support}, including model identifiability and a universal approximation theorem. Our single-index model combines the flexibility of neural networks and the two-piece scaled Student-$t$ distribution, delivering robust mode-based estimation that is resistant to outliers, while retaining clinical interpretability through parametric index coefficients. We demonstrate the performance of our method through simulation studies, and an application to PrD electronic health records obtained from the HealthPartners Institute of Minnesota. The proposed methodology is implemented in the \texttt{R} package \href{https://doi.org/10.32614/CRAN.package.DNNSIM}{\texttt{DNNSIM}}.

stat.ME

Censored broken adaptive ridge rank regression via induced smoothing

Broken adaptive ridge (BAR) penalty approximates $L_0$-regularization through iterative reweighting of L2 penalties. This penalty enjoys both the oracle property and the grouping effect for highly correlated covariates, making it particularly attractive for penalized regression with complex dependence among predictors. In this paper, we develop a BAR-penalized linear rank regression method for the semiparametric accelerated failure time model with right-censored data. Computational tractability is achieved by applying induced smoothing to the nonsmooth Gehan-type rank estimating function, yielding a more stable framework for estimation and inference. For scalable penalization, we develop a cyclic coordinate descent algorithm that minimizes the penalized objective function, and estimates the regression coefficients in a coordinate-wise manner. We further extend the proposed method to more complex survival endpoints, such as multivariate partly interval-censored (PIC) data. Under mild conditions, the proposed estimator satisfies both the oracle property and the grouping effect, and the variance estimator of the informative coefficients can be derived in analytic form. Numerical studies using synthetic data compare our approach to several well-known penalties, and demonstrate its superior selection accuracy and estimation efficiency across various scenarios. Furthermore, applications to right-censored outcomes from primary biliary cirrhosis, and correlated PIC outcomes from colorectal cancer further illustrate the practical utility of the proposed method. The R package aftPenCDA for implementing the method is available on R CRAN.

stat.ME

Semi- and non-parametric approaches to individualized treatment regimes in the presence of causal mediation

Individualized treatment rules (ITRs) map an individual patient's characteristics to their recommended treatment value. Typically, the optimal ITR is defined as the rule which maximizes a mean counterfactual outcome; the resulting ITR maximizes the effect of treatment along all causal pathways to the outcome, including indirect pathways through mediating variables. Although maximizing the total effect is often sufficient, explicitly incorporating causal mediation in an ITR analysis has several potential benefits such as enhanced interpretability, and additional flexibility in targeting specific causal pathways. For this purpose, we introduce novel Bayesian semiparametric and nonparametric estimators for conditional mediation effects in the presence of multiple mediators and show how they can be used to estimate optimal ITRs. We demonstrate the proposed methodology via an application to optimal kidney allocation with hepatitis C positive donors.

stat.ME

Rank estimation for the accelerated failure time model with partially interval-censored data

This paper presents a unified rank-based inferential procedure for fitting the accelerated failure time model to partially interval-censored data. A Gehan-type monotone estimating function is constructed based on the idea of the familiar weighted log-rank test, and an extension to a general class of rank-based estimating functions is suggested. The proposed estimators can be obtained via linear programming and are shown to be consistent and asymptotically normal via standard empirical process theory. Unlike common maximum likelihood-based estimators for partially interval-censored regression models, our approach can directly provide a regression coefficient estimator without involving a complex nonparametric estimation of the underlying residual distribution function. An efficient variance estimation procedure for the regression coefficient estimator is considered. Moreover, we extend the proposed rank-based procedure to the linear regression analysis of multivariate clustered partially interval-censored data. The finite-sample operating characteristics of our approach are examined via simulation studies. Data example from a colorectal cancer study illustrates the practical usefulness of the method.

stat.ME

Asynchronous Distributed ECME Algorithm for Matrix Variate Non-Gaussian Responses

We propose a regression model with matrix-variate skew-t response (REGMVST) for analyzing irregular longitudinal data with skewness, symmetry, or heavy tails. REGMVST models matrix-variate responses and predictors, with rows indexing longitudinal measurements per subject. It uses the matrix-variate skew-t (MVST) distribution to handle skewness and heavy tails, a damped exponential correlation (DEC) structure for row-wise dependencies across irregular time profiles, and leaves the column covariance unstructured. For estimation, we initially develop an ECME algorithm for parameter estimation and further mitigate its computational bottleneck via an asynchronous and distributed ECME (ADECME) extension. ADECME accelerates the E-step through parallelization, and retains the simplicity of the conditional M-step, enabling scalable inference. Simulations using synthetic data and a case study exploring matrix-variate periodontal disease endpoints derived from electronic health records demonstrate ADECME's superiority in efficiency and convergence, over the alternatives. We also provide theoretical support for our empirical observations and identify regularity assumptions for ADECME's optimal performance. An accompanying R package is available at https://github.com/rh8liuqy/STMATREG.

stat.ME

Rapid and Cost-Effective In situ Fabrication of Nanoporous Membranes in Microfluidic Devices for Biomedical and Environmental Applications

Micro/nanoporous membranes have been used extensively in biological and medical applications associated with filtration, particle sorting, cell separation, and real-time sensing. Expanding its applications to microfluidics offers several advantages, such as precise control over fluid flow rates, regulated reaction times, and efficient product extraction. This study demonstrates a simple, cost-effective, rapid, single-step method for the synthesis of nanoporous membranes within a microfluidic device. The membrane is prepared by filling up the microchamber with cellulose acetate dissolved in N, N-dimethylformamide, and 1-hexanol, followed by the continuous flow of water as an antisolvent. Mixing CA-DMF with antisolvent water leads to an in-situ synthesis of cellulose acetate nanoparticles (CANPs) at the Water-DMF interface. The continuous flow of water ensures the coagulation of CANPs, forming nanoporous membranes. The thickness, porosity, and wettability of the membrane are dependent on the flow rate of water and the concentration of CA dissolved in DMF. Apart from its wide application in cell and biomolecule separation, this membrane can be used to detect biomarkers or pathogens in clinical samples. Additionally, these nanoporous membranes can be used to detect and filter environmental contaminants such as heavy metals, pesticides, and emulsified oil from water. In summary, the synthesis of CANP nanoporous membranes within microfluidic channels with adjustable porosity enhances the potential uses of micro/nanomembranes in both healthcare and environmental contexts.

physics.flu-dyn

Individualized treatment regimens under correlated data with multiple outcomes

Precision medicine involves developing individualized treatment regimes (ITRs) which allow for treatment decisions to be tailored to patient characteristics. Naturally, the identification of the optimal regime, that is, the rule which maximizes patient outcomes, is of interest. Several procedures for estimating optimal ITRs from observational data have been proposed; however, relatively few methods exist for estimating optimal ITRs in the presence of competing risks. Previous approaches either target one particular cause of failure, or rely on singly-robust estimators. We propose a novel doubly-robust regression-based method for estimating optimal ITRs which accounts for the uncertainty related to the unobserved cause of failure by averaging over all possible causes, or targeting the most likely cause. Our approach is straightforward to implement, and we demonstrate an extension to incorporate clustering, motivated by the question of for whom kidney transplantation with hepatitis C virus (HCV)-positive donors is safe, using data from the Organ Procurement and Transplantation Network. Our analysis suggests that a large portion of HCV-negative kidney recipients would see their overall survival unchanged if they were instead provided a kidney from an HCV-positive donor. The estimated treatment rules could be used to provide more efficient allocation of HCV-positive kidneys, increasing the donor pool.

stat.ME

An Interpretable Single-Index Mixed-Effects Model for Non-Gaussian National Survey Data

This manuscript presents an innovative statistical model to quantify periodontal disease in the context of complex medical data. A mixed-effects model incorporating skewed random effects and heavy-tailed residuals is introduced, ensuring robust handling of non-normal data distributions. The fixed effect is modeled as a combination of a slope parameter and a single index function, constrained to be monotonic increasing for meaningful interpretation. This approach captures different dimensions of periodontal disease progression by integrating Clinical Attachment Level (CAL) and Pocket Depth (PD) biomarkers within a unified analytical framework. A variable selection method based on the grouped horseshoe prior is employed, addressing the relatively high number of risk factors. Furthermore, survey weight information typically provided with large survey data is incorporated to ensure accurate inference. This comprehensive methodology significantly advances the statistical quantification of periodontal disease, offering a nuanced and precise assessment of risk factors and disease progression. The proposed methodology is implemented in the \textsf{R} package \href{https://cran.r-project.org/package=MSIMST}{\textsc{MSIMST}}.

stat.ME

A monotone single index model for spatially referenced multistate current status data

Assessment of multistate disease progression is commonplace in biomedical research, such as, in periodontal disease (PD). However, the presence of multistate current status endpoints, where only a single snapshot of each subject's progression through disease states is available at a random inspection time after a known starting state, complicates the inferential framework. In addition, these endpoints can be clustered, and spatially associated, where a group of proximally located teeth (within subjects) may experience similar PD status, compared to those distally located. Motivated by a clinical study recording PD progression, we propose a Bayesian semiparametric accelerated failure time model with an inverse-Wishart proposal for accommodating (spatial) random effects, and flexible errors that follow a Dirichlet process mixture of Gaussians. For clinical interpretability, the systematic component of the event times is modeled using a monotone single index model, with the (unknown) link function estimated via a novel integrated basis expansion and basis coefficients endowed with constrained Gaussian process priors. In addition to establishing parameter identifiability, we present scalable computing via a combination of elliptical slice sampling, fast circulant embedding techniques, and smoothing of hard constraints, leading to straightforward estimation of parameters, and state occupation and transition probabilities. Using synthetic data, we study the finite sample properties of our Bayesian estimates, and their performance under model misspecification. We also illustrate our method via application to the real clinical PD dataset.

stat.ME

Zero & $N$-inflated overdispersed binomial models for sum-constrained Poisson count processes

A frequent challenge encountered with compositional ecological data is how to interpret and model data with a high proportion of zeros and $N$'s. Such data frequently occur in ecological applications where counts of species are collected until a pre-specified total imposed (typically) by sampling cost is reached. In the bivariate count (two-species) setting we focus on in this article, zero-inflation of one species will result in $N$-inflation of the other. This can lead to species absence being attributed to an unsuitable habitat as opposed to missingness by chance. Similarly, an excess of $N$'s will lead to misleading inferences about habitat preference and abundance estimates. Our contribution is to identify that two independent zero-inflated Poisson processes subject to a sum constraint provide a novel biologically-motivated generating mechanism for the occurrence of binomial count data exhibiting zero and $N$-inflation. We identify an extension to the model to capture additional overdispersion within the data resulting in a novel zero and $N$-inflated beta-binomial model. We consider two motivating datasets, one involving a pesticide treatment for an invasive species, and a second involving the abundance of two plant species. We demonstrate that incorporation of covariates in each case enable learning about sources of zero and $N$-inflation as well as abundance. We show that the models result in improved understanding of underlying biological processes as well as improved predictive performance.

stat.ME

Scalable Efficient Inference in Complex Surveys through Targeted Resampling of Weights

Survey data often arises from complex sampling designs, such as stratified or multistage sampling, with unequal inclusion probabilities. When sampling is informative, traditional inference methods yield biased estimators and poor coverage. Classical pseudo-likelihood based methods provide accurate asymptotic inference but lack finite-sample uncertainty quantification and the ability to integrate prior information. Existing Bayesian approaches, like the Bayesian pseudo-posterior estimator and weighted Bayesian bootstrap, have limitations; the former struggles with uncertainty quantification, while the latter is computationally intensive and sensitive to bootstrap replicates. To address these challenges, we propose the Survey-adjusted Weighted Likelihood Bootstrap (S-WLB), which resamples weights from a carefully chosen distribution centered around the underlying sampling weights. S-WLB is computationally efficient, theoretically consistent, and delivers finite-sample uncertainty intervals which are proven to be asymptotically valid. We demonstrate its performance through simulations and applications to nationally representative survey datasets like NHANES and NSDUH.

stat.ME

Off-Policy Evaluation with Irregularly-Spaced, Outcome-Dependent Observation Times

While the classic off-policy evaluation (OPE) literature commonly assumes decision time points to be evenly spaced for simplicity, in many real-world scenarios, such as those involving user-initiated visits, decisions are made at irregularly-spaced and potentially outcome-dependent time points. For a more principled evaluation of the dynamic policies, this paper constructs a novel OPE framework, which concerns not only the state-action process but also an observation process dictating the time points at which decisions are made. The framework is closely connected to the Markov decision process in computer science and with the renewal process in the statistical literature. Within the framework, two distinct value functions, derived from cumulative reward and integrated reward respectively, are considered, and statistical inference for each value function is developed under revised Markov and time-homogeneous assumptions. The validity of the proposed method is further supported by theoretical results, simulation studies, and a real-world application from electronic health records (EHR) evaluating periodontal disease treatments.

stat.ME

Joint modelling of time-to-event and longitudinal response using robust skew normal-independent distributions

Joint modelling of longitudinal observations and event times continues to remain a topic of considerable interest in biomedical research. For example, in HIV studies, the longitudinal bio-marker such as CD4 cell count in a patient's blood over follow up months is jointly modelled with the time to disease progression, death or dropout via a random intercept term mostly assumed to be Gaussian. However, longitudinal observations in these kinds of studies often exhibit non-Gaussian behavior (due to high degree of skewness), and parameter estimation is often compromised under violations of the Gaussian assumptions. In linear mixed-effects model assumptions, the distributional assumption for the subject-specific random-effects is taken as Gaussian which may not be true in many situations. Further, this assumption makes the model extremely sensitive to outlying observations. We address these issues in this work by devising a joint model which uses a robust distribution in a parametric setup along with a conditional distributional assumption that ensures dependency of two processes in case the subject-specific random effects is given.

stat.ME

Computationally Scalable Bayesian SPDE Modeling for Censored Spatial Responses

Observations of groundwater pollutants, such as arsenic or Perfluorooctane sulfonate (PFOS), are riddled with left censoring. These measurements have impact on the health and lifestyle of the populace. Left censoring of these spatially correlated observations are usually addressed by applying Gaussian processes (GPs), which have theoretical advantages. However, this comes with a challenging computational complexity of $\mathcal{O}(n^3)$, which is impractical for large datasets. Additionally, a sizable proportion of the data being left-censored creates further bottlenecks, since the likelihood computation now involves an intractable high-dimensional integral of the multivariate Gaussian density. In this article, we tackle these two problems simultaneously by approximating the GP with a Gaussian Markov random field (GMRF) approach that exploits an explicit link between a GP with Matérn correlation function and a GMRF using stochastic partial differential equations (SPDEs). We introduce a GMRF-based measurement error into the model, which alleviates the likelihood computation for the censored data, drastically improving the speed of the model while maintaining admirable accuracy. Our approach demonstrates robustness and substantial computational scalability, compared to state-of-the-art methods for censored spatial responses across various simulation settings. Finally, the fit of this fully Bayesian model to the concentration of PFOS in groundwater available at 24,959 sites across California, where 46.62\% responses are censored, produces prediction surface and uncertainty quantification in real time, thereby substantiating the applicability and scalability of the proposed method. Code for implementation is made available via GitHub.

stat.ME

Pseudo-value regression of clustered multistate current status data with informative cluster sizes

Multistate current status (CS) data presents a more severe form of censoring due to the single observation of study participants transitioning through a sequence of well-defined disease states at random inspection times. Moreover, these data may be clustered within specified groups, and informativeness of the cluster sizes may arise due to the existing latent relationship between the transition outcomes and the cluster sizes. Failure to adjust for this informativeness may lead to a biased inference. Motivated by a clinical study of periodontal disease (PD), we propose an extension of the pseudo-value approach to estimate covariate effects on the state occupation probabilities (SOP) for these clustered multistate CS data with informative cluster or intra-cluster group sizes. In our approach, the proposed pseudo-value technique initially computes marginal estimators of the SOP utilizing nonparametric regression. Next, the estimating equations based on the corresponding pseudo-values are reweighted by functions of the cluster sizes to adjust for informativeness. We perform a variety of simulation studies to study the properties of our pseudo-value regression based on the nonparametric marginal estimators under different scenarios of informativeness. For illustration, the method is applied to the motivating PD dataset, which encapsulates the complex data-generation mechanism.

stat.ME

Efficient Estimation of the Additive Risks Model for Interval-Censored Data

In contrast to the popular Cox model which presents a multiplicative covariate effect specification on the time to event hazards, the semiparametric additive risks model (ARM) offers an attractive additive specification, allowing for direct assessment of the changes or the differences in the hazard function for changing value of the covariates. The ARM is a flexible model, allowing the estimation of both time-independent and time-varying covariates. It has a nonparametric component and a regression component identified by a finite-dimensional parameter. This chapter presents an efficient approach for maximum-likelihood (ML) estimation of the nonparametric and the finite-dimensional components of the model via the minorize-maximize (MM) algorithm for case-II interval-censored data. The operating characteristics of our proposed MM approach are assessed via simulation studies, with illustration on a breast cancer dataset via the R package MMIntAdd. It is expected that the proposed computational approach will not only provide scalability to the ML estimation scenario but may also simplify the computational burden of other complex likelihoods or models.

stat.ME

Alleviating Spatial Confounding in Spatial Frailty Models

Spatial confounding is how is called the confounding between fixed and spatial random effects. It has been widely studied and it gained attention in the past years in the spatial statistics literature, as it may generate unexpected results in modeling. The projection-based approach, also known as restricted models, appears as a good alternative to overcome the spatial confounding in generalized linear mixed models. However, when the support of fixed effects is different from the spatial effect one, this approach can no longer be applied directly. In this work, we introduce a method to alleviate the spatial confounding for the spatial frailty models family. This class of models can incorporate spatially structured effects and it is usual to observe more than one sample unit per area which means that the support of fixed and spatial effects differs. In this case, we introduce a two folded projection-based approach projecting the design matrix to the dimension of the space and then projecting the random effect to the orthogonal space of the new design matrix. To provide fast inference in our analysis we employ the integrated nested Laplace approximation methodology. The method is illustrated with an application with lung and bronchus cancer in California - US that confirms that the methodology efficiency.

stat.ME

Efficient Estimation of Mixture Cure Frailty Model for Clustered Current Status Data

Current status data abounds in the field of epidemiology and public health, where the only observable data for a subject is the random inspection time and the event status at inspection. Motivated by such a current status data from a periodontal study where data are inherently clustered, we propose a unified methodology to analyze such complex data. We allow the time-to-event to follow the semiparametric GOR model with a cure fraction, and develop a unified estimation scheme powered by the EM algorithm. The within-subject correlation is accounted for by a random (frailty) effect, and the non-parametric component of the GOR model is approximated via penalized splines, with a set of knot points that increases with the sample size. Proposed methodology is accompanied by a rigorous asymptotic theory, and the related semiparametric efficiency. The finite sample performance of our model parameters are assessed via simulation studies. Furthermore, the proposed methodology is illustrated via application to the oral health data, accompanied by diagnostic checks to identify influential observations. An easy to use R package CRFCSD is also available for implementation.

stat.ME