Searcharxiv⌕ Search

arXiv subjects

Nilanjana Laha

Publications and source records attributed to Nilanjana Laha.

12 recordsLinked to original sources

On Fisher Consistency of Surrogate Losses for Optimal Dynamic Treatment Regimes with Multiple Categorical Treatments per Stage

Patients with chronic diseases often receive treatments at multiple time points, or stages. Our goal is to learn the optimal dynamic treatment regime (DTR) from longitudinal patient data. When both the number of stages and the number of treatment levels per stage are arbitrary, estimating the optimal DTR reduces to a sequential, weighted, multiclass classification problem (Kosorok and Laber, 2019). In this paper, we aim to solve this classification problem simultaneously across all stages using Fisher consistent surrogate losses. Although computationally feasible Fisher consistent surrogates exist in special cases, e.g., the binary treatment setting, a unified theory of Fisher consistency remains largely unexplored. We establish necessary and sufficient conditions for DTR Fisher consistency within the class of non-negative, stagewise separable surrogate losses. To our knowledge, this is the first result in the DTR literature to provide necessary conditions for Fisher consistency within a non-trivial surrogate class. Furthermore, we show that many convex surrogate losses fail to be Fisher consistent for the DTR classification problem, and we formally establish this inconsistency for smooth, permutation equivariant, and relative-margin-based convex losses. Building on this, we propose SDSS (Simultaneous Direct Search with Surrogates), which uses smooth, non-concave surrogate losses to learn the optimal DTR. We develop a computationally efficient, gradient-based algorithm for SDSS. When the optimization error is small, we establish a sharp upper bound on SDSS's regret decay rate. We evaluate the numerical performance of SDSS through simulations and demonstrate its real-world applicability by estimating optimal fluid resuscitation strategies for severe septic patients using electronic health record data.

math.ST↗

Revisiting the two-sample location shift model with a log-concavity assumption

In this paper, we consider the two-sample location shift model, a classic semiparametric model introduced by Stein (1956). This model is known for its adaptive nature, enabling nonparametric estimation with full parametric efficiency. Existing nonparametric estimators of the location shift often depend on external tuning parameters, which restricts their practical applicability (Van der Vaart and Wellner, 2021). We demonstrate that introducing an additional assumption of log-concavity on the underlying density can alleviate the need for tuning parameters. We propose a one step estimator for location shift estimation, utilizing log-concave density estimation techniques to facilitate tuning-free estimation of the efficient influence function. While we employ a truncated version of the one step estimator for theoretical adaptivity, our simulations indicate that the one step estimators perform best with zero truncation, eliminating the need for tuning during practical implementation.

math.ST↗

Adaptive estimation in symmetric location model under log-concavity constraint

We revisit the problem of estimating the center of symmetry $θ$ of an unknown symmetric density $f$. Although stone (1975), Eden (1970), and Sacks (1975) constructed adaptive estimators of $θ$ in this model, their estimators depend on external tuning parameters. In an effort to reduce the burden of tuning parameters, we impose an additional restriction of log-concavity on $f$. We construct truncated one-step estimators which are adaptive under the log-concavity assumption. Our simulations suggest that the untruncated version of the one step estimator, which is tuning parameter free, is also asymptotically efficient. We also study the maximum likelihood estimator (MLE) of $θ$ in the shape-restricted model.

math.ST↗

Finding the Optimal Dynamic Treatment Regime Using Smooth Fisher Consistent Surrogate Loss

Large health care data repositories such as electronic health records (EHR) open new opportunities to derive individualized treatment strategies for complicated diseases such as sepsis. In this paper, we consider the problem of estimating sequential treatment rules tailored to a patient's individual characteristics, often referred to as dynamic treatment regimes (DTRs). Our main objective is to find the optimal DTR that maximizes a discontinuous value function through direct maximization of Fisher consistent surrogate loss functions. In this regard, we demonstrate that a large class of concave surrogates fails to be Fisher consistent -- a behavior that differs from the classical binary classification problems. We further characterize a non-concave family of Fisher consistent smooth surrogate functions, which is amenable to gradient-descent type optimization algorithms. Compared to the existing direct search approach under the support vector machine framework (Zhao et al., 2015), our proposed DTR estimation via surrogate loss optimization (DTRESLO) method is more computationally scalable to large sample sizes and allows for broader functional classes for treatment policies. We establish theoretical properties for our proposed DTR estimator and obtain a sharp upper bound on the regret corresponding to our DTRESLO method. The finite sample performance of our proposed estimator is evaluated through extensive simulations. Finally, we illustrate the working principles and benefits of our method for estimating an optimal DTR for treating sepsis using EHR data from sepsis patients admitted to intensive care units.

math.ST↗

Improved inference for vaccine-induced immune responses via shape-constrained methods

We study the performance of shape-constrained methods for evaluating immune response profiles from early-phase vaccine trials. The motivating problem for this work involves quantifying and comparing the IgG binding immune responses to the first and second variable loops (V1V2 region) arising in HVTN 097 and HVTN 100 HIV vaccine trials. We consider unimodal and log-concave shape-constrained methods to compare the immune profiles of the two vaccines, which is reasonable because the data support that the underlying densities of the immune responses could have these shapes. To this end, we develop novel shape-constrained tests of stochastic dominance and shape-constrained plug-in estimators of the Hellinger distance between two densities. Our techniques are either tuning parameter free, or rely on only one tuning parameter, but their performance is either better (the tests of stochastic dominance) or comparable with the nonparametric methods (the estimators of Hellinger distance). The minimal dependence on tuning parameters is especially desirable in clinical contexts where analyses must be prespecified and reproducible. Our methods are supported by theoretical results and simulation studies.

stat.ME↗

On Support Recovery with Sparse CCA: Information Theoretic and Computational Limits

In this paper we consider asymptotically exact support recovery in the context of high dimensional and sparse Canonical Correlation Analysis (CCA). Our main results describe four regimes of interest based on information theoretic and computational considerations. In regimes of "low" sparsity we describe a simple, general, and computationally easy method for support recovery, whereas in a regime of "high" sparsity, it turns out that support recovery is information theoretically impossible. For the sake of information theoretic lower bounds, our results also demonstrate a non-trivial requirement on the "minimal" size of the non-zero elements of the canonical vectors that is required for asymptotically consistent support recovery. Subsequently, the regime of "moderate" sparsity is further divided into two sub-regimes. In the lower of the two sparsity regimes, using a sharp analysis of a coordinate thresholding (Deshpande and Montanari, 2014) type method, we show that polynomial time support recovery is possible. In contrast, in the higher end of the moderate sparsity regime, appealing to the "Low Degree Polynomial" Conjecture (Kunisky et al., 2019), we provide evidence that polynomial time support recovery methods are inconsistent. Finally, we carry out numerical experiments to compare the efficacy of various methods discussed.

math.ST↗

On the Existence of Universal Lottery Tickets

The lottery ticket hypothesis conjectures the existence of sparse subnetworks of large randomly initialized deep neural networks that can be successfully trained in isolation. Recent work has experimentally observed that some of these tickets can be practically reused across a variety of tasks, hinting at some form of universality. We formalize this concept and theoretically prove that not only do such universal tickets exist but they also do not require further training. Our proofs introduce a couple of technical innovations related to pruning for strong lottery tickets, including extensions of subset sum results and a strategy to leverage higher amounts of depth. Our explicit sparse constructions of universal function families might be of independent interest, as they highlight representational benefits induced by univariate convolutional architectures.

cs.LG↗

On Statistical Inference with High Dimensional Sparse CCA

We consider asymptotically exact inference on the leading canonical correlation directions and strengths between two high dimensional vectors under sparsity restrictions. In this regard, our main contribution is the development of a loss function, based on which, one can operationalize a one-step bias-correction on reasonable initial estimators. Our analytic results in this regard are adaptive over suitable structural restrictions of the high dimensional nuisance parameters, which, in this set-up, correspond to the covariance matrices of the variables of interest. We further supplement the theoretical guarantees behind our procedures with extensive numerical studies.

math.ST↗

Semi-Supervised Off Policy Reinforcement Learning

Reinforcement learning (RL) has shown great success in estimating sequential treatment strategies which take into account patient heterogeneity. However, health-outcome information, which is used as the reward for reinforcement learning methods, is often not well coded but rather embedded in clinical notes. Extracting precise outcome information is a resource intensive task, so most of the available well-annotated cohorts are small. To address this issue, we propose a semi-supervised learning (SSL) approach that efficiently leverages a small sized labeled data with true outcome observed, and a large unlabeled data with outcome surrogates. In particular, we propose a semi-supervised, efficient approach to Q-learning and doubly robust off policy value estimation. Generalizing SSL to sequential treatment regimes brings interesting challenges: 1) Feature distribution for Q-learning is unknown as it includes previous outcomes. 2) The surrogate variables we leverage in the modified SSL framework are predictive of the outcome but not informative to the optimal policy or value function. We provide theoretical results for our Q-function and value function estimators to understand to what degree efficiency can be gained from SSL. Our method is at least as efficient as the supervised approach, and moreover safe as it robust to mis-specification of the imputation models.

cs.LG↗

Bi-$s^*$-Concave Distributions

We introduce new shape-constrained classes of distribution functions on R, the bi-$s^*$-concave classes. In parallel to results of Dümbgen, Kolesnyk, and Wilke (2017) for what they called the class of bi-log-concave distribution functions, we show that every $s$-concave density $f$ has a bi-$s^*$-concave distribution function $F$ for $s^*\leq s/(s+1)$. Confidence bands building on existing nonparametric bands, but accounting for the shape constraint of bi-$s^*$-concavity, are also considered. The new bands extend those developed by Dümbgen et al. (2017) for the constraint of bi-log-concavity. We also make connections between bi-$s^*$-concavity and finiteness of the Csörgő - Révész constant of $F$ which plays an important role in the theory of quantile processes.

math.ST↗

Location estimation for symmetric log-concave densities

We revisit the problem of estimating the center of symmetry $θ$ of an unknown symmetric density $f$. Although Stone (1975), Van Eden (1970), and Sacks (1975) constructed adaptive estimators of $θ$ in this model, their estimators depend on tuning parameters. In an effort to circumvent the dependence on tuning parameters, we impose an additional assumption of log-concavity on $f$. We show that in this shape-restricted model, the maximum likelihood estimator (MLE) of $θ$ exists. We also study some truncated one-step estimators and show that they are $\sqrt{n}-$consistent, and nearly achieve the asymptotic efficiency bound. We also show that the rate of convergence for the MLE is $O_p(n^{-2/5})$. Furthermore, we show that our estimators are robust with respect to the violation of the log-concavity assumption. In fact, we show that the one step estimators are still $\sqrt{n}$-consistent under some mild conditions. These analytical conclusions are supported by simulation studies.

math.ST↗

Bi-$s^*$-concave distributions

We introduce a new shape-constrained class of distribution functions on R, the bi-$s^*$-concave class. In parallel to results of Dümbgen, Kolesnyk, and Wilke (2017) for what they called the class of bi-log-concave distribution functions, we show that every s-concave density f has a bi-$s^*$-concave distribution function $F$ and that every bi-$s^*$-concave distribution function satisfies $γ(F) \le 1/(1+s)$ where finiteness of $$ γ(F) \equiv \sup_{x} F(x) (1-F(x)) \frac{| f' (x)|}{f^2 (x)}, $$ the Csörgő - Révész constant of F, plays an important role in the theory of quantile processes on $R$.

math.ST↗