SearcharxivSearch

arXiv subjects

Min Qian

Publications and source records attributed to Min Qian.

At least 19 recordsLinked to original sources

Reluctant Transfer Learning in Penalized Regressions for Individualized Treatment Rules under Effect Heterogeneity

Estimating individualized treatment rules (ITRs) is fundamental to precision medicine, where the goal is to tailor treatment decisions to individual patient characteristics. While numerous methods have been developed for ITR estimation, there is limited research on model updating that accounts for shifted treatment-covariate relationships in the ITR setting. In practice, models trained on source data must be updated for new (target) datasets that exhibit shifts in treatment effects. To address this challenge, we propose a Reluctant Transfer Learning (RTL) framework that enables efficient model adaptation by selectively transferring essential model components (e.g., regression coefficients) from source to target data, without requiring access to individual-level source data. Leveraging the principle of reluctant modeling, the RTL approach incorporates model adjustments only when they improve performance on the target dataset, thereby controlling complexity and enhancing generalizability. Our method supports multi-armed treatment settings, performs variable selection for interpretability, and provides a regret bound for the difference in value of the optimal ITR and that of the estimated ITR. Through simulation studies and an application to a real data example from the Best Apnea Interventions for Research (BestAIR) trial, we demonstrate that RTL outperforms existing alternatives. The proposed framework offers an efficient, practically feasible approach to adaptive treatment decision-making under evolving treatment effect conditions.

stat.ME

Leveraging Two-Phase Data for Improved Prediction of Survival Outcomes with Application to Nasopharyngeal Cancer

Accurate survival predicting models are essential for improving targeted cancer therapies and clinical care among cancer patients. In this article, we investigate and develop a method to improve predictions of survival in cancer by leveraging two-phase data with expert knowledge and prognostic index. Our work is motivated by two-phase data in nasopharyngeal cancer (NPC), where traditional covariates are readily available for all subjects, but the primary viral factor, Human Papillomavirus (HPV), is substantially missing. To address this challenge, we propose an expert guided method that incorporates prognostic index based on the observed covariates and clinical importance of key factors. The proposed method makes efficient use of available data, not simply discarding patients with unknown HPV status. We apply the proposed method and evaluate it against other existing approaches through a series of simulation studies and real data example of NPC patients. Under various settings, the proposed method consistently outperforms competing methods in terms of c-index, calibration slope, and integrated Brier score. By efficiently leveraging two-phase data, the model provides a more accurate and reliable predictive ability of survival models.

stat.ME

Validity of Web-based, Self-directed, NeuroCognitive Performance Test in MCI

Digital cognitive tests offer several potential advantages over established paper-pencil tests but have not yet been fully evaluated for the clinical evaluation of mild cognitive impairment. The NeuroCognitive Performance Test (NCPT) is a web-based, self-directed, modular battery intended for repeated assessments of multiple cognitive domains. Our objective was to examine its relationship with the ADAS-Cog and MMSE as well as with established paper-pencil tests of cognition and daily functioning in MCI. We used Spearman correlations, regressions and principal components analysis followed by a factor analysis (varimax rotated) to examine our objectives. In MCI subjects, the NCPT composite is significantly correlated with both a composite measure of established tests (r=0.78, p<0.0001) as well as with the ADAS-Cog (r=0.55, p<0.0001). Both NCPT and paper-pencil test batteries had a similar factor structure that included a large g component with a high eigenvalue. The correlation for the analogous tests (e.g. Trails A and B, learning memory tests) were significant (p<0.0001). Further, both the NCPT and established tests significantly (p< 0.01) predicted the University of California San Diego Performance-Based Skills Assessment and Functional Activities Questionnaire, measures of daily functioning. The NCPT, a web-based, self-directed, computerized test, shows high concurrent validity with established tests and hence offers promise for use as a research or clinical tool in MCI. Despite limitations such as a relatively small sample, absence of control group and cross-sectional nature, these findings are consistent with the growing literature on the promise of self-directed, web-based cognitive assessments for MCI.

q-bio.NC

Analysis of N-of-1 trials using Bayesian distributed lag model with autocorrelated errors

An N-of-1 trial is a multi-period crossover trial performed in a single individual, with a primary goal to estimate treatment effect on the individual instead of population-level mean responses. As in a conventional crossover trial, it is critical to understand carryover effects of the treatment in an N-of-1 trial, especially when no washout periods between treatment periods are instituted to reduce trial duration. To deal with this issue in situations where high volume of measurements is made during the study, we introduce a novel Bayesian distributed lag model that facilitates the estimation of carryover effects, while accounting for temporal correlations using an autoregressive model. Specifically, we propose a prior variance-covariance structure on the lag coefficients to address collinearity caused by the fact that treatment exposures are typically identical on successive days. A connection between the proposed Bayesian model and penalized regression is noted. Simulation results demonstrate that the proposed model substantially reduces the root mean squared error in the estimation of carryover effects and immediate effects when compared to other existing methods, while being comparable in the estimation of the total effects. We also apply the proposed method to assess the extent of carryover effects of light therapies in relieving depressive symptoms in cancer survivors.

stat.AP

Weakly Supervised-Based Oversampling for High Imbalance and High Dimensionality Data Classification

With the abundance of industrial datasets, imbalanced classification has become a common problem in several application domains. Oversampling is an effective method to solve imbalanced classification. One of the main challenges of the existing oversampling methods is to accurately label the new synthetic samples. Inaccurate labels of the synthetic samples would distort the distribution of the dataset and possibly worsen the classification performance. This paper introduces the idea of weakly supervised learning to handle the inaccurate labeling of synthetic samples caused by traditional oversampling methods. Graph semi-supervised SMOTE is developed to improve the credibility of the synthetic samples' labels. In addition, we propose cost-sensitive neighborhood components analysis for high dimensional datasets and bootstrap based ensemble framework for highly imbalanced datasets. The proposed method has achieved good classification performance on 8 synthetic datasets and 3 real-world datasets, especially for high imbalance and high dimensionality problems. The average performances and robustness are better than the benchmark methods.

cs.LG

Personalized Policy Learning using Longitudinal Mobile Health Data

We address the personalized policy learning problem using longitudinal mobile health application usage data. Personalized policy represents a paradigm shift from developing a single policy that may prescribe personalized decisions by tailoring. Specifically, we aim to develop the best policy, one per user, based on estimating random effects under generalized linear mixed model. With many random effects, we consider new estimation method and penalized objective to circumvent high-dimension integrals for marginal likelihood approximation. We establish consistency and optimality of our method with endogenous app usage. We apply our method to develop personalized push ("prompt") schedules in 294 app users, with a goal to maximize the prompt response rate given past app usage and other contextual factors. We found the best push schedule given the same covariates varied among the users, thus calling for personalized policies. Using the estimated personalized policies would have achieved a mean prompt response rate of 23% in these users at 16 weeks or later: this is a remarkable improvement on the observed rate (11%), while the literature suggests 3%-15% user engagement at 3 months after download. The proposed method compares favorably to existing estimation methods including using the R function "glmer" in a simulation study.

stat.ME

A Sequential Significance Test for Treatment by Covariate Interactions

Due to patient heterogeneity in response to various aspects of any treatment program, biomedical and clinical research is gradually shifting from the traditional "one-size-fits-all" approach to the new paradigm of personalized medicine. An important step in this direction is to identify the treatment by covariate interactions. We consider the setting in which there are potentially a large number of covariates of interest. Although a number of novel machine learning methodologies have been developed in recent years to aid in treatment selection in this setting, few, if any, have adopted formal hypothesis testing procedures. In this article, we present a novel testing procedure based on m-out-of-n bootstrap that can be used to sequentially identify variables that interact with treatment. We study the theoretical properties of the method and show that it is more effective in controlling the type I error rate and achieving a satisfactory power as compared to competing methods, via extensive simulations. Furthermore, the usefulness of the proposed method is illustrated using real data examples, both from a randomized trial and from an observational study.

stat.ME

Linear Mixed Models for Comparing Dynamic Treatment Regimens on a Longitudinal Outcome in Sequentially Randomized Trials

A dynamic treatment regimen (DTR) is a pre-specified sequence of decision rules which maps baseline or time-varying measurements on an individual to a recommended intervention or set of interventions. Sequential multiple assignment randomized trials (SMARTs) represent an important data collection tool for informing the construction of effective DTRs. A common primary aim in a SMART is the marginal mean comparison between two or more of the DTRs embedded in the trial. This manuscript develops a mixed effects modeling and estimation approach for these primary aim comparisons based on a continuous, longitudinal outcome. The method is illustrated using data from a SMART in autism research.

stat.ME

Generalization error for decision problems

In this entry we review the generalization error for classification and single-stage decision problems. We distinguish three alternative definitions of the generalization error which have, at times, been conflated in the statistics literature and show that these definitions need not be equivalent even asymptotically. Because the generalization error is a non-smooth functional of the underlying generative model, standard asymptotic approximations, e.g., the bootstrap or normal approximations, cannot guarantee correct frequentist operating characteristics without modification. We provide simple data-adaptive procedures that can be used to construct asymptotically valid confidence sets for the generalization error. We conclude the entry with a discussion of extensions and related problems.

stat.ME

Stochastic robustness and relative stability of multiple pathways in biological networks

Multiple dynamic pathways always exist in biological networks, but their robustness against internal fluctuations and relative stability have not been well recognized and carefully analyzed yet. Here we try to address these issues through an illustrative example, namely the Siah-1/beta-catenin/p14/19 ARF loop of protein p53 dynamics. Its deterministic Boolean network model predicts that two parallel pathways with comparable magnitudes of attractive basins should exist after the protein p53 is activated when a cell becomes harmfully disturbed. Once the low but non-neglectable intrinsic fluctuations are incorporated into the model, we show that a phase transition phenomenon is emerged: in one parameter region the probability weights of the normal pathway, reported in experimental literature, are comparable with the other pathway which is seemingly abnormal with the unknown functions, whereas, in some other parameter regions, the probability weight of the abnormal pathway can even dominate and become globally attractive. The theory of exponentially perturbed Markov chains is applied and further generalized in order to quantitatively explain such a phase transition phenomenon, in which the nonequilibrium "activation energy barriers" along each transiting trajectory between the parallel pathways and the number of "optimal transition paths" play a central part. Our theory can also determine how the transition time and the number of optimal transition paths between the parallel pathways depend on each interaction's strength, and help to identify those possibly more crucial interactions in the biological network.

q-bio.MN

Stochastic Dynamics of Electrical Membrane with Voltage-Dependent Ion Channel Fluctuations

Brownian ratchet like stochastic theory for the electrochemical membrane system of Hodgkin-Huxley (HH) is developed. The system is characterized by a continuous variable $Q_m(t)$, representing mobile membrane charge density, and a discrete variable $K_t$ representing ion channel conformational dynamics. A Nernst-Planck-Nyquist-Johnson type equilibrium is obtained when multiple conducting ions have a common reversal potential. Detailed balance yields a previously unknown relation between the channel switching rates and membrane capacitance, bypassing Eyring-type explicit treatment of gating charge kinetics. From a molecular structural standpoint, membrane charge $Q_m$ is a more natural dynamic variable than potential $V_m$; our formalism treats $Q_m$-dependent conformational transition rates $λ_{ij}$ as intrinsic parameters. Therefore in principle, $λ_{ij}$ vs. $V_m$ is experimental protocol dependent,e.g., different from voltage or charge clamping measurements. For constant membrane capacitance per unit area $C_m$ and neglecting membrane potential induced by gating charges, $V_m=Q_m/C_m$, and HH's formalism is recovered. The presence of two types of ions, with different channels and reversal potentials, gives rise to a nonequilibrium steady state with positive entropy production $e_p$. For rapidly fluctuating channels, an expression for $e_p$ is obtained.

physics.bio-ph

Statistical Inference in Dynamic Treatment Regimes

Dynamic treatment regimes are of growing interest across the clinical sciences as these regimes provide one way to operationalize and thus inform sequential personalized clinical decision making. A dynamic treatment regime is a sequence of decision rules, with a decision rule per stage of clinical intervention; each decision rule maps up-to-date patient information to a recommended treatment. We briefly review a variety of approaches for using data to construct the decision rules. We then review an interesting challenge, that of nonregularity that often arises in this area. By nonregularity, we mean the parameters indexing the optimal dynamic treatment regime are nonsmooth functionals of the underlying generative distribution. A consequence is that no regular or asymptotically unbiased estimator of these parameters exists. Nonregularity arises in inference for parameters in the optimal dynamic treatment regime; we illustrate the effect of nonregularity on asymptotic bias and via sensitivity of asymptotic, limiting, distributions to local perturbations. We propose and evaluate a locally consistent Adaptive Confidence Interval (ACI) for the parameters of the optimal dynamic treatment regime. We use data from the Adaptive Interventions for Children with ADHD study as an illustrative example. We conclude by highlighting and discussing emerging theoretical problems in this area.

stat.ME

Circular Stochastic Fluctuations in SIS Epidemics with Heterogeneous Contacts Among Sub-populations

The conceptual difference between equilibrium and non-equilibrium steady state (NESS) is well established in physics and chemistry. This distinction, however, is not widely appreciated in dynamical descriptions of biological populations in terms of differential equations in which fixed point, steady state, and equilibrium are all synonymous. We study NESS in a stochastic SIS (susceptible-infectious-susceptible) system with heterogeneous individuals in their contact behavior represented in terms of subgroups. In the infinite population limit, the stochastic dynamics yields a system of deterministic evolution equations for population densities; and for very large but finite system a diffusion process is obtained. We report the emergence of a circular dynamics in the diffusion process, with an intrinsic frequency, near the endemic steady state. The endemic steady state is represented by a stable node in the deterministic dynamics; As a NESS phenomenon, the circular motion is caused by the intrinsic heterogeneity within the subgroups, leading to a broken symmetry and time irreversibility.

q-bio.PE

Performance guarantees for individualized treatment rules

Because many illnesses show heterogeneous response to treatment, there is increasing interest in individualizing treatment to patients [Arch. Gen. Psychiatry 66 (2009) 128--133]. An individualized treatment rule is a decision rule that recommends treatment according to patient characteristics. We consider the use of clinical trial data in the construction of an individualized treatment rule leading to highest mean response. This is a difficult computational problem because the objective function is the expectation of a weighted indicator function that is nonconcave in the parameters. Furthermore, there are frequently many pretreatment variables that may or may not be useful in constructing an optimal individualized treatment rule, yet cost and interpretability considerations imply that only a few variables should be used by the individualized treatment rule. To address these challenges, we consider estimation based on $l_1$-penalized least squares. This approach is justified via a finite sample upper bound on the difference between the mean response due to the estimated individualized treatment rule and the mean response due to the optimal individualized treatment rule.

math.ST

Sensitivity Amplification in the Phosphorylation-Dephosphorylation Cycle: Nonequilibrium steady states, chemical master equation and temporal cooperativity

A new type of cooperativity termed temporal cooperativity [Biophys. Chem. 105 585-593 (2003), Annu. Rev. Phys. Chem. 58 113-142 (2007)], emerges in the signal transduction module of phosphorylation-dephosphorylation cycle (PdPC). It utilizes multiple kinetic cycles in time, in contrast to allosteric cooperativity that utilizes multiple subunits in a protein. In the present paper, we thoroughly investigate both the deterministic (microscopic) and stochastic (mesoscopic) models, and focus on the identification of the source of temporal cooperativity via comparing with allosteric cooperativity. A thermodynamic analysis confirms again the claim that the chemical equilibrium state exists if and only if the phosphorylation potential $\triangle G=0$, in which case the amplification of sensitivity is completely abolished. Then we provide comprehensive theoretical and numerical analysis with the first-order and zero-order assumptions in phosphorylation-dephosphorylation cycle respectively. Furthermore, it is interestingly found that the underlying mathematics of temporal cooperativity and allosteric cooperativity are equivalent, and both of them can be expressed by "dissociation constants", which also characterizes the essential differences between the simple and ultrasensitive PdPC switches. Nevertheless, the degree of allosteric cooperativity is restricted by the total number of sites in a single enzyme molecule which can not be freely regulated, while temporal cooperativity is only restricted by the total number of molecules of the target protein which can be regulated in a wide range and gives rise to the ultrasensitivity phenomenon.

physics.chem-ph

Boolean Network Approach to Negative Feedback Loops of the p53 Pathways: Synchronized Dynamics and Stochastic Limit Cycles

Deterministic and stochastic Boolean network models are build for the dynamics of negative feedback loops of the p53 pathways. It is shown that the main function of the negative feedback in the p53 pathways is to keep p53 at a low steady state level, and each sequence of protein states in the negative feedback loops, is globally attracted to a closed cycle of the p53 dynamics after being perturbed by outside signal (e.g. DNA damage). Our theoretical and numerical studies show that both the biological stationary state and the biological oscillation after being perturbed are stable for a wide range of noise level. Applying the mathematical circulation theory of Markov chains, we investigate their stochastic synchronized dynamics and by comparing the network dynamics of the stochastic model with its corresponding deterministic network counterpart, a dominant circulation in the stochastic model is the natural generalization of the deterministic limit cycle in the deterministic system. Moreover, the period of the main peak in the power spectrum, which is in common use to characterize the synchronized dynamics, perfectly corresponds to the number of states in the main cycle with dominant circulation. Such a large separation in the magnitude of the circulations, between a dominant, main cycle and the rest, gives rise to the stochastic synchronization phenomenon.

q-bio.MN

Synchronized Dynamics and Nonequilibrium Steady States in a Stochastic Yeast Cell-Cycle Network

Applying the mathematical circulation theory of Markov chains, we investigate the synchronized stochastic dynamics of a discrete network model of yeast cell-cycle regulation where stochasticity has been kept rather than being averaged out. By comparing the network dynamics of the stochastic model with its corresponding deterministic network counterpart, we show that the synchronized dynamics can be soundly characterized by a dominant circulation in the stochastic model, which is the natural generalization of the deterministic limit cycle in the deterministic system. Moreover, the period of the main peak in the power spectrum, which is in common use to characterize the synchronized dynamics, perfectly corresponds to the number of states in the main cycle with dominant circulation. Such a large separation in the magnitude of the circulations, between a dominant, main cycle and the rest, gives rise to the stochastic synchronization phenomenon.

q-bio.MN

Holomorphic harmonic analysis on complex reductive groups

We define the holomorphic Fourier transform of holomorphic functions on complex reductive groups, prove some properties like the Fourier inversion formula, and give some applications. The definition of the holomorphic Fourier transform makes use of the notion of $K$-admissible measures. We prove that $K$-admissible measures are abundant, and the definition of holomorphic Fourier transform is independent of the choice of $K$-admissible measures.

math.GR