SearcharxivSearch

arXiv subjects

Ran Dai

Publications and source records attributed to Ran Dai.

18 recordsLinked to original sources

Dynamic Modeling and Parameter Estimation for Origami Structure Reconfiguration Process

The reconfiguration of origami during the folding and unfolding process is governed through a sequence of panel deformations and hinge orientations. To develop an effective model for representing the reconfiguration process, this paper introduces planar straight-line graphs and a novel consensus protocol for reaching the target origami configuration. The convergence and stability properties of the proposed consensus protocol are subsequently analyzed. Furthermore, to account for aggregate material and structural effects in the proposed consensus-based reconfiguration model, effective parameters embedded in the consensus protocol are identified from trajectory data using a fitting algorithm. Lastly, the effectiveness of the proposed modeling approach is shown using simulations of the two-panel structure and the Kresling origami pattern reconfiguration process.

math.DS

Online multi-layer FDR control

When hypotheses are tested in a stream and real-time decision-making is needed, online sequential hypothesis testing procedures are needed. Furthermore, these hypotheses are commonly partitioned into groups by their nature. For example, the RNA nanocapsules can be partitioned based on therapeutic nucleic acids (siRNAs) being used, as well as the delivery nanocapsules. When selecting effective RNA nanocapsules, simultaneous false discovery rate control at multiple partition levels is needed. In this paper, we develop hypothesis testing procedures which controls false discovery rate (FDR) simultaneously for multiple partitions of hypotheses in an online fashion. We provide rigorous proofs on their FDR or modified FDR (mFDR) control properties and use extensive simulations to demonstrate their performance.

stat.ME

Model-free High Dimensional Mediator Selection with False Discovery Rate Control

There is a challenge in selecting high-dimensional mediators when the mediators have complex correlation structures and interactions. In this work, we frame the high-dimensional mediator selection problem into a series of hypothesis tests with composite nulls, and develop a method to control the false discovery rate (FDR) which has mild assumptions on the mediation model. We show the theoretical guarantee that the proposed method and algorithm achieve FDR control. We present extensive simulation results to demonstrate the power and finite sample performance compared with existing methods. Lastly, we demonstrate the method for analyzing the Alzheimer's Disease Neuroimaging Initiative (ADNI) data, in which the proposed method selects the volume of the hippocampus and amygdala, as well as some other important MRI-derived measures as mediators for the relationship between gender and dementia progression.

stat.ME

Towards an Adaptable and Generalizable Optimization Engine in Decision and Control: A Meta Reinforcement Learning Approach

Sampling-based model predictive control (MPC) has found significant success in optimal control problems with non-smooth system dynamics and cost function. Many machine learning-based works proposed to improve MPC by a) learning or fine-tuning the dynamics/ cost function, or b) learning to optimize for the update of the MPC controllers. For the latter, imitation learning-based optimizers are trained to update the MPC controller by mimicking the expert demonstrations, which, however, are expensive or even unavailable. More significantly, many sequential decision-making problems are in non-stationary environments, requiring that an optimizer should be adaptable and generalizable to update the MPC controller for solving different tasks. To address those issues, we propose to learn an optimizer based on meta-reinforcement learning (RL) to update the controllers. This optimizer does not need expert demonstration and can enable fast adaptation (e.g., few-shots) when it is deployed in unseen control tasks. Experimental results validate the effectiveness of the learned optimizer regarding fast adaptation.

cs.LG

Variable selection with FDR control for noisy data -- an application to screening metabolites that are associated with breast and colorectal cancer

The rapidly expanding field of metabolomics presents an invaluable resource for understanding the associations between metabolites and various diseases. However, the high dimensionality, presence of missing values, and measurement errors associated with metabolomics data can present challenges in developing reliable and reproducible methodologies for disease association studies. Therefore, there is a compelling need to develop robust statistical methods that can navigate these complexities to achieve reliable and reproducible disease association studies. In this paper, we focus on developing such a methodology with an emphasis on controlling the False Discovery Rate during the screening of mutual metabolomic signals for multiple disease outcomes. We illustrate the versatility and performance of this procedure in a variety of scenarios, dealing with missing data and measurement errors. As a specific application of this novel methodology, we target two of the most prevalent cancers among US women: breast cancer and colorectal cancer. By applying our method to the Wome's Health Initiative data, we successfully identify metabolites that are associated with either or both of these cancers, demonstrating the practical utility and potential of our method in identifying consistent risk factors and understanding shared mechanisms between diseases.

stat.ME

Controlling FDR in selecting group-level simultaneous signals from multiple data sources with application to the National Covid Collaborative Cohort data

One challenge in exploratory association studies using observational data is that the associations between the predictors and the outcome are potentially weak and rare, and the candidate predictors have complex correlation structures. False discovery rate (FDR) controlling procedures can provide important statistical guarantees for replicability in predictor identification in exploratory research. In the recently established National COVID Collaborative Cohort (N3C), electronic health record (EHR) data on the same set of candidate predictors are independently collected in multiple different sites, offering opportunities to identify true associations by combining information from different sources. This paper presents a general knockoff-based variable selection algorithm to identify associations from unions of group-level conditional independence tests (simultaneous signals) with exact FDR control guarantees under finite sample settings. This algorithm can work with general regression settings, allowing heterogeneity of both the predictors and the outcomes across multiple data sources. We demonstrate the performance of this method with extensive numerical studies and an application to the N3C data.

stat.ME

FDR Controlled Multiple Testing for Union Null Hypotheses: A Knockoff-based Approach

False discovery rate (FDR) controlling procedures provide important statistical guarantees for the replicability in signal identification based on multiple hypotheses testing. In many fields of study, FDR controlling procedures are used in high-dimensional (HD) analyses to discover features that are truly associated with the outcome. In some recent applications, data on the same set of candidate features are independently collected in multiple different studies. For example, gene expression data are collected at different facilities and with different cohorts, to identify the genetic biomarkers of multiple types of cancers. These studies provide us opportunities to identify signals by considering information from different sources (with potential heterogeneity) jointly. This paper is about how to provide FDR control guarantees for the tests of union null hypotheses of conditional independence. We present a knockoff-based variable selection method (\textit{Simultaneous knockoffs}) to identify mutual signals from multiple independent data sets, providing exact FDR control guarantees under finite sample settings. This method can work with very general model settings and test statistics. We demonstrate the performance of this method with extensive numerical studies and two real data examples.

stat.ME

Post-selection inference on high-dimensional varying-coefficient quantile regression model

Quantile regression has been successfully used to study heterogeneous and heavy-tailed data. Varying-coefficient models are frequently used to capture changes in the effect of input variables on the response as a function of an index or time. In this work, we study high-dimensional varying-coefficient quantile regression models and develop new tools for statistical inference. We focus on development of valid confidence intervals and honest tests for nonparametric coefficients at a fixed time point and quantile, while allowing for a high-dimensional setting where the number of input variables exceeds the sample size. Performing statistical inference in this regime is challenging due to the usage of model selection techniques in estimation. Nevertheless, we can develop valid inferential tools that are applicable to a wide range of data generating processes and do not suffer from biases introduced by model selection. We performed numerical simulations to demonstrate the finite sample performance of our method, and we also illustrated the application with a real data example.

stat.ME

On High Dimensional Covariate Adjustment for Estimating Causal Effects in Randomized Trials with Survival Outcomes

The purpose of this work is to improve the efficiency in estimating the average causal effect (ACE) on the survival scale where right-censoring exists and high-dimensional covariate information is available. We propose new estimators using regularized survival regression and survival random forests (SRF) to make the adjustment for the high dimensional covariates to improve efficiency. We study the behavior of the adjusted estimator under mild assumptions and show theoretical guarantees that the proposed estimators are more efficient than the unadjusted ones asymptotically when using SRF for adjustment. In addition, these adjusted estimators are $\sqrt{n}$- consistent and asymptotically normally distributed. The finite sample behavior of our methods are studied by simulation, and the results are in agreement with the theoretical results. We also illustrate our methods by analyzing the real data from transplant research to identify the relative effectiveness of identical sibling donors compared to unrelated donors with the adjustment of cytogenetic abnormalities.

stat.ME

Convergence guarantee for the sparse monotone single index model

We consider a high-dimensional monotone single index model (hdSIM), which is a semiparametric extension of a high-dimensional generalize linear model (hdGLM), where the link function is unknown, but constrained with monotone and non-decreasing shape. We develop a scalable projection-based iterative approach, the "Sparse Orthogonal Descent Single-Index Model" (SOD-SIM), which alternates between sparse-thresholded orthogonalized "gradient-like" steps and isotonic regression steps to recover the coefficient vector. Our main contribution is that we provide finite sample estimation bounds for both the coefficient vector and the link function in high-dimensional settings under very mild assumptions on the design matrix $\mathbf{X}$, the error term $ε$, and their dependence. The convergence rate for the link function matched the low-dimensional isotonic regression minimax rate up to some poly-log terms ($n^{-1/3}$). The convergence rate for the coefficients is also $n^{-1/3}$ up to some poly-log terms. This method can be applied to many real data problems, including GLMs with misspecified link, classification with mislabeled data, and classification with positive-unlabeled (PU) data. We study the performance of this method via both numerical studies and also an application on a rocker protein sequence data.

math.ST

Angle-Based Sensor Network Localization

This paper studies angle-based sensor network localization (ASNL) in a plane, which is to determine locations of all sensors in a sensor network, given locations of partial sensors (called anchors) and angle measurements obtained in the local coordinate frame of each sensor. Firstly it is shown that a framework with a non-degenerate bilateration ordering must be angle fixable, implying that it can be uniquely determined by angles between edges up to translations, rotations, reflections and uniform scaling. Then ASNL is proved to have a unique solution if and only if the grounded framework is angle fixable and anchors are not all collinear. Subsequently, ASNL is solved in centralized and distributed settings, respectively. The centralized ASNL is formulated as a rank-constrained semi-definite program (SDP) in either a noise-free or a noisy scenario, with a decomposition approach proposed to deal with large-scale ASNL. The distributed protocol for ASNL is designed based on inter-sensor communications. Graphical conditions for equivalence of the formulated rank-constrained SDP and a linear SDP, decomposition of the SDP, as well as the effectiveness of the distributed protocol, are proposed, respectively. Finally, simulation examples demonstrate our theoretical results.

eess.SY

Convex and Non-convex Approaches for Statistical Inference with Class-Conditional Noisy Labels

We study the problem of estimation and testing in logistic regression with class-conditional noise in the observed labels, which has an important implication in the Positive-Unlabeled (PU) learning setting. With the key observation that the label noise problem belongs to a special sub-class of generalized linear models (GLM), we discuss convex and non-convex approaches that address this problem. A non-convex approach based on the maximum likelihood estimation produces an estimator with several optimal properties, but a convex approach has an obvious advantage in optimization. We demonstrate that in the low-dimensional setting, both estimators are consistent and asymptotically normal, where the asymptotic variance of the non-convex estimator is smaller than the convex counterpart. We also quantify the efficiency gap which provides insight into when the two methods are comparable. In the high-dimensional setting, we show that both estimation procedures achieve $\ell_2$-consistency at the minimax optimal $\sqrt{s\log p/n}$ rates under mild conditions. Finally, we propose an inference procedure using a de-biasing approach. We validate our theoretical findings through simulations and a real-data example.

stat.ME

The bias of isotonic regression

We study the bias of the isotonic regression estimator. While there is extensive work characterizing the mean squared error of the isotonic regression estimator, relatively little is known about the bias. In this paper, we provide a sharp characterization, proving that the bias scales as $O(n^{-β/3})$ up to log factors, where $1 \leq β\leq 2$ is the exponent corresponding to H{ö}lder smoothness of the underlying mean. Importantly, this result only requires a strictly monotone mean and that the noise distribution has subexponential tails, without relying on symmetric noise or other restrictive assumptions.

math.ST

Learning-based Hamilton-Jacobi-Bellman Methods for Optimal Control

Many optimal control problems are formulated as two point boundary value problems (TPBVPs) with conditions of optimality derived from the Hamilton-Jacobi-Bellman (HJB) equations. In most cases, it is challenging to solve HJBs due to the difficulty of guessing the adjoint variables. This paper proposes two learning-based approaches to find the initial guess of adjoint variables in real-time, which can be applied to solve general TPBVPs. For cases with database of solutions and corresponding adjoint variables of a TPBVP under varying boundary conditions, a supervised learning method is applied to learn the HJB solutions off-line. After obtaining a trained neural network from supervised learning, we are able to find proper initial adjoint variables for given boundary conditions in real-time. However, when validated solutions of TPBVPs are not available, the reinforcement learning method is applied to solve HJB by constructing a neural network, defining a reward function, and setting appropriate super parameters. The reinforcement learning based HJB method can learn how to find accurate adjoint variables via an updating neural network. Finally, both learning approaches are implemented in classical optimal control problems to verify the effectiveness of the learning based HJB methods.

math.OC

ADMM for combinatorial graph problems

We investigate a class of general combinatorial graph problems, including MAX-CUT and community detection, reformulated as quadratic objectives over nonconvex constraints and solved via the alternating direction method of multipliers (ADMM). We propose two reformulations: one using vector variables and a binary constraint, and the other further reformulating the Burer-Monteiro form for simpler subproblems. Despite the nonconvex constraint, we prove the ADMM iterates converge to a stationary point in both formulations, under mild assumptions. Additionally, recent work suggests that in this latter form, when the matrix factors are wide enough, local optimum with high probability is also the global optimum. To demonstrate the scalability of our algorithm, we include results for MAX-CUT, community detection, and image segmentation benchmark and simulated examples.

math.OC

Instrumental Variable with Competing Risk Model

In this paper, we discuss causal inference on the efficacy of a treatment or medication on a time-to-event outcome with competing risks. Although the treatment group can be randomized, there can be confoundings between the compliance and the outcome. Unmeasured confoundings may exist even after adjustment for measured co- variates. Instrumental variable (IV) methods are commonly used to yield consistent estimations of causal parameters in the presence of unmeasured confoundings. Based on a semi-parametric additive hazard model for the subdistribution hazard, we pro- pose an instrumental variable estimator to yield consistent estimation of efficacy in the presence of unmeasured confoundings for competing risk settings. We derived the asymptotic properties for the proposed estimator. The estimator is shown to be well per- formed under finite sample size according to simulation results. We applied our method to a real transplant data example and showed that the unmeasured confoundings lead to significant bias in the estimation of the effect (about 50% attenuated).

stat.ME

An Iterative Method for Nonconvex Quadratically Constrained Quadratic Programs

This paper examines the nonconvex quadratically constrained quadratic programming (QCQP) problems using an iterative method. One of the existing approaches for solving nonconvex QCQP problems relaxes the rank one constraint on the unknown matrix into semidefinite constraint to obtain the bound on the optimal value without finding the exact solution. By reconsidering the rank one matrix, an iterative rank minimization (IRM) method is proposed to gradually approach the rank one constraint. Each iteration of IRM is formulated as a convex problem with semidefinite constraints. An augmented Lagrangian method, named extended Uzawa algorithm, is developed to solve the subproblem at each iteration of IRM for improved scalability and computational efficiency. Simulation examples are presented using the proposed method and comparative results obtained from the other methods are provided and discussed.

math.OC

The knockoff filter for FDR control in group-sparse and multitask regression

We propose the group knockoff filter, a method for false discovery rate control in a linear regression setting where the features are grouped, and we would like to select a set of relevant groups which have a nonzero effect on the response. By considering the set of true and false discoveries at the group level, this method gains power relative to sparse regression methods. We also apply our method to the multitask regression problem where multiple response variables share similar sparsity patterns across the set of possible features. Empirically, the group knockoff filter successfully controls false discoveries at the group level in both settings, with substantially more discoveries made by leveraging the group structure.

stat.ME