SearcharxivSearch

arXiv subjects

Qian Peng

Publications and source records attributed to Qian Peng.

14 recordsLinked to original sources

DynoSys: A Dynamic Systems Framework for Multimodal Integration of Genetic, Environmental, and Neurobiological Signals

Understanding the development of adolescent behavioral and mental health outcomes requires integrating genetic predisposition, environmental exposures, and neurobiological processes over time. Here, we present a unified quantitative framework that models the human body as a dynamic system, where genetic factors form the foundational state, environmental exposures act as time-varying inputs, the brain might serve as a mediation processor, and behavioral phenotypes emerge as system outputs. Using longitudinal data from the Adolescent Brain Cognitive Development (ABCD) Study, we construct harmonized multi-domain representations across six phenotypes: externalizing behavior, internalizing behavior, and four substance use initiation outcomes (alcohol, nicotine, cannabis, and any substance use). We integrate polygenic risk scores (PRS), multi-domain environmental features, and multimodal neuroimaging representations derived through stability selection and dimensionality reduction. Our framework supports both continuous longitudinal modeling and survival-based event modeling through a unified data structure. We further develop interpretable domain-level representations using principal components, weighted risk scores, and cluster-based summaries. These representations enable downstream modeling using survival analysis, state-space models, and machine learning approaches. This work establishes a scalable and interpretable framework for studying how genetic and environmental factors interact over time to shape behavioral outcomes, providing a foundation for identifying modifiable risk factors and informing early intervention strategies.

q-bio.OT

Time-Varying Environmental and Polygenic Predictors of Substance Use Initiation in Youth: A Survival and Causal Modeling Study in the ABCD Cohort

Early substance use initiation is associated with later substance use disorders, but the relative contributions of environmental, behavioral, and genetic factors during adolescence remain unclear. We analyzed 2,366 participants of European genetic ancestry from the Adolescent Brain Cognitive Development Study. Four outcomes were examined through four years of follow-up: initiation of alcohol, nicotine, cannabis, and any substance use. Time-varying Cox models screened predictors, followed by LASSO-selected multivariable Cox models adjusted for age, sex, ancestry principal components, and study site. Marginal structural models evaluated selected predictors. Environmental and behavioral factors were the strongest and most consistent predictors, including youth rule-breaking, impulsivity, parental alcohol-related problems, financial adversity, family context, phone use, negative life events, peer or romantic experiences, sleep, caffeine exposure, and parental monitoring. Polygenic risk scores for alcohol, cannabis, nicotine, and substance use disorder showed weaker and less consistent associations. Although some appeared in screening models, none remained robust after correction, multivariable selection, or causal modeling. In marginal structural models, youth rule-breaking was consistently associated with all four outcomes. Other outcome-specific signals included sensation seeking, lack of planning, parental alcohol-related problems, phone use, sleep duration, and parental monitoring. Early substance use initiation was more strongly associated with dynamic environmental, behavioral, family, and psychosocial factors than with common-variant genetic liability. These findings highlight modifiable developmental pathways and support prevention strategies focused on proximal behavioral and environmental risk contexts.

q-bio.QM

A Joint Survival Modeling and Therapy Knowledge Graph Framework to Characterize Opioid Use Disorder Trajectories

Motivation: Opioid use disorder (OUD) often arises after prescription opioid exposure and follows transitions among onset, remission, and relapse. Linked EHR-survey resources such as the All of Us Research Program enable stage-specific risk modeling and connection to intervention options. Results: We built a multi-stage framework to model time-to-onset, time-to-remission, and time-to-relapse after remission using All of Us EHR and survey data. For each participant we derived longitudinal predictors from clinical conditions and survey concepts, including recent (1/3/12-month) event counts, cumulative exposures, and time since last event. We fit regularized Cox models for each transition and aggregated selection frequencies and hazard ratios to identify a compact set of high-confidence predictors. Pain, mental health, and polysubstance use contributed across stages: chronic pain syndromes, tobacco/nicotine dependence, anxiety and depressive disorders, and cannabis dependence prominently predicted onset and relapse, whereas tobacco dependence during remission and other remission-coded conditions were strongly associated with transition to remission. To support therapeutic prioritization, we constructed a therapy knowledge graph integrating genetic targets, biological pathways, and published evidence to map identified risk factors to candidate treatments in recent OUD studies and clinical guidelines.

q-bio.QM

Spiking Heterogeneous Graph Attention Networks

Real-world graphs or networks are usually heterogeneous, involving multiple types of nodes and relationships. Heterogeneous graph neural networks (HGNNs) can effectively handle these diverse nodes and edges, capturing heterogeneous information within the graph, thus exhibiting outstanding performance. However, most methods of HGNNs usually involve complex structural designs, leading to problems such as high memory usage, long inference time, and extensive consumption of computing resources. These limitations pose certain challenges for the practical application of HGNNs, especially for resource-constrained devices. To mitigate this issue, we propose the Spiking Heterogeneous Graph Attention Networks (SpikingHAN), which incorporates the brain-inspired and energy-saving properties of Spiking Neural Networks (SNNs) into heterogeneous graph learning to reduce the computing cost without compromising the performance. Specifically, SpikingHAN aggregates metapath-based neighbor information using a single-layer graph convolution with shared parameters. It then employs a semantic-level attention mechanism to capture the importance of different meta-paths and performs semantic aggregation. Finally, it encodes the heterogeneous information into a spike sequence through SNNs, simulating bioinformatic processing to derive a binarized 1-bit representation of the heterogeneous graph. Comprehensive experimental results from three real-world heterogeneous graph datasets show that SpikingHAN delivers competitive node classification performance. It achieves this with fewer parameters, quicker inference, reduced memory usage, and lower energy consumption. Code is available at https://github.com/QianPeng369/SpikingHAN.

cs.NE

Hillclimb-Causal Inference: A Data-Driven Approach to Identify Causal Pathways Among Parental Behaviors, Genetic Risk, and Externalizing Behaviors in Children

Motivation: Externalizing behaviors in children, such as aggression, hyperactivity, and defiance, are influenced by complex interplays between genetic predispositions and environmental factors, particularly parental behaviors. Unraveling these intricate causal relationships can benefit from the use of robust data-driven methods. Methods: We developed a method called Hillclimb-Causal Inference, a causal discovery approach that integrates the Hill Climb Search algorithm with a customized Linear Gaussian Bayesian Information Criterion (BIC). This method was applied to data from the Adolescent Brain Cognitive Development (ABCD) Study, which included parental behavior assessments, children's genotypes, and externalizing behavior measures. We performed dimensionality reduction to address multicollinearity among parental behaviors and assessed children's genetic risk for externalizing disorders using polygenic risk scores (PRS), which were computed based on GWAS summary statistics from independent cohorts. Once the causal pathways were identified, we employed structural equation modeling (SEM) to quantify the relationships within the model. Results: We identified prominent causal pathways linking parental behaviors to children's externalizing outcomes. Parental alcohol misuse and broader behavioral issues exhibited notably stronger direct effects (0.33 and 0.20, respectively) compared to children's polygenic risk scores (0.07). Moreover, when considering both direct and indirect paths, parental substance misuse (alcohol, drug, and tobacco) collectively resulted in a total effect exceeding 1.1 on externalizing behaviors. Bootstrap and sensitivity analyses further validated the robustness of these findings.

q-bio.QM

Conformal Prediction with Cellwise Outliers: A Detect-then-Impute Approach

Conformal prediction is a powerful tool for constructing prediction intervals for black-box models, providing a finite sample coverage guarantee for exchangeable data. However, this exchangeability is compromised when some entries of the test feature are contaminated, such as in the case of cellwise outliers. To address this issue, this paper introduces a novel framework called detect-then-impute conformal prediction. This framework first employs an outlier detection procedure on the test feature and then utilizes an imputation method to fill in those cells identified as outliers. To quantify the uncertainty in the processed test feature, we adaptively apply the detection and imputation procedures to the calibration set, thereby constructing exchangeable features for the conformal prediction interval of the test label. We develop two practical algorithms, PDI-CP and JDI-CP, and provide a distribution-free coverage analysis under some commonly used detection and imputation procedures. Notably, JDI-CP achieves a finite sample $1-2\alpha$ coverage guarantee. Numerical experiments on both synthetic and real datasets demonstrate that our proposed algorithms exhibit robust coverage properties and comparable efficiency to the oracle baseline.

stat.ML

Scalable Reinforcement Learning for Virtual Machine Scheduling

Recent advancements in reinforcement learning (RL) have shown promise for optimizing virtual machine scheduling (VMS) in small-scale clusters. The utilization of RL to large-scale cloud computing scenarios remains notably constrained. This paper introduces a scalable RL framework, called Cluster Value Decomposition Reinforcement Learning (CVD-RL), to surmount the scalability hurdles inherent in large-scale VMS. The CVD-RL framework innovatively combines a decomposition operator with a look-ahead operator to adeptly manage representation complexities, while complemented by a Top-$k$ filter operator that refines exploration efficiency. Different from existing approaches limited to clusters of $10$ or fewer physical machines (PMs), CVD-RL extends its applicability to environments encompassing up to $50$ PMs. Furthermore, the CVD-RL framework demonstrates generalization capabilities that surpass contemporary SOTA methodologies across a variety of scenarios in empirical studies. This breakthrough not only showcases the framework's exceptional scalability and performance but also represents a significant leap in the application of RL for VMS within complex, large-scale cloud infrastructures. The code is available at https://anonymous.4open.science/r/marl4sche-D0FE.

cs.LG

V2I-Calib++: A Multi-terminal Spatial Calibration Approach in Urban Intersections for Collaborative Perception

Urban intersections, dense with pedestrian and vehicular traffic and compounded by GPS signal obstructions from high-rise buildings, are among the most challenging areas in urban traffic systems. Traditional single-vehicle intelligence systems often perform poorly in such environments due to a lack of global traffic flow information and the ability to respond to unexpected events. Vehicle-to-Everything (V2X) technology, through real-time communication between vehicles (V2V) and vehicles to infrastructure (V2I), offers a robust solution. However, practical applications still face numerous challenges. Calibration among heterogeneous vehicle and infrastructure endpoints in multi-end LiDAR systems is crucial for ensuring the accuracy and consistency of perception system data. Most existing multi-end calibration methods rely on initial calibration values provided by positioning systems, but the instability of GPS signals due to high buildings in urban canyons poses severe challenges to these methods. To address this issue, this paper proposes a novel multi-end LiDAR system calibration method that does not require positioning priors to determine initial external parameters and meets real-time requirements. Our method introduces an innovative multi-end perception object association technique, utilizing a new Overall Distance metric (oDist) to measure the spatial association between perception objects, and effectively combines global consistency search algorithms with optimal transport theory. By this means, we can extract co-observed targets from object association results for further external parameter computation and optimization. Extensive comparative and ablation experiments conducted on the simulated dataset V2X-Sim and the real dataset DAIR-V2X confirm the effectiveness and efficiency of our method. The code for this method can be accessed at: https://github.com/MassimoQu/v2i-calib.

cs.RO

VMAgent: Scheduling Simulator for Reinforcement Learning

A novel simulator called VMAgent is introduced to help RL researchers better explore new methods, especially for virtual machine scheduling. VMAgent is inspired by practical virtual machine (VM) scheduling tasks and provides an efficient simulation platform that can reflect the real situations of cloud computing. Three scenarios (fading, recovering, and expansion) are concluded from practical cloud computing and corresponds to many reinforcement learning challenges (high dimensional state and action spaces, high non-stationarity, and life-long demand). VMAgent provides flexible configurations for RL researchers to design their customized scheduling environments considering different problem features. From the VM scheduling perspective, VMAgent also helps to explore better learning-based scheduling solutions.

cs.LG

$α$-Satellite: An AI-driven System and Benchmark Datasets for Hierarchical Community-level Risk Assessment to Help Combat COVID-19

The novel coronavirus and its deadly outbreak have posed grand challenges to human society: as of March 26, 2020, there have been 85,377 confirmed cases and 1,293 reported deaths in the United States; and the World Health Organization (WHO) characterized coronavirus disease (COVID-19) - which has infected more than 531,000 people with more than 24,000 deaths in at least 171 countries - a global pandemic. A growing number of areas reporting local sub-national community transmission would represent a significant turn for the worse in the battle against the novel coronavirus, which points to an urgent need for expanded surveillance so we can better understand the spread of COVID-19 and thus better respond with actionable strategies for community mitigation. By advancing capabilities of artificial intelligence (AI) and leveraging the large-scale and real-time data generated from heterogeneous sources (e.g., disease related data from official public health organizations, demographic data, mobility data, and user geneated data from social media), in this work, we propose and develop an AI-driven system (named $α$-Satellite}, as an initial offering, to provide hierarchical community-level risk assessment to assist with the development of strategies for combating the fast evolving COVID-19 pandemic. More specifically, given a specific location (either user input or automatic positioning), the developed system will automatically provide risk indexes associated with it in a hierarchical manner (e.g., state, county, city, specific location) to enable individuals to select appropriate actions for protection while minimizing disruptions to daily life to the extent possible. The developed system and the generated benchmark datasets have been made publicly accessible through our website. The system description and disclaimer are also available in our website.

cs.SI

General Approach To Compute Phosphorescent OLED Efficiency

Phosphorescent organic light-emitting diodes (PhOLEDs) are widely used in the display industry. In PhOLEDs, cyclometalated Ir(III) complexes are the most widespread triplet emitter dopants to attain red, e.g., Ir(piq)3 (piq = 1-phenylisoquinoline), and green, e.g., Ir(ppy)3 (ppy = 2-phenylpyridine), emissions, whereas obtaining operative deep-blue emitters is still one of the major challenges. When designing new emitters, two main characteristics besides colors should be targeted: high photostability and large photoluminescence efficiencies. To date, these are very often optimized experimentally in a trial-and-error manner. Instead, accurate predictive tools would be highly desirable. In this contribution, we present a general approach for computing the photoluminescence lifetimes and efficiencies of Ir(III) complexes by considering all possible competing excited-state deactivation processes and importantly explicitly including the strongly temperature-dependent ones. This approach is based on the combination of state-of-the-art quantum chemical calculations and excited-state decay rate formalism with kinetic modeling, which is shown to be an efficient and reliable approach for a broad palette of Ir(III) complexes, i.e., from yellow/orange to deep-blue emitters.

physics.chem-ph

Prevalent Intrinsic Emission from Nonaromatic Amino Acids and Poly(Amino Acids)

Nonaromatic amino acids are generally believed to be nonemissive, owing to their lack of apparently remarkable conjugation within individual molecules. Here we report the intrinsic visible emission of nonaromatic amino acids and poly(amino acids) in concentrated solutions and solid powders. This unique and widespread luminescent characteristic can be well rationalized by the clustering-triggered emission (CTE) mechanism, namely the clustering of nonconventional chromophores (i.e. amino, carbonyl, and hydroxyl) and subsequent electron cloud overlap with simultaneously conformation rigidification. Such CTE mechanism is further supported by the single crystal structure analysis. Besides prompt fluorescence, room temperature phosphorescence (RTP) are also detected from the solids. Moreover, persistent RTP is observed in the powders of exampled poly(amino acid) of ε-poly-L-lysine (ε-PLL) after ceasing UV irradiation. These results not only illustrate the feasibility of employing the building blocks of nonaromatic amino acids in the exploration of new luminescent biomolecules, but also provide significant implications for the RTP of peptides and proteins at aggregated or crystalline states.

physics.chem-ph

Revised Version of a JCIT Paper-Comparison of Feature Point Extraction Algorithms for Vision Based Autonomous Aerial Refueling

This is a revised version of our paper published in Journal of Convergence Information Technology(JCIT): "Comparison of Feature Point Extraction Algorithms for Vision Based Autonomous Aerial Refueling". We corrected some errors including measurement unit errors, spelling errors and so on. Since the published papers in JCIT are not allowed to be modified, we submit the revised version to arXiv.org to make the paper more rigorous and not to confuse other researchers.

cs.OH

Coordinate-choice independent expression for drift orbit flux and flux-force relation in neoclassical toroidal viscosity theory

A coordinate-choice independent expression does not depend how the magnetic surface is parametrized by (θ,ζ). Flux-force relation in neoclassical toroidal viscosity(NTV) theory has been generalized in a coordinate-choice independent way. The expression for the surface averaged drift orbit flux in 1/νregime is derived without the requirement of straight field line coordinates. The resulted formula is insensitive to how the magnetic surface is parametrized and broadens the cases where flux-force relation can be applied. Construction of straight field line coordinates is avoided when the formula is used for numerical computation.

physics.plasm-ph