SearcharxivSearch

arXiv subjects

Kai Kang

Publications and source records attributed to Kai Kang.

At least 19 recordsLinked to original sources

Bioinfoysis Technical Report

Large language model agents have shown promise in bioinformatics, but most existing systems focus primarily on producing final answers, treating planning, tool use, and code execution as transient interactions. This design is poorly suited to long-horizon bioinformatics tasks, where conclusions must remain connected to the data, computations, and intermediate evidence that support them. We introduce \textbf{Bioinfoysis}, a multi-agent harness that represents each request as a persistent, artifact-grounded analysis run. Bioinfoysis combines global planning with step-wise, evidence-driven replanning: the planner maintains an executable checklist and revises pending steps using structured handoffs returned after each worker execution. These handoffs bind intermediate results to their responsible agent, checklist step, and plan generation, preventing stale evidence from being silently reused after replanning. A controlled runtime validates generated scripts, tables, and figures before they are used in downstream analysis or reporting, while role-specific context, persistent memory, and governed bioinformatics skills support reliable execution over long analysis trajectories. We evaluate Bioinfoysis on BixBench and two question-answering tracks of LAB-Bench 2. On BixBench, Bioinfoysis achieves state-of-the-art accuracy of 82.4\%. Across four underlying language models, Bioinfoysis increases average accuracy from 27.81\% to 64.13\% on SeqQA2 and from 3.13\% to 31.25\% on DbQA2. These results demonstrate that reliable bioinformatics automation depends not only on model capability, but also on the harness that governs planning, execution, memory, and evidence flow. We hope that the emergence of Bioinfoysis will play a driving and leading role in the development of the bioinformatics community. Our demo website can be seen in https://report.bioinfoysis.com/.

cs.AI

Endogeneity-Aware Cognitive Diagnostic Model for Multidomain Ordinal Assessments

Multidomain assessment batteries generate ordinal item responses that are often summarized through latent attribute profiles. Conventional cognitive diagnostic models (CDMs) provide interpretable measurement models for such profiles, but they typically do not represent directed dependence among latent attributes from distinct domains. We propose an endogeneity-aware cognitive diagnostic model (EACDM) for multivariate ordinal assessments. The model combines a block-structured diagnostic measurement component, in which item groups are linked to domain-specific binary attributes through a block-diagonal Q-matrix, with a logistic structural component, in which one attribute block is regressed on another block and subject-level covariates while accounting for latent classification uncertainty. This formulation yields a parsimonious framework for studying endogenous relationships among diagnostic attributes without collapsing domain-specific measurement structure. We establish identifiability conditions for the Q-matrix, effective loadings, latent-profile probabilities, and structural coefficients, and develop a Markov chain Monte Carlo algorithm for joint estimation of the measurement and structural components. Simulation studies demonstrate accurate recovery of item parameters, latent structures, and structural coefficients for the proposed EACDM, whereas conventional CDMs can fail to recover the ground truth when endogeneity is present. We apply the proposed method to Parkinson's disease data to examine how non-motor latent traits relate to motor impairment profiles.

stat.ME

Semiparametric Functional Multistate Modeling of Alzheimer's Disease Progression with Imaging Biomarkers

Medical imaging provides rich information for predicting Alzheimer's disease progression, but existing imaging-based methods typically focus on a single survival endpoint and treat transition times as exactly observed or right-censored. Motivated by the Alzheimer's Disease Neuroimaging Initiative (ADNI), we develop a predictive framework that represents disease progression as an intermittently observed multistate process with interval-censored transition times and predicts future progression from any current disease state. We incorporate imaging biomarkers as functional covariates in a semiparametric proportional intensity model and combine functional principal component analysis with nonparametric maximum pseudo likelihood estimation. We further develop a profile score test for assessing the overall association between the imaging covariate and the multistate process. We establish the asymptotic properties of the proposed estimators and test statistic, and simulation studies demonstrate satisfactory finite-sample performance. In the ADNI application, baseline lateral ventricular morphology is strongly associated with Alzheimer's disease progression. The proposed functional multistate model also achieves the best overall predictive performance among the competing methods.

stat.AP

Dispatch-Embedded Long-Term Tail Risk Assessment and Mitigation via CVaR for Renewable Power Systems

Renewable energy (RE) generation exhibits pronounced seasonality and variability, and neglecting these features can lead to significant underestimation of long-term power system risks in power supply. While long-term dispatch strategies are essential for evaluating and mitigating tail risks, they are often excluded from existing models due to their complexity. This paper proposes a long-term tail risk assessment and mitigation framework for renewable power systems, explicitly embedding dispatch strategies. A representative scenario generation method is designed, combining multi-timescale Copula modeling to capture RE's long-range variability and correlation. Building on these scenarios, an evolution-based risk assessment model is established, where Conditional Value-at-Risk (CVaR) is employed as a robust metric to quantify tail risks. Finally, a controlled evolution-based risk mitigation scheme is introduced to refine long-term dispatch strategies for mitigating tail risks. Case studies on a modified IEEE-39 bus system incorporating real-world data substantiate the efficacy of the proposed method.

eess.SY

KoopmanFlow: Spectrally Decoupled Generative Control Policy via Koopman Structural Bias

Generative Control Policies (GCPs) show immense promise in robotic manipulation but struggle to simultaneously model stable global motions and high-frequency local corrections. While modern architectures extract multi-scale spatial features, their underlying Probability Flow ODEs apply a uniform temporal integration schedule. Compressed to a single step for real-time Receding Horizon Control (RHC), uniform ODE solvers mathematically smooth over sparse, high-frequency transients entangled within low-frequency steady states. To decouple these dynamics without accumulating pipelined errors, we introduce KoopmanFlow, a parameter-efficient generative policy guided by a Koopman-inspired structural inductive bias. Operating in a unified multimodal latent space with visual context, KoopmanFlow bifurcates generation at the terminal stage. Because visual conditioning occurs before spectral decomposition, both branches are visually guided yet temporally specialized. A macroscopic branch anchors slow-varying trajectories via single-step Consistency Training, while a transient branch uses Flow Matching to isolate high-frequency residuals stimulated by sudden visual cues (e.g., contacts or occlusions). Guided by an explicit spectral prior and optimized via a novel asymmetric consistency objective, KoopmanFlow establishes a fused co-training mechanism. This allows the variant branch to absorb localized dynamics without multi-stage error accumulation. Extensive experiments show KoopmanFlow significantly outperforms state-of-the-art baselines in contact-rich tasks requiring agile disturbance rejection. By trading a surplus latency buffer for a richer structural prior, KoopmanFlow achieves superior control fidelity and parameter efficiency within real-time deployment limits.

cs.RO

Controlled Evolution-Based Day-Ahead Robust Dispatch Considering Frequency Security with Frequency Regulation Loads and Curtailable Loads

With the extensive integration of volatile and uncertain renewable energy, power systems face significant challenges in primary frequency regulation due to instantaneous power fluctuations. However, the maximum frequency deviation constraint is inherently non-convex, and commonly used two-stage dispatch methods overlook causality, potentially resulting in infeasible day-ahead decisions. This paper presents a controlled evolution-based day-ahead robust dispatch method to address these issues. First, we suggest the convex relaxation technique to transform the maximum frequency deviation constraint to facilitate optimization. Then, an evolution-based robust dispatch framework is introduced to align day-ahead decisions with intraday strategies, ensuring both frequency security and power supply reliability. Additionally, a novel controlled evolution-based algorithm is developed to solve this framework efficiently. Case studies on a modified IEEE 14-bus system demonstrate the superiority of the proposed method in enhancing frequency security and system reliability.

eess.SY

Tritiated methane reduction in the PandaX-4T experiment via purge and cryogenic distillation processes

Tritium from tritiated methane (CH$_3$T) calibration is a significant impurity that restricts the sensitivity of the PandaX-4T dark matter detection experiment in the low-energy region. The CH$_3$T removal is essential for PandaX-4T and other liquid xenon dark matter direct detection experiments, as CH$_3$T serves as a critical component for low-energy calibration. To eliminate CH$_3$T, the xenon in the detector is suitably recuperated, leaving 1.8 bar of xenon gas inside, and the detector is flushed with heated xenon gas. Concurrently, leveraging the lower boiling point of methane relative to xenon, the PandaX-4T cryogenic distillation system is effectively utilized to extract CH$_3$T from xenon after optimizing the operational parameters. Following the commissioning run, 5.7 tons of xenon are purified via the distillation method. Recent data indicate that the CH$_3$T concentration reduces from $3.6\times10^{-24}$ mol/mol to $5.9\times10^{-25}$ mol/mol, demonstrating that gas purging and distillation are effective in removing CH$_3$T, even at concentrations on the order of $10^{-24}$ mol/mol.

physics.ins-det

Heterogeneous immune recovery after viral response through a dynamical model of feedback-driven persistence and clearance

Viral infections trigger complex immune responses with heterogeneous outcomes shaped by nonlinear feedbacks. An ordinary differential equation model is developed to investigate immune response dynamics during viral infection, incorporating six modules: viral load, innate immunity, cellular immunity, humoral immunity, immune suppression, and IL-6 levels. Bifurcation analysis reveals that under continuous viral exposure, when viral clearance rate and intrinsic viral death rate satisfy specific conditions, the system exhibits up to five stable equilibria. This indicates that different health and disease states may coexist depending on initial conditions, while severe inflammation mainly arises from strong activation of cellular immunity, highlighting the complexity of immune responses. Simulations of finite-time viral exposure demonstrate multi-timescale recovery characteristics: viral load and IL-6 levels decline rapidly, whereas humoral immune activation and immunosuppression show delayed and sustained patterns. Furthermore, analysis of infectious period and disease duration also indicates that during transition from early acute response to chronic disease, viral replication rate plays a critical role, while immune response intensity is sensitive to both viral clearance and immune self-activation. Subsystem analysis identifies the three-component subsystem of viral load, innate immunity, and cellular immunity as core drivers of bistability and oscillations, while humoral immunity, immune suppression, and IL-6 primarily modulate response amplitude and timing. This work establishes a theoretical framework for analyzing immune response and chronic risks through feedback dynamical modelling, providing insights for intervention strategies.

q-bio.QM

Evaluation and LLM-Guided Learning of ICD Coding Rationales

ICD coding is the process of mapping unstructured text from Electronic Health Records (EHRs) to standardised codes defined by the International Classification of Diseases (ICD) system. In order to promote trust and transparency, existing explorations on the explainability of ICD coding models primarily rely on attention-based rationales and qualitative assessments conducted by physicians, yet lack a systematic evaluation across diverse types of rationales using consistent criteria and high-quality rationale-annotated datasets specifically designed for the ICD coding task. Moreover, dedicated methods explicitly trained to generate plausible rationales remain scarce. In this work, we present evaluations of the explainability of rationales in ICD coding, focusing on two fundamental dimensions: faithfulness and plausibility -- in short how rationales influence model decisions and how convincing humans find them. For plausibility, we construct a novel, multi-granular rationale-annotated ICD coding dataset, based on the MIMIC-IV database and the updated ICD-10 coding system. We conduct a comprehensive evaluation across three types of ICD coding rationales: entity-level mentions automatically constructed via entity linking, LLM-generated rationales, and rationales based on attention scores of ICD coding models. Building upon the strong plausibility exhibited by LLM-generated rationales, we further leverage them as distant supervision signals to develop rationale learning methods. Additionally, by prompting the LLM with few-shot human-annotated examples from our dataset, we achieve notable improvements in the plausibility of rationale generation in both the teacher LLM and the student rationale learning models.

cs.AI

LAMIC: Layout-Aware Multi-Image Composition via Scalability of Multimodal Diffusion Transformer

In controllable image synthesis, generating coherent and consistent images from multiple references with spatial layout awareness remains an open challenge. We present LAMIC, a Layout-Aware Multi-Image Composition framework that, for the first time, extends single-reference diffusion models to multi-reference scenarios in a training-free manner. Built upon the MMDiT model, LAMIC introduces two plug-and-play attention mechanisms: 1) Group Isolation Attention (GIA) to enhance entity disentanglement; and 2) Region-Modulated Attention (RMA) to enable layout-aware generation. To comprehensively evaluate model capabilities, we further introduce three metrics: 1) Inclusion Ratio (IN-R) and Fill Ratio (FI-R) for assessing layout control; and 2) Background Similarity (BG-S) for measuring background consistency. Extensive experiments show that LAMIC achieves state-of-the-art performance across most major metrics: it consistently outperforms existing multi-reference baselines in ID-S, BG-S, IN-R and AVG scores across all settings, and achieves the best DPG in complex composition tasks. These results demonstrate LAMIC's superior abilities in identity keeping, background preservation, layout control, and prompt-following, all achieved without any training or fine-tuning, showcasing strong zero-shot generalization ability. By inheriting the strengths of advanced single-reference models and enabling seamless extension to multi-image scenarios, LAMIC establishes a new training-free paradigm for controllable multi-image composition. As foundation models continue to evolve, LAMIC's performance is expected to scale accordingly. Our implementation is available at: https://github.com/Suchenl/LAMIC.

cs.CV

Extreme Scenario Characterization for High Renewable Energy Penetrated Power Systems over Long Time Scales

Power systems with high renewable energy penetration are highly influenced by weather conditions, often facing significant challenges such as persistent power shortages and severe power fluctuations over long time scales. This paper addresses the critical need for effective characterization of extreme scenarios under these situations. First, novel risk indices are proposed to quantify the severity of continuous power shortages and substantial power fluctuations over long-term operations. These indices are independent of specific scheduling strategies and incorporate the system's resource regulation capabilities. By employing a filtering-based approach, the proposed indices focus on retaining key characteristics of continuous power shortages and fluctuation events, enabling the identification of extreme scenarios on long time scales. Secondly, an extreme scenario generation method is developed using Gaussian mixture models and sequential Monte Carlo simulation. Especially, this method periodically evaluates the severity of generated scenarios based on the defined risk indices, retaining extreme scenarios while discarding less critical ones. Finally, case studies based on real-world data demonstrate the efficacy of the proposed method. The results confirm that integrating the identified extreme scenarios significantly enhances the system's ability to ensure long-term security and reliability under high renewable energy penetration.

eess.SY

Rooms from Motion: Un-posed Indoor 3D Object Detection as Localization and Mapping

We revisit scene-level 3D object detection as the output of an object-centric framework capable of both localization and mapping using 3D oriented boxes as the underlying geometric primitive. While existing 3D object detection approaches operate globally and implicitly rely on the a priori existence of metric camera poses, our method, Rooms from Motion (RfM) operates on a collection of un-posed images. By replacing the standard 2D keypoint-based matcher of structure-from-motion with an object-centric matcher based on image-derived 3D boxes, we estimate metric camera poses, object tracks, and finally produce a global, semantic 3D object map. When a priori pose is available, we can significantly improve map quality through optimization of global 3D boxes against individual observations. RfM shows strong localization performance and subsequently produces maps of higher quality than leading point-based and multi-view 3D object detection methods on CA-1M and ScanNet++, despite these global methods relying on overparameterization through point clouds or dense volumes. Rooms from Motion achieves a general, object-centric representation which not only extends the work of Cubify Anything to full scenes but also allows for inherently sparse localization and parametric mapping proportional to the number of objects in a scene.

cs.CV

SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding

We introduce SlowFast-LLaVA-1.5 (abbreviated as SF-LLaVA-1.5), a family of video large language models (LLMs) offering a token-efficient solution for long-form video understanding. We incorporate the two-stream SlowFast mechanism into a streamlined training pipeline, and perform joint video-image training on a carefully curated data mixture of only publicly available datasets. Our primary focus is on highly efficient model scales (1B and 3B), demonstrating that even relatively small Video LLMs can achieve state-of-the-art performance on video understanding, meeting the demand for mobile-friendly models. Experimental results demonstrate that SF-LLaVA-1.5 achieves superior performance on a wide range of video and image tasks, with robust results at all model sizes (ranging from 1B to 7B). Notably, SF-LLaVA-1.5 achieves state-of-the-art results in long-form video understanding (e.g., LongVideoBench and MLVU) and excels at small scales across various video benchmarks.

cs.CV

MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs

Multimodal large language models (MLLMs) excel at 2D visual understanding but remain limited in their ability to reason about 3D space. In this work, we leverage large-scale high-quality 3D scene data with open-set annotations to introduce 1) a novel supervised fine-tuning dataset and 2) a new evaluation benchmark, focused on indoor scenes. Our Cubify Anything VQA (CA-VQA) data covers diverse spatial tasks including spatial relationship prediction, metric size and distance estimation, and 3D grounding. We show that CA-VQA enables us to train MM-Spatial, a strong generalist MLLM that also achieves state-of-the-art performance on 3D spatial understanding benchmarks, including our own. We show how incorporating metric depth and multi-view inputs (provided in CA-VQA) can further improve 3D understanding, and demonstrate that data alone allows our model to achieve depth perception capabilities comparable to dedicated monocular depth estimation models.

cs.CV

Conditional Success of Adaptive Therapy: The Role of Treatment-Holiday Thresholds and Non-Existence of Optimal Strategies Revealed by Mathematical Modelling and Optimal Control

Adaptive therapy improves cancer treatment by controlling the competition between sensitive and resistant cells through treatment holidays. This study highlights the critical role of treatment-holiday thresholds in adaptive therapy for tumors composed of drug-sensitive and resistant cells. Using a Lotka-Volterra model, adaptive therapy outcomes are compared with maximum tolerated dose therapy and intermittent therapy outcomes, showing that adaptive therapy success depends critically on the threshold for pausing and resuming treatment and on competitive interactions between cell populations. Three comparison scenarios between adaptive therapy and other therapies emerge: uniform-decline where adaptive therapy underperforms regardless of threshold, conditional-improve where efficacy requires threshold optimization, and uniform-improve where adaptive therapy consistently outperforms alternatives. Tumor composition including initial burden and resistant cell proportion influences outcomes. Threshold adjustments enable adaptive therapy to suppress resistant subclones while preserving sensitive cells, extending progression-free survival. Crucially, this work establishes an optimal control problem for time-to-progression and mathematically proves that under biological constraints like neutral competition or low initial burden, the theoretically optimal strategy is unrealizable as it requires infinitely many treatment holidays, rendering it clinically impractical. These findings emphasize personalized treatment strategies for enhancing long-term therapeutic outcomes.

q-bio.OT

Using Drone Swarm to Stop Wildfire: A Predict-then-optimize Approach

Drone swarms coupled with data intelligence can be the future of wildfire fighting. However, drone swarm firefighting faces enormous challenges, such as the highly complex environmental conditions in wildfire scenes, the highly dynamic nature of wildfire spread, and the significant computational complexity of drone swarm operations. We develop a predict-then-optimize approach to address these challenges to enable effective drone swarm firefighting. First, we construct wildfire spread prediction convex neural network (Convex-NN) models based on real wildfire data. Then, we propose a mixed-integer programming (MIP) model coupled with dynamic programming (DP) to enable efficient drone swarm task planning. We further use chance-constrained robust optimization (CCRO) to ensure robust firefighting performances under varying situations. The formulated model is solved efficiently using Benders Decomposition and Branch-and-Cut algorithms. After 75 simulated wildfire environments training, the MIP+CCRO approach shows the best performance among several testing sets, reducing movements by 37.3\% compared to the plain MIP. It also significantly outperformed the GA baseline, which often failed to fully extinguish the fire. Eventually, we will conduct real-world fire spread and quenching experiments in the next stage for further validation.

cs.CY

A Blockwise Mixed Membership Model for Multivariate Longitudinal Data: Discovering Clinical Heterogeneity and Identifying Parkinson's Disease Subtypes

Current diagnosis and prognosis for Parkinson's disease (PD) face formidable challenges due to the heterogeneous nature of the disease course, including that (i) the impairment severity varies hugely between patients, (ii) whether a symptom occur independently or co-occurs with related symptoms differs significantly, and (iii) repeated symptom measurements exhibit substantial temporal dependence. To tackle these challenges, we propose a novel blockwise mixed membership model (BM3) to systematically unveil between-patient, between-symptom, and between-time clinical heterogeneity within PD. The key idea behind BM3 is to partition multivariate longitudinal measurements into distinct blocks, enabling measurements within each block to share a common latent membership while allowing latent memberships to vary across blocks. Consequently, the heterogeneous PD-related measurements across time are divided into clinically homogeneous blocks consisting of correlated symptoms and consecutive time. From the analysis of Parkinson's Progression Markers Initiative data (n=1,531), we discover three typical disease profiles (stages), four symptom groups (i.e., autonomic function, tremor, left-side and right-side motor function), and two periods, advancing the comprehension of PD heterogeneity. Moreover, we identify several clinically meaningful PD subtypes by summarizing the blockwise latent memberships, paving the way for developing more precise and targeted therapies to benefit patients. Our findings are validated using external variables, successfully reproduced in validation datasets, and compared with existing methods. Theoretical results of model identifiability further ensures the reliability and reproducibility of latent structure discovery in PD.

stat.AP

MapComp: A Secure View-based Collaborative Analytics Framework for Join-Group-Aggregation

Join-group-aggregation (JGA) queries are fundamental to data analytics, yet executing them collaboratively across different parties poses significant privacy risks. Secure multi-party computation (MPC) offers a cryptographic solution. However, existing MPC-based JGA approaches consider only a one-time query paradigm and suffer from significant performance bottlenecks. It executes expensive join operations from scratch across multiple queries and employs inefficient group-aggregation (GA) protocols, both of which hinder their practical use for scalable, real-time analysis. This paper introduces MapComp, a novel view-based framework to facilitate JGA queries for secure collaborative analytics. Through specially crafted materialized views for join and novel design of GA protocols, MapComp removes duplicate join workload and expedites subsequent GA, improving the efficiency of JGA query execution. To address the challenge of continuous data updates, our materialized view offers payload-independence feature and provides significant efficiency improvements in view refreshing with free MPC overhead. This feature, on the other hand, also allows further acceleration for GA, where we devise multiple novel protocols that outperform prior works. Notably, our work represents the first endeavor to expedite secure collaborative JGA queries using materialized views. Our rigorous experiments demonstrate a significant advantage of MapComp, achieving up to a 308.9x improvement in efficiency over the baseline in real-world query simulations. Moreover, our optimized GA protocols achieve up to a 1140.5x improvement compared to prior oblivious sorting-based solutions.

cs.CR