SearcharxivSearch

arXiv subjects

Ning Xie

Publications and source records attributed to Ning Xie.

At least 19 recordsLinked to original sources

Operationalizing open-ended biological discovery across single-cell representations

Single-cell studies are typically initiated from predefined research questions, leaving much of the biological information encoded within existing data unexplored. We formalize open-ended discovery as an analytical paradigm, in which data-derived signals are identified before biological context is interrogated and subsequently evaluated according to their potential to justify prospective experimental investment. Here we develop PROSPECTor, an end-to-end framework that searches for reproducible biological structures across conventional expression representations and diverse foundation-model embeddings, translating robust signals into quantitatively testable candidate hypotheses. Projection into unseen datasets then evaluates their generalizability and phenotype association, providing a scalable screen for candidates that warrant prospective validation. Supported signals emerged from different representation spaces and search strategies. PROSPECTor-nominated hypotheses were then examined in independent biological settings: fibroblast extracellular-matrix programmes demonstrated transferability to an independent mouse cohort with an intervention context, while a patient-resolved gastric-cancer T-cell programme recurred across single-cell, bulk and spatial cohorts. PROSPECTor establishes an auditable framework for systematically revisiting single-cell datasets across expanding representation spaces, turning retrospective collections into prospective resources for biological discovery that can motivate new research questions.

q-bio.QM

Link-Based Multimodal Traffic Dynamics Model in Continuous-Time Framework

Computationally efficient models for multimodal traffic flows with inter-modal interactions are foundational for coordinated traffic management. However, such models are lacking in the current literature. This study introduces a multimodal link transmission model (M-LTM) accommodating both continuous road traffic and discrete tramway traffic at the network level. M-LTM builds on a link model that captures the inter-modal interactions through the moving bottleneck theory. We further develop node models describing the flow transfer, tram operations, and mode interactions at tram stops and intersections. Specifically, the tram dwell process and dwell-induced congestion of road traffic are simulated at the stop node, and the intersection node model can reproduce queue spillback and tram priority. Case studies are conducted on two synthetic one-node networks, a synthetic arterial network, and the real network of Dresden to demonstrate the predictive power and the applicability of the proposed model in network traffic analysis and management. The GEH statistics of road traffic are below 4, and the average tram arrival time absolute errors are less than 10 seconds.

math.OC

Geometry-Aware Multi-Armed Bandits for Antenna Beam Selection on Spheres, Tori, $\SO(3)$, and Reconfigurable Intelligent Surfaces

Beam alignment in mmWave phased arrays and RIS-assisted links is a stochastic bandit under both short TTI budgets and Doppler-induced non-stationarity. The arm space is a Riemannian manifold: $\sphere^2$ for steering, $\torus^n$ for phase combining, $\SO(3)$ for panel orientation, or the discrete torus $(\mathbb Z_B)^M$ with up to $K\!\sim\!10^{90}$ configurations for $B$-level RIS ($B\!=\!2^b$, $b$ bits/element); the intrinsic Mat\'ern kernel of Borovitskiy et al.\ provides the base GP. We contribute two algorithmic pieces. \textbf{(C1)} A Kronecker-factorised intrinsic-product Mat\'ern kernel on $(\mathbb Z_B)^M$ evaluating in $O(M)$ table lookups, making GP-UCB tractable at $K\sim 10^{90}$ where the extrinsic alternative is infeasible. \textbf{(C2)} AdaptiveGP-v2, an online sliding-window controller that selects $W$ by per-sample marginal likelihood, with predictive-variance and drift $z$-score reset triggers and a post-reset $\beta$-boost. On a four-speed ($v\!\in\!\{0.02,0.08,0.12,0.20\}$~km/h), $20$-seed paired campaign at $T\!=\!3000$, AdaptiveGP-v2 is statistically indistinguishable from the hand-tuned fixed-window oracle at every speed (Holm--Bonferroni-corrected paired differences cross zero); the operational benefit is the absence of a deployment-time per-speed calibration step, not a mean-regret improvement. On four static 3GPP-style mmWave benchmarks, intrinsic-kernel GP-UCB reduces cumulative regret by $25$--$45\%$ vs.\ codebook UCB1/Thompson and by $10$--$33\%$ vs.\ Euclidean-ambient GP-UCB on the toroidal arm spaces; a wideband OFDM ablation on a $100$~MHz channel confirms the advantage persists under frequency-selective fading ($\sim\!32$~Mbps/UE at initial access vs.\ UCB1). A third-party-simulator sanity check on Sionna CDL is reported in Section~V.

eess.SP

Manifold-Aware Information Gain and Lower Bounds for Gaussian-Process Bandits on Riemannian Quotient Spaces

We prove a regret lower bound for Gaussian-process bandits on a smooth compact Riemannian manifold $\M$ of dimension $d$ with intrinsic Mat\'ern-$\nu$ kernel ($\nu>d/2$) that exposes how the geometry of the arm space enters the constant. For any algorithm and time horizon $T$ exceeding an explicit threshold, the worst-case expected regret over the RKHS-ball $\|f\|_{\Hil_{k_\nu}}\!\le\!B$ satisfies \begin{multline*} \E[R_T(f)]\;\ge\;c_*(d,\nu)\,B^{d/(2\nu+d)}\,\sigma_n^{2\nu/(2\nu+d)} \\ \cdot\,\vol_g(\M)^{\nu/(2\nu+d)}\,T^{(\nu+d)/(2\nu+d)}(\log T)^{\nu/(2\nu+d)}. \end{multline*} The exponent matches the Vakili--Khezeli--Picheny upper bound \cite{vakili2021information}; the $\vol_g(\M)^{\nu/(2\nu+d)}$ factor is, to our knowledge, the first explicit volume-dependent geometric constant in a manifold GP-bandit lower bound. We extend the analysis in five directions: (i)~a companion Assouad-style proof gives a different lower bound with a strictly smaller $T$-exponent $(2\nu+3d)/(4(\nu+d))$ but with a polylog factor of the form $1/(\log\log T)^{(2\nu+d)/(4(\nu+d))}$, sharpening the $(\log T)^{\nu/(2\nu+d)}$ Fano polylog of Theorem~\ref{thm:main}; (ii)~we prove a $|G|^{1/2}$ upper bound on the regret of an extrinsic-kernel GP-UCB algorithm on a quotient space $\M=\Mt/G$, plus a bracketing theorem (Theorem~\ref{thm:gauge-bracket}); the precise constant is conjectured to take the modulated form $(1+(|G|-1)h(\rinj/\kappa))^{1/2}$ (Conjecture~\ref{conj:gauge-modulated}), validated numerically on $\SO(3)$; (iii)~we write the leading constant $c_*(d,\nu)$ out fully; (iv)~we extract a curvature dependence $1+O(K\eps_T^2)$ via Bishop--Gromov; (v)~we transfer the bound to the Bayesian regret framework via the Yang--Barron / Castillo et al.\ Bayesian-Fano transfer.

eess.SP

Frozen-Tag-Based Physical-Layer Authentication Against User Interference

Tag-based physical layer authentication (PLA) has garnered significant attention due to its low complexity and enhanced security. However, existing PLA schemes encounter two challenges. First, unintended user interference, which overlaps with the authentication signal, corrupts the tag and degrades authentication performance. Second, the vulnerability introduced by direct embedding of the raw tag exposes the tag to the adversary and degrades the security. To address these challenges, this paper proposes a novel frozen-tag-based PLA framework. Different from typical schemes that directly embed the uncoded tag into the signal, a well-designed frozen tag is inserted for authentication, where the frozen tag is generated based on the concept of polar codes with the anchor information as information bits and raw tags as frozen bits. Accordingly, the proposed PLA framework offers two principal advantages. First, the authentication performance is improved since the legitimate receiver can decode the frozen tag and mitigate unintended user interference. Second, the authentication process becomes indecipherable to the illegitimate receiver due to the concealment of the raw tags. Furthermore, we conduct a comprehensive analysis of the proposed framework in terms of robustness, security, and compatibility. Theoretical analysis and simulation demonstrate that the proposed frozen-tag-based PLA framework not only enhances the detection performance but also significantly degrades Eve's capability to estimate the raw tags.

cs.IT

How did the Urban Network Flow Adapt to the Collapse of the Carola Bridge?

The unexpected collapse of the Carola Bridge in Dresden, Germany, provides a rare opportunity to characterise how urban network traffic adapts to an unexpected infrastructure disruption. This study develops a data-driven analytical framework using traffic data from the Dresden traffic management system to assess the short-term impacts of the disruption. By combining statistical comparisons of pre- and post-collapse motorised traffic distributions, peak-hour shifts, and Park-and-Ride data analyses, the framework reveals how traffic dynamics and traveller choices adjust under infrastructure disruption. Results reveal that the two closest bridges, the Albert and Marien Bridges, absorb the majority of the diverted motorised traffic. In particular, the daily traffic volume on the Albert bridge increases by up to 81%, which is equivalent to 3.5 hours of traffic operating with maximum flow. Peak hours on critical links are significantly prolonged, reaching up to 250 minutes. Besides redistribution, the overall daily motorised traffic crossing the Elbe river declines by approximately 8,000 vehicles, while Park-and-Ride usage increases by up to 188%, suggesting a potential travel mode shift after the disruption. The study reveals the patterns of traffic redistribution following an unexpected disruption and provides insights for resilience planning and emergency traffic management.

nlin.AO

Efficient and Interpretable Multi-Agent LLM Routing via Ant Colony Optimization

Large Language Model (LLM)-driven Multi-Agent Systems (MAS) have demonstrated strong capability in complex reasoning and tool use, and heterogeneous agent pools further broaden the quality--cost trade-off space. Despite these advances, real-world deployment is often constrained by high inference cost, latency, and limited transparency, which hinders scalable and efficient routing. Existing routing strategies typically rely on expensive LLM-based selectors or static policies, and offer limited controllability for semantic-aware routing under dynamic loads and mixed intents, often resulting in unstable performance and inefficient resource utilization. To address these limitations, we propose AMRO-S, an efficient and interpretable routing framework for Multi-Agent Systems (MAS). AMRO-S models MAS routing as a semantic-conditioned path selection problem, enhancing routing performance through three key mechanisms: First, it leverages a supervised fine-tuned (SFT) small language model for intent inference, providing a low-overhead semantic interface for each query; second, it decomposes routing memory into task-specific pheromone specialists, reducing cross-task interference and optimizing path selection under mixed workloads; finally, it employs a quality-gated asynchronous update mechanism to decouple inference from learning, optimizing routing without increasing latency. Extensive experiments on five public benchmarks and high-concurrency stress tests demonstrate that AMRO-S consistently improves the quality--cost trade-off over strong routing baselines, while providing traceable routing evidence through structured pheromone patterns.

cs.AI

Global-Local Dual Perception for MLLMs in High-Resolution Text-Rich Image Translation

Text Image Machine Translation (TIMT) aims to translate text embedded in images in the source-language into target-language, requiring synergistic integration of visual perception and linguistic understanding. Existing TIMT methods, whether cascaded pipelines or end-to-end multimodal large language models (MLLMs),struggle with high-resolution text-rich images due to cluttered layouts, diverse fonts, and non-textual distractions, resulting in text omission, semantic drift, and contextual inconsistency. To address these challenges, we propose GLoTran, a global-local dual visual perception framework for MLLM-based TIMT. GLoTran integrates a low-resolution global image with multi-scale region-level text image slices under an instruction-guided alignment strategy, conditioning MLLMs to maintain scene-level contextual consistency while faithfully capturing fine-grained textual details. Moreover, to realize this dual-perception paradigm, we construct GLoD, a large-scale text-rich TIMT dataset comprising 510K high-resolution global-local image-text pairs covering diverse real-world scenarios. Extensive experiments demonstrate that GLoTran substantially improves translation completeness and accuracy over state-of-the-art MLLMs, offering a new paradigm for fine-grained TIMT under high-resolution and text-rich conditions.

cs.CV

MM-R1: Unleashing the Power of Unified Multimodal Large Language Models for Personalized Image Generation

Multimodal Large Language Models (MLLMs) with unified architectures excel across a wide range of vision-language tasks, yet aligning them with personalized image generation remains a significant challenge. Existing methods for MLLMs are frequently subject-specific, demanding a data-intensive fine-tuning process for every new subject, which limits their scalability. In this paper, we introduce MM-R1, a framework that integrates a cross-modal Chain-of-Thought (X-CoT) reasoning strategy to unlock the inherent potential of unified MLLMs for personalized image generation. Specifically, we structure personalization as an integrated visual reasoning and generation process: (1) grounding subject concepts by interpreting and understanding user-provided images and contextual cues, and (2) generating personalized images conditioned on both the extracted subject representations and user prompts. To further enhance the reasoning capability, we adopt Grouped Reward Proximal Policy Optimization (GRPO) to explicitly align the generation. Experiments demonstrate that MM-R1 unleashes the personalization capability of unified MLLMs to generate images with high subject fidelity and strong text alignment in a zero-shot manner.

cs.CV

Training-Free ANN-to-SNN Conversion for High-Performance Spiking Transformer

Leveraging the event-driven paradigm, Spiking Neural Networks (SNNs) offer a promising approach for energy-efficient Transformer architectures.While ANN-to-SNN conversion avoids the high training cost of directly trained Spiking Transformers, existing approaches still struggle to handle the nonlinear operations within Transformer blocks, and often require additional fine-tuning of pretrained ANNs.To address these limitations, we propose a training-free and high-performance ANN-to-SNN conversion framework tailored for Transformer architectures. Specifically, we introduce a Multi-basis Exponential Decay (MBE) neuron that combines exponential decay with a multi-basis encoding strategy to effectively approximate nonlinear operations, eliminating the need for weight modifications in pretrained ANNs.Extensive experiments across diverse tasks (CV, NLU, NLG) and mainstream Transformer architectures (ViT, RoBERTa, GPT-2) demonstrate that our method achieves near-lossless conversion accuracy with significantly lower latency. This provides a promising pathway for the efficient and scalable deployment of Spiking Transformers in real-world applications.

cs.LG

DIMT25@ICDAR2025: HW-TSC's End-to-End Document Image Machine Translation System Leveraging Large Vision-Language Model

This paper presents the technical solution proposed by Huawei Translation Service Center (HW-TSC) for the "End-to-End Document Image Machine Translation for Complex Layouts" competition at the 19th International Conference on Document Analysis and Recognition (DIMT25@ICDAR2025). Leveraging state-of-the-art open-source large vision-language model (LVLM), we introduce a training framework that combines multi-task learning with perceptual chain-of-thought to develop a comprehensive end-to-end document translation system. During the inference phase, we apply minimum Bayesian decoding and post-processing strategies to further enhance the system's translation capabilities. Our solution uniquely addresses both OCR-based and OCR-free document image translation tasks within a unified framework. This paper systematically details the training methods, inference strategies, LVLM base models, training data, experimental setups, and results, demonstrating an effective approach to document image machine translation.

cs.CV

Evaluating Menu OCR and Translation: A Benchmark for Aligning Human and Automated Evaluations in Large Vision-Language Models

The rapid advancement of large vision-language models (LVLMs) has significantly propelled applications in document understanding, particularly in optical character recognition (OCR) and multilingual translation. However, current evaluations of LVLMs, like the widely used OCRBench, mainly focus on verifying the correctness of their short-text responses and long-text responses with simple layout, while the evaluation of their ability to understand long texts with complex layout design is highly significant but largely overlooked. In this paper, we propose Menu OCR and Translation Benchmark (MOTBench), a specialized evaluation framework emphasizing the pivotal role of menu translation in cross-cultural communication. MOTBench requires LVLMs to accurately recognize and translate each dish, along with its price and unit items on a menu, providing a comprehensive assessment of their visual understanding and language processing capabilities. Our benchmark is comprised of a collection of Chinese and English menus, characterized by intricate layouts, a variety of fonts, and culturally specific elements across different languages, along with precise human annotations. Experiments show that our automatic evaluation results are highly consistent with professional human evaluation. We evaluate a range of publicly available state-of-the-art LVLMs, and through analyzing their output to identify the strengths and weaknesses in their performance, offering valuable insights to guide future advancements in LVLM development. MOTBench is available at https://github.com/gitwzl/MOTBench.

cs.LG

Self-interaction effects on the Kerr black hole superradiance and their observational implications

Through the black hole (BH) superradiance, ultralight bosons can form dense clouds around rotating Kerr BHs. Certain ultralight bosons, such as axions and axion-like particles (promising dark matter candidates), naturally possess self-interactions, and thus may significantly modify the dynamics of the superradiance process. Previous studies on the detection or constraint of ultralight bosons through superradiance have usually neglected the self-interaction effects of bosons. In this work, we investigate the formation and evolution of self-interacting boson clouds in the full Kerr spacetime during BH superradiance. Using numerical methods, we compute the superradiant growth rate of boson clouds with self-interactions around Kerr BHs and quantitatively evaluate how the self-interaction strength of scalar bosons affects the growth rate. We also assess the evolution of the BH's mass and spin. Our results reveal that, in addition to the superradiance-imposed upper bound on the boson cloud mass, self-interaction of ultralight bosons introduces a new, lower critical mass limit, beyond which the growth rate of the boson cloud approaches zero. This implies that the superradiance process terminates earlier when self-interaction is considered. Furthermore, we explore how self-interaction affects both the oscillation frequency of boson clouds in gravitational atoms and the frequency of gravitational wave (GW) emitted through cloud annihilation. The anticipated frequency shift might be detectable by the GW observatories. Given that self-interaction substantially alters the evolution of BH superradiance, its effects can significantly relax existing constraints on scalar bosons derived from superradiance. Taking the spin measurements from GW190412 and GW190517 as examples, we discuss the impact of self-interaction on constraint results in details.

hep-ph

InPK: Infusing Prior Knowledge into Prompt for Vision-Language Models

Prompt tuning has become a popular strategy for adapting Vision-Language Models (VLMs) to zero/few-shot visual recognition tasks. Some prompting techniques introduce prior knowledge due to its richness, but when learnable tokens are randomly initialized and disconnected from prior knowledge, they tend to overfit on seen classes and struggle with domain shifts for unseen ones. To address this issue, we propose the InPK model, which infuses class-specific prior knowledge into the learnable tokens during initialization, thus enabling the model to explicitly focus on class-relevant information. Furthermore, to mitigate the weakening of class information by multi-layer encoders, we continuously reinforce the interaction between learnable tokens and prior knowledge across multiple feature levels. This progressive interaction allows the learnable tokens to better capture the fine-grained differences and universal visual concepts within prior knowledge, enabling the model to extract more discriminative and generalized text features. Even for unseen classes, the learned interaction allows the model to capture their common representations and infer their appropriate positions within the existing semantic structure. Moreover, we introduce a learnable text-to-vision projection layer to accommodate the text adjustments, ensuring better alignment of visual-text semantics. Extensive experiments on 11 recognition datasets show that InPK significantly outperforms state-of-the-art methods in multiple zero/few-shot image classification tasks.

cs.CV

Zeitgebers-Based User Time Perception Analysis and Data-Driven Modeling via Transformer in VR

Virtual Reality (VR) creates a highly realistic and controllable simulation environment that can manipulate users' sense of space and time. While the sensation of "losing track of time" is often associated with enjoyable experiences, the link between time perception and user experience in VR and its underlying mechanisms remains largely unexplored. This study investigates how different zeitgebers-light color, music tempo, and task factor-influence time perception. We introduced the Relative Subjective Time Change (RSTC) method to explore the relationship between time perception and user experience. Additionally, we applied a data-driven approach called the Time Perception Modeling Network (TPM-Net), which integrates Convolutional Neural Network (CNN) and Transformer architectures to model time perception based on multimodal physiological and zeitgebers data. With 56 participants in a between-subject experiment, our results show that task factors significantly influence time perception, with red light and slow-tempo music further contributing to time underestimation. The RSTC method reveals that underestimating time in VR is strongly associated with improved user experience, presence, and engagement. Furthermore, TPM-Net shows potential for modeling time perception in VR, enabling inference of relative changes in users' time perception and corresponding changes in user experience. This study provides insights into the relationship between time perception and user experience in VR, with applications in VR-based therapy and specialized training.

cs.HC

Multimodal Instruction Tuning with Hybrid State Space Models

Handling lengthy context is crucial for enhancing the recognition and understanding capabilities of multimodal large language models (MLLMs) in applications such as processing high-resolution images or high frame rate videos. The rise in image resolution and frame rate substantially increases computational demands due to the increased number of input tokens. This challenge is further exacerbated by the quadratic complexity with respect to sequence length of the self-attention mechanism. Most prior works either pre-train models with long contexts, overlooking the efficiency problem, or attempt to reduce the context length via downsampling (e.g., identify the key image patches or frames) to decrease the context length, which may result in information loss. To circumvent this issue while keeping the remarkable effectiveness of MLLMs, we propose a novel approach using a hybrid transformer-MAMBA model to efficiently handle long contexts in multimodal applications. Our multimodal model can effectively process long context input exceeding 100k tokens, outperforming existing models across various benchmarks. Remarkably, our model enhances inference efficiency for high-resolution images and high-frame-rate videos by about 4 times compared to current models, with efficiency gains increasing as image resolution or video frames rise. Furthermore, our model is the first to be trained on low-resolution images or low-frame-rate videos while being capable of inference on high-resolution images and high-frame-rate videos, offering flexibility for inference in diverse scenarios.

cs.CV

Approaches to Simultaneously Solving Variational Quantum Eigensolver Problems

The variational quantum eigensolver (VQE), a type of variational quantum algorithm, is a hybrid quantum-classical algorithm to find the lowest-energy eigenstate of a particular Hamiltonian. We investigate ways to optimize the VQE solving process on multiple instances of the same problem, by observing the process on one instance of the problem to inform initialization for other processes. We aim to take advantage of the VQE solution process to obtain useful information while disregarding information which we can predict to not be very useful. In particular, we find that the solution process produces lots of data with very little new information. Therefore, we can safely disregard much of this repetitive information with little effect on the outcome of the solution process.

quant-ph

Efficient Circuit Wire Cutting Based on Commuting Groups

Current quantum devices face challenges when dealing with large circuits due to error rates as circuit size and the number of qubits increase. The circuit wire-cutting technique addresses this issue by breaking down a large circuit into smaller, more manageable subcircuits. However, the exponential increase in the number of subcircuits and the complexity of reconstruction as more cuts are made poses a great practical challenge. Inspired by ancilla-assisted quantum process tomography and the MUBs-based grouping technique for simultaneous measurement, we propose a new approach that can reduce subcircuit running overhead. The approach first uses ancillary qubits to transform all quantum input initializations into quantum output measurements. These output measurements are then organized into commuting groups for the purpose of simultaneous measurement, based on MUBs-based grouping. This approach significantly reduces the number of necessary subcircuits as well as the total number of shots. Lastly, we provide numerical experiments to demonstrate the complexity reduction.

quant-ph