SearcharxivSearch

arXiv subjects

Xingjian Zhang

Publications and source records attributed to Xingjian Zhang.

At least 19 recordsLinked to original sources

EventMemAgent: Hierarchical Event-Centric Memory for Online Video Understanding with Adaptive Tool Use

Online video understanding requires models to perform continuous perception and long-range reasoning within potentially infinite visual streams. Its fundamental challenge lies in the conflict between the unbounded nature of streaming media input and the limited context window of Multimodal Large Language Models (MLLMs). Current methods primarily rely on passive processing, which often face a trade-off between maintaining long-range context and capturing the fine-grained details necessary for complex tasks. To address this, we introduce EventMemAgent, an active online video agent framework based on a hierarchical memory module. Our framework employs a dual-layer strategy for online videos: short-term memory detects event boundaries and utilizes event-granular reservoir sampling to process streaming video frames within a fixed-length buffer dynamically; long-term memory structuredly archives past observations on an event-by-event basis. Furthermore, we integrate a multi-granular perception toolkit for active, iterative evidence capture and employ Agentic Reinforcement Learning (Agentic RL) to end-to-end internalize reasoning and tool-use strategies into the agent's intrinsic capabilities. Experiments show that EventMemAgent achieves competitive results on online video benchmarks. The code will be released here: https://github.com/lingcco/EventMemAgent.

cs.CV

Exponential speedup of polarization stabilization for long distance DWDM quantum networks

Fibre-based quantum networks distributing polarization entanglement require a stable and uninterrupted transmission basis for reliable operation. Bright classical reference light enables rapid polarization feedback but can introduce noise into quantum channels. Entangled-photon-based feedback avoids this noise, but typically interrupts the target entanglement channel during calibration and becomes prohibitively slow over long distances due to the product loss of fibre links. Here, we overcome both limitations by combining wavelength-bracketed probing with switch-enabled path decomposition. Spectrally adjacent entangled-photon sidebands track the polarization response of the central distribution channel without interrupting its transmission, while optical switches and local reference fibres independently determine the signal and idler network transformations. Transferring the resulting compensation settings to the central channel eliminates calibration-induced downtime and changes the acquisition-time scaling from the product of the link losses to the sum of losses. We demonstrate the method on a 133 km fibre testbed and achieve continuous closed-loop stabilization for more than 24 hours without classical reference light. The decomposition of multi-link quantum feedback into single-link measurements provides a scalable stabilization strategy for wavelength-multiplexed quantum networks.

quant-ph

Size-Independent Robustness in Multipartite Bell Self-Testing

Practical robust self-testing of multipartite entanglement has so far been restricted to small-scale systems due to error bounds that degrade severely with system size. In this work, we establish multipartite self-testing with robustness independent of the size of the quantum network. We derive a fully analytic, device-independent self-testing bound for $n$-qubit Greenberger-Horne-Zeilinger (GHZ) states. The bound scales linearly with the observed violation error and lies universally within a constant factor of two from a theoretical upper bound. Furthermore, the operator-inequality framework reduces the verification of the conjectured optimal bound to a highly efficient numerical check, which we perform up to $n=100$. Consequently, GHZ entanglement can be certified under a fixed noise level in arbitrarily large systems, enabling scalable device-independent verification.

quant-ph

Broadband Characterization of Polarization Mode Dispersion for Quantum Communication Channels

We present a method for characterizing polarization fiber channels carrying broadband quantum signals, where narrowband filtering would waste photon flux. Wavelength-dependent polarization mode dispersion (PMD) maps each input state to a trajectory on the Poincaré sphere; we show that the singular value decomposition of the band-averaged rotation matrix yields, in closed form, the optimal input states, the mutually unbiased measurement bases, and their infidelities. The three singular values provide a compact, bandwidth-dependent channel signature that separates first- from higher-order PMD, and the resulting 5%-infidelity bandwidth gives a practical filtering budget. We characterize deployed fiber links in Masdar City and demonstrate PMD mitigation by concatenating two channels through a single polarization controller.

quant-ph

BehaviorBench: Benchmarking Foundation Models for Behavioral Science Tasks

Foundation models have been increasingly applied to behavioral science domains such as psychology, sociology, and economics. While these models show promise in individual tasks such as survey response prediction and human-subject experiment simulation, there remains no systematic understanding of how well they perform across diverse behavioral science tasks, contexts, and populations. We introduce BehaviorBench, a comprehensive benchmark that evaluates foundation models along four core capabilities: (1) behavior prediction and simulation, (2) strategic decision-making, (3) subject-trait inference, and (4) behavioral knowledge application. Crucially, BehaviorBench evaluates model outputs at both the individual and distributional levels, capturing not only per-subject accuracy but also population-level alignment, an essential requirement for behavioral validity. Leveraging the tasks in BehaviorBench, we further develop Be.FM-1.5, extending the Be.FM family of behavioral foundation models fine-tuned on behavioral data. Our results reveal a considerable gap: proprietary general-purpose models excel at individual-level prediction and knowledge-intensive tasks, whereas behavioral foundation models, fine-tuned on behavioral data, achieve substantially stronger distributional alignment. Notably, Be.FM-1.5 leads on distributional metrics and remains competitive on individual-level metrics, suggesting that proper behavioral adaptation can close the gap. Our results highlight the importance of distributional evaluation, establish BehaviorBench as a foundation for developing and assessing behaviorally aligned AI systems, and demonstrate Be.FM-1.5's potential for a broad range of behavioral science studies. Our BehaviorBench and Be.FM-1.5 models can be accessed via https://umich-foreseer.github.io/behaviorbench/.

cs.CL

Worst-case depth hierarchy for shallow quantum circuits

Circuit depth is a central resource in complexity theory. While bounded-depth classical circuits admit well-understood hierarchy theorems, the internal structure of constant-depth quantum computation remains comparatively unexplored. We prove an explicit depth hierarchy theorem for $\mathsf{QNC}^0$. For each $d\ge 12$, we construct a family of two-round interactive problems on which no depth-$(d-1)$ quantum circuit can achieve near-perfect success, regardless of gate set, circuit size, or ancillary qubits. In contrast, we prove that our construction admits realizations by simple bounded fan-in quantum circuits of depth larger than $d$ by a small constant factor. Moreover, all bounded fan-in classical circuits of sublogarithmic depth (in the input size) fail to achieve perfect success on these tasks for every $d$, yielding a hierarchy of problems that show unconditional quantum advantage of $\mathsf{QNC}^0$ over $\mathsf{NC}^0$. A key obstacle is the scarcity of lower bound techniques for quantum circuits. To address this, we develop methods to analyze how depth affects a circuit's ability to realize nonlocal correlations amongst its output qubits in a fine-grained manner. Our approach exploits the correspondence between constraint systems and nonlocal games, translating group-theoretic constructions into rigid operator-valued constraint systems and then into non-local games. In particular, we construct constraint systems whose unique faithful operator-valued solutions require every perfect strategy, and every near-perfect strategy to a fixed precision, to implement multi-controlled phase operations. This reduces to a nonlocal unitary-synthesis problem, yielding depth lower bounds for both shallow quantum and classical circuits. These results show that increasing depth strictly increases computational power within $\mathsf{QNC}^0$, establishing a genuinely quantum hierarchy.

quant-ph

Stream randomness extraction against quantum side information

Randomness extraction is indispensable for quantum random number generators, serving to eliminate bias and potential information leakage from raw measurement data. Conventional extractors operate in a block-wise fashion, requiring the complete accumulation of raw data before processing. To circumvent the latency and buffering overheads that hinder real-time random number generation, recent work introduced a stream-cipher implementation for the randomness extractor based on the Toeplitz matrix hashing. In this work, we generalize this stream-processing paradigm to the broader family of randomness extractors based on (almost dual) universal$_2$ random hashing. Specifically, we shift the computational burden from a time-consuming block-wise post-processing stage into an offline pre-processing stage that generates a pseudo-random mask. This allows the raw data to be processed by the mask on the fly using a simple bitwise exclusive-OR operation. Crucially, we prove that this stream implementation strictly preserves the security guarantees of the original block-wise protocols. We detail the transformation of three typical constructions -- based on standard Toeplitz, circulant, and modified Toeplitz matrices -- from block to stream implementations, and benchmark their practical performance using realistic quantum experimental data. We anticipate our framework will enhance the efficiency of real-time quantum cryptographic systems.

quant-ph

IRIS: time-structured manifold projections

High-dimensional biomedical data, such as cell-by-gene matrices, are increasingly generated temporally. However, Manifold Learning algorithms, like t-SNE and UMAP, cannot incorporate time-ordering in their layouts, obfuscating the dynamics of cell types or other classes. As a solution, we present IRIS, a new Manifold Learning algorithm that structures layouts both chronologically and by manifold topology. IRIS can visualize a wide range of dynamic biomedical data, including scRNA-seq, comparative metagenomics, and literature.

cs.LG

Scalable self-testing of generic multipartite quantum states

Characterizing large quantum systems with minimal assumptions is a central challenge in quantum information science. Self-testing provides the strongest form of certification by identifying the underlying quantum state solely from observed measurement statistics. However, existing self-testing methods for generic $n$-partite states face a scalability barrier, requiring exponentially many samples in the system size. In this work, we overcome this barrier by introducing a protocol that robustly self-tests almost all $n$-qubit states with only polynomial sample complexity. The key ingredient is an efficient scheme for device-independently evaluating multipartite Pauli measurements, which can be implemented using only a linear number of ancillary Bell pairs together with standard projective and Bell measurements, well within the reach of current quantum technology. Beyond self-testing states, our scheme provides a general framework for implementing a wide range of learning and certification protocols in the device-independent setting, thereby opening a scalable route to device-independent quantum information processing in large-scale quantum networks.

quant-ph

Expanding mmWave Datasets for Human Pose Estimation with Unlabeled Data and LiDAR Datasets

Current millimeter-wave (mmWave) datasets for human pose estimation (HPE) are scarce and lack diversity in both point cloud (PC) attributes and human poses, hindering the generalization ability of their trained models. On the other hand, unlabeled mmWave HPE data and diverse LiDAR HPE datasets are readily available. We propose EMDUL, a novel approach to expand the volume and diversity of an existing mmWave dataset using unlabeled mmWave data and LiDAR datasets. EMDUL consists of two independent modules, namely a pseudo-label estimator to annotate unlabeled mmWave data, and a closed-form converter that translates an annotated LiDAR PC to its mmWave counterpart. Expanding the original dataset with both LiDAR-converted and pseudo-labeled mmWave PCs significantly boosts the performance and generalization ability of all the examined HPE models, reducing 15.1% and 18.9% error for in-domain and out-of-domain settings, respectively. Code is available at https://github.com/Shimmer93/EMDUL.

cs.CV

Distilling the knowledge with quantum neural networks

Quantum Neural Networks (QNNs) are a promising class of quantum machine learning models with potential quantum advantages when implemented on scalable, error-corrected quantum computers. However, as system sizes increase, deploying QNNs becomes challenging. Similar to their classical counterparts, a key obstacle to their practical applications is that large-scale QNNs may not be easily deployed on smaller systems that have limited resources. Here, we tackle this challenge by compressing QNNs via knowledge distillation. We demonstrate how well-trained QNNs on large systems can be distilled into smaller architectures with similar configurations. We numerically show that knowledge distillation helps reduce the training cost of QNNs in terms of the number of qubits and circuit depth. Additionally, we find that a self-knowledge-distillation approach can accelerate training convergence. We believe our results offer new strategies for the efficient compression and practical deployment of QNNs.

quant-ph

Entanglement distribution over 155 km metropolitan fiber using a CMOS-compatible silicon chip

Transmitting entangled states over long distances is crucial for developing quantum networks. Previous demonstrations using satellites or fibers relied on photon pairs generated from bulk crystal arrangements. Polarization entanglement distribution based on CMOS-compatible silicon chips has long been restricted to lab-scale demonstrations spanning only a few meters, due to the difficulty of achieving sufficient off-chip brightness. We report a silicon chip platform that provides an off-chip entangled photon pair brightness ranging from 8,000 to 460,000 pairs per second, exceeding previous reports by three orders of magnitude. The entanglement fidelity reaches 99.85(6)% and 97.90(3)%, respectively. After addressing key challenges in long distance entanglement distribution over deployed fiber, including phase drift and chromatic dispersion, entangled photons were successfully distributed over 155 km (66 dB loss). These results demonstrate that CMOS-compatible silicon chips can perform competitively with bulk crystal sources and represent an important step toward scalable, chip-based quantum networks.

quant-ph

Through the Judge's Eyes: Inferred Thinking Traces Improve Reliability of LLM Raters

Large language models (LLMs) are increasingly used as raters for evaluation tasks. However, their reliability is often limited for subjective tasks, when human judgments involve subtle reasoning beyond annotation labels. Thinking traces, the reasoning behind a judgment, are highly informative but challenging to collect and curate. We present a human-LLM collaborative framework to infer thinking traces from label-only annotations. The proposed framework uses a simple and effective rejection sampling method to reconstruct these traces at scale. These inferred thinking traces are applied to two complementary tasks: (1) fine-tuning open LLM raters; and (2) synthesizing clearer annotation guidelines for proprietary LLM raters. Across multiple datasets, our methods lead to significantly improved LLM-human agreement. Additionally, the refined annotation guidelines increase agreement among different LLM models. These results suggest that LLMs can serve as practical proxies for otherwise unrevealed human thinking traces, enabling label-only corpora to be extended into thinking-trace-augmented resources that enhance the reliability of LLM raters.

cs.AI

FlyAOC: Evaluating Agentic Ontology Curation of Drosophila Scientific Knowledge Bases

Scientific knowledge bases accelerate discovery by curating findings from primary literature into structured, queryable formats for both human researchers and emerging AI systems. Maintaining these resources requires expert curators to search relevant papers, reconcile evidence across documents, and produce ontology-grounded annotations - a workflow that existing benchmarks, focused on isolated subtasks like named entity recognition or relation extraction, do not capture. We present FlyBench to evaluate AI agents on end-to-end agentic ontology curation from scientific literature. Given only a gene symbol, agents must search and read from a corpus of 16,898 full-text papers to produce structured annotations: Gene Ontology terms describing function, expression patterns, and historical synonyms linking decades of nomenclature. The benchmark includes 7,397 expert-curated annotations across 100 genes drawn from FlyBase, the Drosophila (fruit fly) knowledge base. We evaluate four baseline agent architectures: memorization, fixed pipeline, single-agent, and multi-agent. We find that architectural choices significantly impact performance, with multi-agent designs outperforming simpler alternatives, yet scaling backbone models yields diminishing returns. All baselines leave substantial room for improvement. Our analysis surfaces several findings to guide future development; for example, agents primarily use retrieval to confirm parametric knowledge rather than discover new information. We hope FlyBench will drive progress on retrieval-augmented scientific reasoning, a capability with broad applications across scientific domains.

cs.AI

Log Focal Frequency Loss for Bioimage Restoration

Image restoration of biological structures in microscopy poses unique challenges for preserving fine textures and sharp edges. While recent GAN-based image restoration formulations have introduced frequency-domain losses for natural images, microscopy images pose distinct challenges with large dynamic ranges and sparse but critical structures with spatially-variable contrast. Inspired by the principle of logarithmic perception in human vision, we propose a log focal frequency loss (LFFL) tailored for microscopy restoration. This loss combines adaptive spectral weighting from log-space differences with log-dampened error measurement, ensuring balanced reconstruction across all frequency bands while preserving both structural coherence and fine details. We tested our GAN-based framework on two use-cases with real ground-truths: deblurring of fluorescence images of cell nuclei on microgroove substrates and denoising of zebrafish embryo images from the FMD dataset. Compared to training with only spatial-domain losses and with existing frequency-domain losses, our method achieves improvements across several quality metrics. Code is available at github.com/xjzhaang/log-focal-frequency-loss.

q-bio.QM

Your Reasoning Benchmark May Not Test Reasoning: Revealing Perception Bottleneck in Abstract Reasoning Benchmarks

Reasoning benchmarks such as the Abstraction and Reasoning Corpus (ARC) and ARC-AGI are widely used to assess progress in artificial intelligence and are often interpreted as probes of core, so-called ``fluid'' reasoning abilities. Despite their apparent simplicity for humans, these tasks remain challenging for frontier vision-language models (VLMs), a gap commonly attributed to deficiencies in machine reasoning. We challenge this interpretation and hypothesize that the gap arises primarily from limitations in visual perception rather than from shortcomings in inductive reasoning. To verify this hypothesis, we introduce a two-stage experimental pipeline that explicitly separates perception and reasoning. In the perception stage, each image is independently converted into a natural-language description, while in the reasoning stage a model induces and applies rules using these descriptions. This design prevents leakage of cross-image inductive signals and isolates reasoning from perception bottlenecks. Across three ARC-style datasets, Mini-ARC, ACRE, and Bongard-LOGO, we show that the perception capability is the dominant factor underlying the observed performance gap by comparing the two-stage pipeline with against standard end-to-end one-stage evaluation. Manual inspection of reasoning traces in the VLM outputs further reveals that approximately 80 percent of model failures stem from perception errors. Together, these results demonstrate that ARC-style benchmarks conflate perceptual and reasoning challenges and that observed performance gaps may overstate deficiencies in machine reasoning. Our findings underscore the need for evaluation protocols that disentangle perception from reasoning when assessing progress in machine intelligence.

cs.CL

On the physics of nested Markov models: a generalized probabilistic theory perspective

Determining potential probability distributions with a given causal graph is vital for causality studies. To bypass the difficulty in characterizing latent variables in a Bayesian network, the nested Markov model provides an elegant algebraic approach by listing exactly all the equality constraints on the observed variables. However, this algebraically motivated causal model comprises distributions outside Bayesian networks, and its physical interpretation remains vague. In this work, we inspect the nested Markov model through the lens of generalized probabilistic theory, an axiomatic framework to describe general physical theories. We prove that all the equality constraints defining the nested Markov model are valid theory-independently. At the same time, not every distribution within the nested Markov model is implementable, not even via exotic physical theories associated with generalized probability theories (GPTs). To interpret the origin of such a gap, we study three causal models standing between the nested Markov model and the set of all distributions admitting some GPT realization. Each of the successive three models gives a strictly tighter characterization of the physically implementable distribution set; that is, each successive model manifests new types of GPT-inviolable constraints. We further demonstrate each gap through a specially chosen illustrative causal structure. We anticipate our results will enlighten further explorations on the unification of algebraic and physical perspectives of causality.

quant-ph

SIGMA: An AI-Empowered Training Stack on Early-Life Hardware

An increasing variety of AI accelerators is being considered for large-scale training. However, enabling large-scale training on early-life AI accelerators faces three core challenges: frequent system disruptions and undefined failure modes that undermine reliability; numerical errors and training instabilities that threaten correctness and convergence; and the complexity of parallelism optimization combined with unpredictable local noise that degrades efficiency. To address these challenges, SIGMA is an open-source training stack designed to improve the reliability, stability, and efficiency of large-scale distributed training on early-life AI hardware. The core of this initiative is the LUCIA TRAINING PLATFORM (LTP), the system optimized for clusters with early-life AI accelerators. Since its launch in March 2025, LTP has significantly enhanced training reliability and operational productivity. Over the past five months, it has achieved an impressive 94.45% effective cluster accelerator utilization, while also substantially reducing node recycling and job-recovery times. Building on the foundation of LTP, the LUCIA TRAINING FRAMEWORK (LTF) successfully trained SIGMA-MOE, a 200B MoE model, using 2,048 AI accelerators. This effort delivered remarkable stability and efficiency outcomes, achieving 21.08% MFU, state-of-the-art downstream accuracy, and encountering only one stability incident over a 75-day period. Together, these advances establish SIGMA, which not only tackles the critical challenges of large-scale training but also establishes a new benchmark for AI infrastructure and platform innovation, offering a robust, cost-effective alternative to prevailing established accelerator stacks and significantly advancing AI capabilities and scalability. The source code of SIGMA is available at https://github.com/microsoft/LuciaTrainingPlatform.

cs.DC