SearcharxivSearch

arXiv subjects

Xin Wang

Publications and source records attributed to Xin Wang.

At least 19 recordsLinked to original sources

Mr.LHDR: A Benchmark for Multimodal Real-World Long-Horizon Deep Research Agents

Deep research agents are increasingly capable of web search, tool use, multimodal evidence analysis, and information synthesis. However, existing benchmarks mainly evaluate medium-horizon exploration and rarely test whether agents can sustain long, dependency-heavy research processes. We introduce Mr.LHDR (Multimodal real-world Long-Horizon Deep Research), a benchmark for evaluating real-world deep research over long, irreducible chains of interdependent evidence across eight categories. Each question is constructed from a hidden Node-Relation graph and requires an average of 12.1 necessary intermediate conclusions with a mean dependency depth of 10.4 before reaching a short, unique, and verifiable answer. Questions incorporate multimodal evidence, including images, maps, PDFs, logos, charts, tables, and video frames, with at least one non-text element that changes the reasoning state. Mr.LHDR evaluates both final answers and the correctness of intermediate conclusions under annotated dependencies. We evaluate general models, deep research systems, and agent frameworks using Overall Accuracy (OA), Strict Accuracy (SA), Checklist Score (CS), and Dependency-Aware Checklist Score (DACS). Results show that even the strongest system achieves only 43.1% OA and 34.3% SA, indicating that final-answer accuracy substantially overestimates complete research success. Removing images reduces DACS by 12.6 points, demonstrating the importance of multimodal evidence, while SA consistently declines as reasoning chains become longer. These findings reveal sustained, dependency-consistent evidence integration, rather than isolated fact retrieval, as a key bottleneck for current deep research agents.

cs.AI

Pre- and Post-Treatment Brain Metastases Segmentation Using nnU-Net with Post-Processing for BraTS 2026

Brain metastases exhibit high inter-lesion variability in size, enhancement pattern, and post-treatment appearance, making volumetric segmentation of both pre- and post-treatment cases the central challenge of the BraTS 2026 Task 1 (Brain Metastases). We build a pragmatic pipeline on a 5-fold nnU-Net ResEnc-L ensemble, in which each fold is trained independently for 1,000 epochs with the standard Dice + cross-entropy loss on 1,296 four-modality training cases. This ensemble is followed by a rule-based post-processing cascade tuned for the lesion-wise Dice similarity coefficient (LW-DSC), a detection-oriented metric that behaves very differently from the traditional global Dice. The final pipeline reaches an LW-DSC of 0.733 / 0.751 / 0.713 / 0.549 on the enhancing tumour (ET), tumour core (TC), whole tumour (WT), and resection cavity (RC) sub-regions on the official validation leaderboard. Rather than trusting these leaderboard gains, we audit every post-processing stage with a five-fold out-of-fold (OOF) analysis with no model-training leakage over all 1,296 training cases, scored with the official BraTS evaluation code (BraTS_evaluation): it confirms two stages as robust, per-fold-consistent improvements while the third improves only the leaderboard and does not reproduce out-of-fold. We further provide a mechanistic analysis of the LW-DSC metric that explains why recall-recovering post-processing carries low risk whereas component deletion does not, and we report thirteen negative results spanning loss engineering, alternative backbones, and inference-time settings, several of which run counter to widely held intuitions. Source code is released under Apache-2.0 at https://github.com/hornbeamliu/brats2026-met.

cs.CV

An exchange-assisted entangling gate between 87Rb and 171Yb Rydberg atoms

Neutral-atom tweezer arrays support scalable quantum information processing. Dual-species $^{87}\mathrm{Rb}$--$^{171}\mathrm{Yb}$ arrays combine long-lived ytterbium nuclear-spin data qubits with fast, species-selective rubidium ancilla control and readout. However, realizing interspecies gates without inducing destructive Stark mixing in divalent atoms remains an outstanding problem. Here, we identify an optically accessible $S{+}S\leftrightarrow P{+}P$ F\"orster resonance at zero electric field, providing strong dipole-dipole exchange at array pitch. Using a shaped optical pulse under finite control response, we demonstrate a $0.36\,\mu\mathrm{s}$ exchange-assisted controlled-$Z$ gate with an intrinsic fidelity of $99.91\%$, remaining above $99.85\%$ under bounded perturbations. We also identify an auxiliary repulsive van der Waals channel, providing a comprehensive toolbox for hybrid quantum processors.

quant-ph

Learning Quantum Matter through Attention in Complex Space

Magnetic many-electron wavefunctions require amplitude and phase to be optimized together. Whether a complex internal representation improves this variational search is a practical question for neural wavefunction design. We introduce Complex Psiformer for interacting electrons in a magnetic moir\'e continuum, combining complex hidden features and Hermitian-magnitude attention with magnetic boundary conditions and fermionic antisymmetry. After the same number of optimization steps, Complex Psiformer reaches lower energies than Real Psiformer in two finite supercells. Both Psiformers also improve on their respective neural Hartree-Fock references. Across five training seeds in the 25-cell system, the mean Complex advantage is 1.458 meV per electron, with a smaller observed spread. A separately trained two-electron Complex state has a smaller energy gap to a finite configuration interaction reference than its Real counterpart. In the Complex states, flux scans show nonmonotonic density correlations and weaker honeycomb mean-density modulation at higher flux, while connected fluctuations persist. Gauge invariant current maps provide a qualitative comparison of local circulation in the optimized states. These benchmarks support the combined architecture as a variational ansatz for studying energies and charge arrangements in finite magnetic systems.

cond-mat.str-el

Who You Are Adds Nothing Detectable to Where You Go Next: Sociodemographic Conditioning in LLM Next-Location Prediction

Large language models (LLMs) are increasingly used for individual next-location prediction, while sociodemographic conditioning is common in LLM-based travel simulation. Yet the incremental predictive value of sociodemographic attributes remains unclear. To directly test this contribution, sociodemographic records were linked with passively sensed mobility data from 5,000 Shenzhen residents to construct a closed-set benchmark in which models rank 100 candidate destinations. Each prediction instance is evaluated with and without age, gender, occupation and income, while holding mobility history, candidates and all other prompt content fixed. Results show that across four history lengths, the paired change in top-1 accuracy ranges from -0.8 to +0.5 percentage points, with no detectable gain from attributes. This result remains consistent when stay history is withheld, across alternative prediction times, in two additional LLMs and in a supervised reranker trained on the same benchmark. The null does not reflect a lack of model responsiveness to demographic information, as permuted attributes reduce LLM accuracy whereas correctly matched attributes do not improve it. A further asymmetry emerges in the reverse predictive direction, as pre-cut mobility trajectories recover income with an AUC of 0.708, while sociodemographic attributes contribute little to next-location prediction. Beyond demographic conditioning, candidate construction exerts a much larger influence on reported performance. Removing distance raises top-1 accuracy by 7.7 percentage points under proximity sampling but lowers it by 22.3 points under popularity sampling, with the reversal reproduced across all three LLMs. These results distinguish demographic association from incremental predictive usefulness and show that sampled next-location accuracy depends strongly on how candidate alternatives are constructed.

cs.CY

Private communication via zero-private-capacity quantum channels

Private communication over a noisy quantum channel requires reliable transmission to the receiver and secrecy from the environment. Whether two channels with zero private capacity can jointly enable private communication is a longstanding open problem in quantum information theory. Here we resolve this problem by exhibiting a four-level channel and a qubit erasure channel with half erasure probability, each with zero private capacity, whose joint use achieves more than 0.0001903 private bits per product use. The encoding gives the receiver a linear information gain with at most quadratic environmental leakage, enabling privacy through a fixed joint measurement and classical coding. This superactivation, impossible for independent classical memoryless wiretap channels, shows that a channel's private capacity alone does not determine its value for secure communication. The initial activation example was identified through interactions with large language models, and the result has been formalized in Lean 4.

quant-ph

Causal-Class Hierarchies in Coherence-Constrained Channel Transformation

Higher-order quantum transformations allow multiple channel uses to be combined through different causal architectures, from parallel and fixed-order sequential networks to general higher-order processes. Whether this causal freedom improves channel transformation when the higher-order operation is also constrained by a resource theory remains largely unexplored. We study this question in the dynamical resource theory of coherence using a unified semidefinite-programming framework. For two qubit amplitude-damping channels and the identity target, we prove a strict causal hierarchy at every nontrivial damping strength under both maximally incoherent superchannels (MISC) and dephasing-covariant incoherent superchannels (DISC). In contrast, mixed-Pauli channels admit a common teleportation simulation that transfers the channel dependence to Bell-diagonal program states prepared in parallel. The remaining processing can then be absorbed into a single quantum operational, so parallel, fixed-order sequential, and general higher-order strategies achieve the same optimal error for any target. These results identify free program-state parallelisation as a structural obstruction to causal enhancement.

quant-ph

Mask Forcing: Improving Autoregressive Video Diffusion Distillation via Dual-Noise Masking Rollout

Autoregressive (AR) video diffusion models have shown great potential in real-time video generation. Recent methods distill pretrained bidirectional video diffusion models into causal AR students through Distribution Matching Distillation (DMD), but the generated videos often suffer from over-saturation and over-smoothing issues, resulting in limited visual quality and realism. The key contributing factor is the mode-seeking behavior of the reverse KL objective in DMD, which can cause the student distribution to collapse onto only a few modes of the teacher distribution. To address this, we propose Mask Forcing, a Dual-Noise Masking Rollout strategy that perturbs the AR student self-rollout to mitigate mode collapse induced by reverse-KL mode seeking. The core idea is to inject cleaner signals into noisy rollout inputs via random masks along spatial and temporal axes during the self-rollout process of AR diffusion distillation. Such perturbations encourage the student rollouts to explore more regions of the teacher distribution, allowing DMD to provide learning signals beyond the modes already covered by the student. Moreover, the cleaner tokens act as denoising guidance for other noisier tokens, improving the intermediate rollout predictions and reducing error accumulation. Extensive experiments demonstrate that our method improves multiple AR video diffusion distillation methods with higher visual quality efficiently, without incorporating real video data or additional post-training stages.

cs.CV

Tensor network investigation of the monomer-dimer model on the honeycomb lattice

The monomer-dimer model is one of the most well-known unsolved lattice models. In this paper we study the monomer-dimer model on the honeycomb lattice using the tensor network method, in the case that the dimer and monomer activities are 1. The monomer-dimer configurations are exactly mapped into the ground states of the antiferromagnetic Ising model on the Kagom\'e lattice in the critical field $H_{\rm{ex}}=4J$, and the tensor network is constructed based on the local ground states of each Ising triangle. The VUMPS approach is employed to contract the tensor network, providing a high-precision result of the monomer-dimer problem. We also revisit the edge coloring problem on the honeycomb lattice and discuss its relationship to the monomer-dimer model. Finally we formulate the monomer-dimer problem in the language of the sixteen-vertex model, and discuss the non-integrability of the general monomer-dimer model and the integrability of the pure dimer model.

cond-mat.stat-mech

WIDE: Wildcard Inference with Dynamic Expansion for Cross-Modal Generative Retrieval

Generative retrieval has demonstrated significant success by unifying representation learning and search into a single sequence-to-sequence generation task. However, extending this paradigm to cross-modal retrieval reveals a critical challenge arising from the inherent information asymmetry across different modalities, such as the gap between concise text queries and dense visual candidates. This structural mismatch causes the autoregressive decoder to suffer from forced hallucination when generating identifiers via standard trie-constrained beam search, where the model is severely penalized for failing to guess fine-grained details absent from the query, allowing irrelevant candidates to hijack top rankings. To address this issue, we propose Wildcard Inference with Dynamic Expansion (WIDE). WIDE employs Adaptive Entropy Thresholding (AET) to calibrate layer-specific uncertainty boundaries offline. During the decoding generation phase, Asymmetry-aware Wildcard Decoding (AWD) detects semantic blind spots and emits wildcards instead of forced deterministic identifiers, dynamically expanding the search space without incurring log-probability penalties. Finally, Blind-Spot Re-ranking (BSR) evaluates the expanded candidate pool using a hybrid scoring mechanism that combines discrete generation confidence with continuous semantic similarity. Extensive experiments on the M-BEIR benchmark demonstrate that WIDE outperforms state-of-the-art generative retrieval methods, effectively suppressing forced hallucination while maintaining compact index structures.

cs.CV

Directional Optimal Sub-Gamma Scales for Infinitely Divisible Laws

Fixing the quadratic proxy in a sub-gamma bound at the true variance leaves a scale to optimize, and a two-sided infinitely divisible law generally requires different scales in the two directions. For a centered law with finite nonzero variance, we normalize its Kolmogorov canonical measure and multiply the resulting variable by an independent Beta(1,2) variable. The signed remainder obtained in this way gives exact variational formulas for the right and left scales. We prove that a directional scale vanishes exactly when the Levy measure has no jumps in that direction, establish reflection, scaling, convolution, Levy-time, and opposite-jump perturbation rules, and recover the Levy triplet from the remainder law. The formulas give the two Gamma scales for bilateral Gamma laws and, for centered Skellam laws, the exact transition points p_+=(2+sqrt(3))/4 and p_-=(2-sqrt(3))/4 between local and interior control; when positive jumps are rare, the right scale is asymptotic to 1/log(1/p). The pair therefore records jump direction and the mechanism that controls the variance-exact sub-gamma pole.

math.PR

Minimal Radial Sub-Gamma Envelopes for Infinitely Divisible Random Vectors

Let X be a centered infinitely divisible random vector with finite second moment and covariance matrix Sigma. We define the radial pole C_X(t) as the smallest scale in a right sub-gamma bound for whose quadratic proxy is fixed at the true variance t^T Sigma t. A canonical directional measure and an independent Beta(1,2) multiplier give an exact variational formula for C_X. The resulting extended-valued function is positive homogeneous and is pointwise least among all homogeneous denominators compatible with the covariance quadratic form. We prove linear-map, convolution, and Levy-time rules, and show that the full family of directional remainders determines the law of X. Geometrically, C_X lies between the Minkowski functional of the moment-generating-function domain and one third of the positive support function of the Levy measure; the upper constant is sharp, and the zero set is a polar cone. The pole need not be subadditive. It is continuous on the sphere under global exponential moments and positive-definite covariance, whereas finite variance alone permits a jump from zero to infinity in nearby directions. For additive gamma-ray models, C_X equals the domain gauge and has a finite-polytope formula.

math.PR

SolarWM: Open Data and Scalable Training for Long-Horizon Video World Models

We introduce SolarWM, a fully open foundation for building interactive video world models from data preparation through long-horizon inference. Training across heterogeneous data sources and video backbones is challenging: datasets differ in temporal scale, camera geometry, visual quality, motion, and captioning styles, while video generators use distinct representations and architectures. Naive data mixing and model-specific implementations therefore produce inconsistent supervision and make results difficult to reproduce and compare. SolarWM addresses this coupling with a reconfigurable multi-source data engine and a backbone-native adaptation framework. The engine converts 1.43 million canonical clips from 10 datasets into a unified, frame-aligned contract covering visual observations, metric camera geometry, captions, quality metadata, selection decisions, and provenance, while decoupling source processing from mixture construction. Under shared camera-conditioning, training, and inference interfaces, we instantiate four 5B--33B models based on Wan2.2, LTX-2.5, and MiniMax-H3 while preserving their native representations and objectives. A unified three-stage recipe combines bidirectional adaptation, teacher-forced autoregressive initialization, and distribution matching distillation. The resulting causal models enable real-time interaction over rollouts ranging from minutes to hours after being trained on only 5s sequences. By releasing the resulting data, pipeline, recipes, weights, and framework, SolarWM provides a reproducible and extensible foundation for interactive world-model research.

cs.CV

VPID: An Integrated Framework for Vulnerability Prioritization and Intrusion Detection in Enterprise Networks

Small enterprises face increasingly serious threats to their internal networks but often lack the financial resources, computing capacity, and specialist staff required to deploy resource intensive security platforms. This paper designs and implements VPID, a lightweight framework for vulnerability prioritization and intrusion detection that consists of two principal modules: controlled vulnerability validation and intelligent intrusion defense. The first module uses OpenVAS for asset mapping and vulnerability identification, applies a decision tree to prioritize vulnerabilities, and employs a rule engine to generate targeted validation payloads. The second module captures network traffic using Scapy, analyzes it through a detection pipeline that combines a decision tree with multinomial Naive Bayes, verifies traffic assessed as high risk using Snort rules, and performs blocking and alerting through iptables. The evaluation uses 550,000 network flow samples containing normal and attack traffic for detector training, together with 15,000 labeled vulnerability records. On the vulnerability ranking test set, the decision tree achieves a precision of 91.8%, a recall of 89.5%, and an F1 score of 90.6%. On an independent test set containing 55,000 traffic samples, the combined detection pipeline achieves a precision of 94.5%, a recall of 88.3%, and an F1 score of 91.3%, while maintaining a false positive rate below 1.5%.

cs.CE

On-the-Fly3R: Towards Robust Online 3D Reconstruction with Feed-Forward 3R Models for Large-Scale UAV Scenarios

While feed-forward 3D reconstruction (3R) offers efficient end-to-end modeling, its application in large-scale UAV mapping is hindered by the prohibitive memory cost of Transformer attention. Current scalable streaming 3R methods assume temporally and spatially continuous inputs, rendering them ineffective for the weakly ordered or unordered image streams common in cross-strip UAV operations. To address this, we propose On-the-Fly3R, a training-free, progressive online 3D reconstruction framework for large-scale UAV images that upgrades various 3R backbones for large-scale UAV scenarios. Our method enables reconstruction from unordered inputs via retrieval-guided dynamic subset construction, which adaptively selects spatially relevant images. To further improve the robustness, a validation-rejection-retry mechanism is designed to guarantee global consistency, performing a pre-integration consistency check and automatically rejecting misaligned images and retrying with alternative subset. Finally, inspired by VSLAM, pose graph optimization based on the retrieval loop closure is employed to mitigate camera drift. Evaluations on several UAV benchmarks show that our On-the-Fly3R successfully scales various 3R models to over 5,000 images across square-kilometer UAV scenes, delivering substantially superior accuracy compared to several SOTA streaming 3R methods. Code is available at https://github.com/Sh1nZzz/On_the_Fly3R

cs.CV

Exact Virtual Channel Programming with Vanishing Excess Overhead

A finite-dimensional physical processor cannot exactly program a continuous family of distinct unitary channels. We show that this obstruction becomes quantitative when the target channel is stored in a normalized Choi state and its output observables are reconstructed by sampling physical channels and classically post-processing their measurement outcomes. For arbitrary $d$-dimensional channels, we construct a target-independent exact reconstruction protocol and prove the optimal one-copy sampling overhead, which grows quadratically with system dimension. We further prove the sharp fixed-$d$ law that the excess overhead vanishes inversely with the number of identical Choi programs. The upper bound combines deterministic port-based teleportation with a quasi-decomposition that corrects its depolarizing distortion. The converse maps any low-overhead reconstruction protocol to a physical learner of unknown unitaries and uses local quantum estimation to recover the same leading coefficient. These results recast the universal no-programming obstruction as a quantitative trade-off between quantum program memory and classical sampling, with a leading cost that reflects the locally learnable unitary degrees of freedom.

quant-ph

Textual Acoustic Grounding for Generalizable LLM-Based Deepfake Voice Detection

Deepfake voice detection suffers from poor generalization across unseen domains. While Audio Large Language Models (ALLMs) show promise, the modality gap between continuous audio embeddings which capture the subtle acoustic details necessary for deepfake detection and the semantic space of LLMs remains a critical, underexplored bottleneck. We address this by benchmarking diverse audio encoders integrated with Qwen LLMs (0.5B to 7B parameters). First, we demonstrate that fine-tuning the LLM alone risks out-of-domain overfitting, making a frozen LLM a stronger, resource-efficient baseline. Second, to explicitly bridge the modality gap, we introduce a cross-modal prompting strategy that injects linguistic-knowledge-driven acoustic features (via openSMILE) as structured text tokens. This explicit textual grounding not only enhances the frozen baseline but also makes LLM fine-tuning more effective. Ultimately, our approach demonstrates state-of-the-art resilience on the out-of-domain ITW and MLAAD benchmarks, yielding over \textbf{16.2\%} absolute improvement in Macro-F1 over existing ALLM baselines while maintaining competitive in-domain performance. All models reported in this work are \href{https://huggingface.co/01Yassine/AudioLLM-Deepfake-Detection}{publicly available}.

cs.SD

Deterministic Minimum-Output-Entropy Nonadditivity via Haagerup's Inequality and Near-Free Permutation Representations

We give a deterministic realization of the finite-dimensional quadratic certificate underlying Collins's mixed-unitary proof of minimum-output-entropy nonadditivity. For every fixed integer $K\ge 2$ and rational $\eta>0$ satisfying $\log K>2(3+\eta)^2$, a deterministic polynomial-time algorithm, for every sufficiently large target size $N$, outputs $K$ permutations on $N'=N+o_{K,\eta}(N)$ points. Restricting their permutation matrices to the nontrivial standard representation yields real orthogonal Stinespring blocks and a channel $\Phi_{N'}:M_{N'-1}(\mathbb{C})\to M_K(\mathbb{C})$ such that \[ 2H_{\min}(\Phi_{N'}) -H_{\min}(\Phi_{N'}^{\otimes 2}) \ge \frac{\log K}{K} -2\log\left(1+\frac{(3+\eta)^2}{K}\right) >0. \] The construction combines Haagerup's length-two inequality with the simultaneous deterministic spectral approximation of O'Donnell and Wu. We further show that the constant $3$ is asymptotically sharp on the relevant Hermitian zero-diagonal coefficient class and that the finite spectral transfer is nearly saturated, thereby isolating the finer geometry of the full output body as the natural next level of refinement beyond the scalar-radius method. Finally, a standard covariant extension converts the same deterministic entropy gap exactly into self-tensor superadditivity of the one-shot Holevo quantity.

quant-ph