SearcharxivSearch

arXiv subjects

Weijun Zeng

Publications and source records attributed to Weijun Zeng.

11 recordsLinked to original sources

CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence

Multimodal Large Language Models (MLLMs) have significantly advanced document understanding, yet current Doc-VQA evaluations score only the final answer and leave the supporting evidence unchecked. This answer-only approach masks a critical failure mode: a model can land on the correct answer while grounding it in the wrong passage -- a critical risk in high-stakes domains like law, finance, and medicine, where every conclusion must be traceable to a specific source region. To address this, we introduce CiteVQA, a benchmark that requires models to return element-level bounding-box citations alongside each answer, evaluating both jointly. CiteVQA comprises 1,897 questions across 711 PDFs spanning seven domains and two languages, averaging 40.6 pages per document. To ensure fidelity and scalability, the ground-truth citations are generated by an automated pipeline-which identifies crucial evidence via masking ablation-and are subsequently validated through expert review. At the core of our evaluation is Strict Attributed Accuracy (SAA), which credits a prediction only when the answer and the cited region are both correct. Auditing 20 MLLMs reveals a pervasive Attribution Hallucination: models frequently produce the right answer while citing the wrong region. The strongest system (Gemini-3.1-Pro-Preview) achieves an SAA of only 76.0, and the strongest open-source MLLM reaches just 22.5. Ultimately, towards trustworthy document intelligence, CiteVQA exposes a reliability gap that answer-only evaluations overlook, providing the instrumentation needed to close it. Our repository is available at https://github.com/opendatalab/CiteVQA.

cs.CL

MinerU2.5-Pro: Pushing the Limits of Data-Centric Document Parsing at Scale

Current document parsing methods advance primarily through model architecture innovation, while systematic engineering of training data remains underexplored. Yet state-of-the-art models spanning diverse architectures and parameter scales exhibit highly consistent failure patterns on the same set of hard samples, suggesting that the performance bottleneck stems from shared deficiencies in training data rather than from architectural differences. Building on this finding, we present MinerU2.5-Pro, which advances the state of the art purely through data engineering and training strategy design while retaining the 1.2B-parameter architecture of MinerU2.5 unchanged. At its core is a Data Engine co-designed around coverage, informativeness, and annotation accuracy: Diversity-and-Difficulty-Aware Sampling expands training data from under 10M to 65.5M samples while mitigating distribution shift; Cross-Model Consistency Verification leverages output consensus among heterogeneous models to assess sample difficulty and generate reliable annotations; the Judge-and-Refine pipeline improves annotation quality for hard samples through render-then-verify iterative correction. A three-stage progressive training strategy--large-scale pre-training, hard sample fine-tuning, and GRPO alignment--sequentially exploits these data at different quality tiers. On the evaluation front, we rectify element-matching biases in OmniDocBench v1.5 and introduce a Hard subset, establishing the more discriminative OmniDocBench v1.6 protocol. Without any architectural modification, MinerU2.5-Pro achieves 95.69 on OmniDocBench v1.6, improving over the same-architecture baseline by 2.71 points and surpassing all existing methods, including those based on models with over 200x more parameters.

cs.CV

MinerU-Diffusion: Rethinking Document OCR as Inverse Rendering via Diffusion Decoding

Optical character recognition (OCR) has evolved from line-level transcription to structured document parsing, requiring models to recover long-form sequences containing layout, tables, and formulas. Despite recent advances in vision-language models, most existing systems rely on autoregressive decoding, which introduces sequential latency and amplifies error propagation in long documents. In this work, we revisit document OCR from an inverse rendering perspective, arguing that left-to-right causal generation is an artifact of serialization rather than an intrinsic property of the task. Motivated by this insight, we propose MinerU-Diffusion, a unified diffusion-based framework that replaces autoregressive sequential decoding with parallel diffusion denoising under visual conditioning. MinerU-Diffusion employs a block-wise diffusion decoder and an uncertainty-driven curriculum learning strategy to enable stable training and efficient long-sequence inference. Extensive experiments demonstrate that MinerU-Diffusion consistently improves robustness while achieving up to 3.2x faster decoding compared to autoregressive baselines. Evaluations on the proposed Semantic Shuffle benchmark further confirm its reduced dependence on linguistic priors and stronger visual OCR capability.

cs.CV

GFM4GA: Graph Foundation Model for Group Anomaly Detection

Group anomaly detection is crucial in many network applications, but faces challenges due to diverse anomaly patterns. Motivated by the success of large language models (LLMs) in natural language processing, graph foundation models (GFMs) is proposed to handle few-shot learning task with fewer labeling efforts. GFMs have been successfully applied to detection of individual anomalies but cannot be generalized to group anomalies, as group anomaly patterns must be detected as a whole and individuals in an abnormal group can look rather normal. Therefore, we propose GFM4GA, a novel graph foundation model for group anomaly detection. The pipeline is pretrained via dual-level contrastive learning based on feature-based estimation and group extraction, to capture potential group anomaly structure and feature inconsistencies. In the downstream tasks, the pipeline is finetuned in parameter-constrained and group-anomaly-proportion weighted few-shot settings, and its adaptive ability to unseen group anomalies expanded via group contexts determined by labeled anomaly neighbors. Experiments show that GFM4GA surpasses group anomaly detectors and GFMs for individual anomalies, achieving average improvements of 2.85% in AUROC and 2.55% in AUPRC.

cs.AI

The RoboSense Challenge: Sense Anything, Navigate Anywhere, Adapt Across Platforms

Autonomous systems are increasingly deployed in open and dynamic environments -- from city streets to aerial and indoor spaces -- where perception models must remain reliable under sensor noise, environmental variation, and platform shifts. However, even state-of-the-art methods often degrade under unseen conditions, highlighting the need for robust and generalizable robot sensing. The RoboSense 2025 Challenge is designed to advance robustness and adaptability in robot perception across diverse sensing scenarios. It unifies five complementary research tracks spanning language-grounded decision making, socially compliant navigation, sensor configuration generalization, cross-view and cross-modal correspondence, and cross-platform 3D perception. Together, these tasks form a comprehensive benchmark for evaluating real-world sensing reliability under domain shifts, sensor failures, and platform discrepancies. RoboSense 2025 provides standardized datasets, baseline models, and unified evaluation protocols, enabling large-scale and reproducible comparison of robust perception methods. The challenge attracted 143 teams from 85 institutions across 16 countries, reflecting broad community engagement. By consolidating insights from 23 winning solutions, this report highlights emerging methodological trends, shared design principles, and open challenges across all tracks, marking a step toward building robots that can sense reliably, act robustly, and adapt across platforms in real-world environments.

cs.RO

1/f frequency fluctuations due to kinetic inductance in CoSi$_2$ microwave cavities

Cobalt disilicide provides a promising nearly-epitaxial superconducting material on silicon, which is compatible with high-density integrated circuit technology. We have characterized CoSi$_{2}$ superconducting microwave cavities around 5.5 GHz for resonance frequency fluctuations at temperatures 10 - 200 mK. We found relatively weak fluctuations $(δf/f)^2$ following the spectral density $A/f^γ $, with $A \simeq 6 \times 10^{-16}$ and $γ$ slightly below 1 at an average number of photons of $10^4$; the noise decreased with measurement power as $1/P^{1/2}$. We identify the noise as arising from kinetic inductance fluctuations and discuss possible origins of such fluctuations.

cond-mat.supr-con

Kinetic inductance in superconducting CoSi$_2$ coplanar microwave transmission lines

We have looked into cobalt disilicide (CoSi$_2$) as a potential building block for superconducting quantum circuits. In order to achieve this, we annealed a thin layer of Co to create 10-105 nm thick microwave cavities from CoSi$_2$ embedded in the silicon substrate. The cavity properties were measured as a function of temperature and power. In films measuring 10 and 25 nm, we find a significant kinetic inductance $L_\mathrm{K}$ with a non-BCS power-law variation $δL_\mathrm{K} \propto T^{4.3 \pm 0.2}$ at low temperatures. The quality factor of the studied microwave resonances increased almost linearly with thickness, with two-level systems having very little effect. The power dependence of kinetic inductance was analyzed in terms of heat flow due to electron-phonon coupling, which was found stronger than estimated for heat relaxation by regular quasiparticles.

cond-mat.supr-con

Suppression of $1/f$ noise in graphene due to non-scalar mobility fluctuations induced by impurity motion

Low frequency resistance variations due to mobility fluctuations is one of the key factors of $1/f$ noise in metallic conductors. According to theory, such noise in a two-dimensional (2D) device can be suppressed to zero at small magnetic fields, implying important technological benefits for low noise 2D devices. In this work, we provide direct evidence of anisotropic mobility fluctuations by demonstrating a strong field-induced suppression of noise in a high-mobility graphene Corbino disk, even though the device displays only a tiny amount of $1/f$ noise inherently. The suppression of the $1/f$ noise depends on charge density, showing less non-uniform mobility fluctuations away from the Dirac point with charge puddles. We model our results using a new approach based on impurity clustering dynamics and find our results consistent with the $1/f$ noise induced by scattering of carriers on mobile impurities forming clusters.

cond-mat.mes-hall

Electrical low-frequency $1/f^γ$ noise due to surface diffusion of scatterers on an ultra low noise graphene platform

Low-frequency $1/f^γ$ noise is ubiquitous, even in high-end electronic devices. For qubits such noise results in decrease of their coherence times. Recently, it was found that adsorbed O$_2$ molecules provide the dominant contribution to flux noise in superconducting quantum interference devices. To clarify the basic principles of such adsorbant noise, we have investigated the formation of low-frequency noise while the mobility of surface adsorbants is varied by temperature. In our experiments, we measured low-frequency current noise in suspended monolayer graphene samples under the influence of adsorbed Ne atoms. Owing to the extremely small intrinsic noise of graphene in suspended Corbino geometry, we could resolve a combination of $1/f^γ$ and Lorentzian noise spectra induced by the presence of Ne. We find that the $1/f^γ$ noise is caused by surface diffusion of Ne atoms and by temporary formation of few-Ne-atom clusters. Our results support the idea that clustering dynamics of defects is relevant for understanding of $1/f$ noise in general metallic systems.

cond-mat.mes-hall

Turbo DPSK in Bi-directional Relaying

In this paper, iterative differential phase-shift keying (DPSK) demodulation and channel decoding scheme is investigated for the Joint Channel decoding and physical layer Network Coding (JCNC) approach in two-way relaying systems. The Bahl, Cocke, Jelinek, and Raviv (BCJR) algorithm for both coherent and noncoherent detection is derived for soft-in soft-out decoding of DPSK signalling over the two-user multiple-access channel with Rayleigh fading. Then, we propose a pragmatic approach with the JCNC scheme for iteratively exploiting the extrinsic information of the outer code. With coherent detection, we show that DPSK can be well concatenated with simple convolutional codes to achieve excellent coding gain just like in traditional point-to-point communication scenarios. The proposed noncoherent detection, which essentially requires that the channel keeps constant over two consecutive symbols, can work without explicit channel estimation. Simulation results show that the iterative processing converges very fast and most of the coding gain is obtained within two iterations.

cs.IT

Joint Network and LDPC Coding for Bi-directional Relaying

In this paper, we consider joint network and LDPC coding for practically implementing the denosie-and-forward protocol over bi-directional relaying. the closed-form expressions for computing the log-likelihood ratios of the network-coded codewords have been derived for both real and complex multiple-access channels. It is revealed that the equivalent channel observed at the relay is an asymmetrical channel, where the channel input is the XOR form of the two source nodes.

cs.IT