Searcharxiv⌕ Search

arXiv subjects

Ming Yan

Publications and source records attributed to Ming Yan.

At least 55 records · Page 3Linked to original sources

MPDocBench-Parse: Benchmarking Practical Multi-page Document Parsing

Document parsing converts visually rich documents into machine-readable structured representations, forming a crucial foundation for information systems. Although many benchmarks have been proposed for document parsing, they remain inadequate for realistic scenarios. Existing benchmarks either focus on specific tasks or assess only single-page, text-centric settings, making them insufficient for practical multi-page parsing. Moreover, they lack fine-grained evaluation of semantic continuity, hierarchical structure recovery, and visual content preservation. To address these gaps, we propose MPDocBench-Parse, a benchmark for multi-page document parsing in real-world applications. It contains 433 manually annotated documents with 3,246 pages, covering 15 document types in English and Chinese, with diverse layout styles, and supports document-level end-to-end evaluation. We further design a comprehensive protocol for content fidelity and logical structure, covering text, table, and formula recognition, truncated text and table merging, figure extraction, reading order, and heading hierarchy recovery. Experiments show that, while existing models perform well on basic text extraction, they still suffer clear limitations in semantic continuity integration, visual content parsing, and hierarchical structure recovery. MPDocBench-Parse provides a unified foundation for advancing document parsing toward more realistic scenarios.

cs.AI↗

STAMP: Training Explicit Memory for Mobile GUI Agents in Controllable and Scalable Virtual Environments

Mobile GUI agents excel at immediate reactive control but frequently fail in realistic, long-horizon tasks that require memory. This failure stems from a fundamental conflict between limited context windows and token-heavy screenshots. To save the limited context, agents must progressively discard older visual history, permanently losing crucial transient information. Furthermore, existing action-centric datasets fail to teach agents what or when to explicitly memorize, and augmenting static real-world data is prohibitively expensive and lacks interactive verification. To resolve this, we present STAMP, a framework that trains explicit memory in mobile agents through controllable virtual environments, where deterministic memory variables are programmatically injected into synthesized tasks to control what must be memorized, when it should be encoded, and when it must later be retrieved, thereby producing verifiable supervised data at scale and enabling online reinforcement learning through environment-driven reward feedback. Evaluated on our newly introduced Memory-World benchmark, the resulting Stamp-GUI agent achieves state-of-the-art performance among GUI-specialized models and sets a new high watermark on our Memory-World benchmark, demonstrating exceptional memory accuracy and task resilience while maintaining strong general mobile navigation capabilities.

cs.CL↗

Single-photon time-stretch infrared spectroscopy

Sensitive mid-infrared (MIR) spectroscopy is highly demanded in various fields ranging from industrial inspection, biomedical diagnosis to astronomical observation. However, the detection sensitivity of conventional MIR spectrometers has been severely limited by excessive noises for existing infrared sensors, which hinders widespread use in photon-scarce scenarios. Here, we devise and implement a broadband MIR single-photon time-stretch spectrometer based on high-fidelity spectral upconversion and time-correlated coincidence counting. Specifically, a nanophotonic supercontinuum illumination covering 2.4-4.2 $μ$m is nonlinearly converted to the near-infrared band, where low-loss single-mode fiber and high-performance silicon detector can be leveraged to facilitate dispersive operation and sensitive detection, respectively. The arrival time for the dispersed upconversion photons is precisely registered with a low-timing-jitter photon counter, which enables us to obtain a high spectral resolution about 0.5 cm$^{-1}$ under a low-light-level illumination down to 0.14 photons/nm/pulse. In comparison to previous MIR upconversion spectrometers, the presented time-stretch architecture favors single-pixel simplicity and high-throughput acquisition for the single-photon spectral measurement. The achieved MIR spectroscopic features of broadband spectral coverage, sub-wavenumber resolution, single-photon sensitivity, and room-temperature operation would stimulate immediate applications in material and life sciences.

physics.optics↗

Wide-field mid-infrared hyperspectral imaging beyond video rate

Mid-infrared hyperspectral imaging has become an indispensable tool to spatially resolve chemical information in a wide variety of samples. However, acquiring three-dimensional data cubes is typically time-consuming due to the limited speed of raster scanning or wavelength tuning, which impedes real-time visualization with high spatial definition across broad spectral bands. Here, we devise and implement a high-speed, wide-field mid-infrared hyperspectral imaging system relying on broadband parametric upconversion of high-brightness supercontinuum illumination at the Fourier plane. The upconverted replica is spectrally decomposed by a rapid acousto-optic tunable filter, which records high-definition monochromatic images at a frame rate of 10 kHz based on a megapixel silicon camera. Consequently, the hyperspectral imager allows us to acquire 100 spectral bands over 2600-4085 cm$^{-1}$ in 10 ms, corresponding to a refreshing rate of 100 Hz. Moreover, the angular dependence of phase matching in the image upconversion is leveraged to realize snapshot operation with spatial multiplexing for multiple spectral channels, which may further boost the spectral imaging rate. The high acquisition rate, wide-field operation, and broadband spectral coverage could open new possibilities for high-throughput characterization of transient processes in material and life sciences.

physics.optics↗

High-resolution mid-infrared single-photon upconversion ranging

Single-photon laser ranging has widespread applications in remote sensing and target recognition. However, highly-sensitive light detection and ranging (LiDAR) has long been restricted in visible or near-infrared bands. An appealing quest is to extend the operation wavelength into the mid-infrared (MIR) region, which calls for an infrared photon counting system at high detection sensitivity and precise temporal resolution. Here, we devise and demonstrate a MIR upconversion LiDAR based on nonlinear asynchronous optical sampling. Specifically, the infrared probe is interrogated in a nonlinear crystal by a train of pump pulses at a slightly different repetition rate, which favors for a temporal optical scanning at a picosecond timing resolution and a kilohertz refreshing rate over $\sim$50 ns. Moreover, the cross-correlation upconversion trace is temporally stretched by a factor of 2$\times$10$^4$, which can thus be recorded by a low-bandwidth silicon detector. In combination with time-correlated photon-counting technique, the achieved effective resolution is about two orders of magnitude better than the timing jitter of the detector itself, which facilitates a ranging precision of 4 $μ$m under a low detected flux of 8$\times$10$^{-5}$ photons per pulse. The presented MIR time-of-flight range finder is featured with single-photon sensitivity and high positioning resolution, which would be particularly useful in infrared sensing and imaging in photon-starved scenarios.

physics.optics↗

Tongyi DeepResearch Technical Report

We present Tongyi DeepResearch, an agentic large language model, which is specifically designed for long-horizon, deep information-seeking research tasks. To incentivize autonomous deep research agency, Tongyi DeepResearch is developed through an end-to-end training framework that combines agentic mid-training and agentic post-training, enabling scalable reasoning and information seeking across complex tasks. We design a highly scalable data synthesis pipeline that is fully automatic, without relying on costly human annotation, and empowers all training stages. By constructing customized environments for each stage, our system enables stable and consistent interactions throughout. Tongyi DeepResearch, featuring 30.5 billion total parameters, with only 3.3 billion activated per token, achieves state-of-the-art performance across a range of agentic deep research benchmarks, including Humanity's Last Exam, BrowseComp, BrowseComp-ZH, WebWalkerQA, xbench-DeepSearch, FRAMES and xbench-DeepSearch-2510. We open-source the model, framework, and complete solutions to empower the community.

cs.CL↗

CorpusQA: A 10 Million Token Benchmark for Corpus-Level Analysis and Reasoning

While large language models now handle million-token contexts, their capacity for reasoning across entire document repositories remains largely untested. Existing benchmarks are inadequate, as they are mostly limited to single long texts or rely on a "sparse retrieval" assumption-that answers can be derived from a few relevant chunks. This assumption fails for true corpus-level analysis, where evidence is highly dispersed across hundreds of documents and answers require global integration, comparison, and statistical aggregation. To address this critical gap, we introduce CorpusQA, a new benchmark scaling up to 10 million tokens, generated via a novel data synthesis framework. By decoupling reasoning from textual representation, this framework creates complex, computation-intensive queries with programmatically guaranteed ground-truth answers, challenging systems to perform holistic reasoning over vast, unstructured text without relying on fallible human annotation. We further demonstrate the utility of our framework beyond evaluation, showing that fine-tuning on our synthesized data effectively enhances an LLM's general long-context reasoning capabilities. Extensive experiments reveal that even state-of-the-art long-context LLMs struggle as input length increases, and standard retrieval-augmented generation systems collapse entirely. Our findings indicate that memory-augmented agentic architectures offer a more robust alternative, suggesting a critical shift is needed from simply extending context windows to developing advanced architectures for global information synthesis.

cs.CL↗

High-speed hyperspectral 3D ghost imaging LiDAR

Light detection and ranging (LiDAR) is widely used in autonomous systems and industrial metrology; however, the simultaneous acquisition of three-dimensional (3D) structure and broadband spectral information remains challenging, as conventional hyperspectral LiDAR relies on wavelength-scanning or spectrometer-based detection that limits speed. Here, we demonstrate a hyperspectral 3D ghost imaging LiDAR that eliminates these bottlenecks. By combining a stochastic broadband laser with single-pixel detection, and integrating spatiotemporal encoding with spectral ghost imaging in a time-of-flight framework, the system enables pulse-resolved recovery of spatial and spectral information. Consequently, we achieve a line-scanning rate of 60.5 MHz (point rate 1.8 GHz) and a ranging precision of 0.02 mm within a 10 μs integration time. Each voxel contains a 1.4 nm resolution spectrum over 1100-1250 nm, enabling simultaneous 3D imaging and chemical identification. This approach provides a route to high-speed hyperspectral LiDAR for environmental monitoring, precision agriculture, and industrial inspection.

physics.optics↗

Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement Learning

Recent advances in Large Language Models(LLMs) have enabled strong performance in long-form writing, but current training paradigms remain limited: Supervised Fine-Tuning (SFT) remains constrained by data saturation and performance ceilings, while Reinforcement Learning with Verifiable Reward (RLVR), though successful in verifiable domains like math and code, cannot be directly migrated to open-ended long-form writing due to a lack of ground-truths. To further advance long-form writing, we present Writing-RL: an Adaptive Curriculum Reinforcement Learning framework to advance long-form writing capabilities beyond SFT. The framework consists of three key components: Margin-aware Data Selection strategy that prioritizes samples with high learning potential, Pairwise Comparison Reward mechanism that provides discriminative learning signals in the absence of verifiable rewards, and Dynamic Reference Scheduling approach, which plays a critical role by adaptively adjusting task difficulty based on evolving model performance. Experiments on 7B-scale writer models show that Writing-RL effectively improves long-form writing performance over strong SFT baselines. Furthermore, we observe that models trained with long-output RL generalize surprisingly well to long-input reasoning tasks, potentially offering a promising perspective for rethinking long-context training.

cs.CL↗

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding

Video temporal understanding is crucial for multimodal large language models (MLLMs) to reason over events in videos. Despite recent advances in general video understanding, current MLLMs still struggle with fine-grained temporal reasoning. While reinforcement learning (RL) has been explored to address this issue recently, existing RL approaches remain limited in performance on time-sensitive tasks. In this work, we propose MUSEG, a novel RL-based method that enhances temporal understanding by introducing timestamp-aware multi-segment grounding. MUSEG enables MLLMs to align queries with multiple relevant video segments, promoting more comprehensive temporal reasoning. To facilitate effective learning, we design a customized RL training recipe with phased rewards that progressively guides the model toward temporally grounded reasoning. Extensive experiments on temporal grounding and time-sensitive video question answering (QA) tasks demonstrate that MUSEG significantly outperforms existing methods and generalizes well across diverse temporal understanding scenarios. View our project at https://github.com/THUNLP-MT/MUSEG.

cs.CV↗

Scaling External Knowledge Input Beyond Context Windows of LLMs via Multi-Agent Collaboration

With the rapid advancement of post-training techniques for reasoning and information seeking, large language models (LLMs) can incorporate a large quantity of retrieved knowledge to solve complex tasks. However, the limited context window of LLMs obstructs scaling the amount of external knowledge input, prohibiting further improvement. Existing context window extension methods inevitably cause information loss. LLM-based multi-agent methods emerge as a new paradigm to handle massive input in a distributional manner, where we identify two core bottlenecks in existing agent orchestration designs. In this work, we develop a multi-agent framework, \textbf{\ExtAgents}, to overcome the bottlenecks and enable better scalability in inference-time knowledge integration without longer-context training. Benchmarked with our enhanced multi-hop question answering test, \textbf{$\boldsymbol{\infty}$Bench+}, and other public test sets including long survey generation, \ExtAgents significantly enhances the performance over existing non-training methods with the same amount of external knowledge input, regardless of whether it falls \emph{within or exceeds the context window}. Moreover, the method maintains efficiency due to high parallelism. We believe further study in the coordination of LLM agents on increasing external knowledge input could benefit real-world applications.

cs.CL↗

Lightweight Geometric Adaptation for Training Physics-Informed Neural Networks

Physics-Informed Neural Networks (PINNs) often suffer from slow convergence, training instability, and reduced accuracy on challenging partial differential equations due to the anisotropic and rapidly varying geometry of their loss landscapes. We propose a lightweight curvature-aware optimization framework that augments existing first-order optimizers with an adaptive predictive correction based on secant information. Consecutive gradient differences are used as a cheap proxy for local geometric change, together with a step-normalized secant curvature indicator to control the correction strength. The framework is plug-and-play, computationally efficient, and broadly compatible with existing optimizers, without explicitly forming second-order matrices. Experiments on diverse PDE benchmarks show consistent improvements in convergence speed, training stability, and solution accuracy over standard optimizers and strong baselines, including on the high-dimensional heat equation, Gray--Scott system, Belousov--Zhabotinsky system, and 2D Kuramoto--Sivashinsky system.

cs.LG↗

Perception-Aware Policy Optimization for Multimodal Reasoning

Reinforcement Learning with Verifiable Rewards (RLVR) has proven to be a highly effective strategy for endowing Large Language Models (LLMs) with robust multi-step reasoning abilities. However, its design and optimizations remain tailored to purely textual domains, resulting in suboptimal performance when applied to multimodal reasoning tasks. In particular, we observe that a major source of error in current multimodal reasoning lies in the perception of visual inputs. To address this bottleneck, we propose PAPO, a novel policy gradient algorithm that encourages the model to learn to perceive while learning to reason. Specifically, we introduce the Implicit Perception Loss in the form of a KL divergence term, which can be seamlessly plugged into mainstream RLVR algorithms such as GRPO and DAPO. Notably, PAPO does not rely on additional data curation, reward models, or stronger teacher models. To further enhance the training stability of PAPO, we introduce the Double Entropy Loss, which effectively regularizes the new KL objective without compromising performance. Despite its simplicity, PAPO yields significant overall improvements of 4.4%-17.5% on diverse multimodal benchmarks. The improvements are more pronounced, approaching 8.0%-19.1%, on tasks with high vision dependency. We also observe a substantial reduction of 30.5% in perception errors, indicating improved perceptual capabilities with PAPO. Overall, our work introduces a deeper integration of perception-aware supervision into core learning objectives and lays the groundwork for a new RL framework that encourages visually grounded reasoning. Code and data will be made publicly available for research purposes. Project page: https://mikewangwzhl.github.io/PAPO.

cs.CL↗

R2-Write: Reflection and Revision for Open-Ended Writing with Deep Reasoning

While deep reasoning with long chain-of-thought has dramatically improved large language models in verifiable domains like mathematics, its effectiveness for open-ended tasks such as writing remains unexplored. In this paper, we conduct a systematic investigation revealing that existing mainstream reasoning models achieve limited gains on open-ended writing tasks. Our further analysis shows that these models lack deep reflection and revision patterns in open-ended writing, resulting in substantially smaller improvements compared to mathematical reasoning tasks. To address this limitation, we introduce R2-Write: an automated framework that synthesizes high-quality thinking trajectories enriched with explicit reflection and revision patterns through iterative writer-judge interaction. To prevent redundant reflections, we design a process reward mechanism that supervises reflection quality during reinforcement learning, improving both performance and token efficiency. Extensive experiments across multiple creative writing and deep-research benchmarks demonstrate significant improvements, validating that explicitly incorporating reflection and revision patterns unlocks deep reasoning capabilities for open-ended writing tasks.

cs.CL↗

Do Phone-Use Agents Respect Your Privacy?

We study whether phone-use agents respect privacy while completing benign mobile tasks. This question has remained hard to answer because privacy-compliant behavior is not operationalized for phone-use agents, and ordinary apps do not reveal exactly what data agents type into which form entries during execution. To make this question measurable, we introduce MyPhoneBench, a verifiable evaluation framework for privacy behavior in mobile agents. We operationalize privacy-respecting phone use as permissioned access, minimal disclosure, and user-controlled memory through a minimal privacy contract, iMy, and pair it with instrumented mock apps plus rule-based auditing that make unnecessary permission requests, deceptive re-disclosure, and unnecessary form filling observable and reproducible. Across five frontier models on 10 mobile apps and 300 tasks, we find that task success, privacy-compliant task completion, and later-session use of saved preferences are distinct capabilities, and no single model dominates all three. Evaluating success and privacy jointly reshuffles the model ordering relative to either metric alone. The most persistent failure mode across models is simple data minimization: agents still fill optional personal entries that the task does not require. These results show that privacy failures arise from over-helpful execution of benign tasks, and that success-only evaluation overestimates the deployment readiness of current phone-use agents. All code, mock apps, and agent trajectories are publicly available at~ https://github.com/FreedomIntelligence/MyPhoneBench.

cs.CR↗

FlashCap: Millisecond-Accurate Human Motion Capture via Flashing LEDs and Event-Based Vision

Precise motion timing (PMT) is crucial for swift motion analysis. A millisecond difference may determine victory or defeat in sports competitions. Despite substantial progress in human pose estimation (HPE), PMT remains largely overlooked by the HPE community due to the limited availability of high-temporal-resolution labeled datasets. Today, PMT is achieved using high-speed RGB cameras in specialized scenarios such as the Olympic Games; however, their high costs, light sensitivity, bandwidth, and computational complexity limit their feasibility for daily use. We developed FlashCap, the first flashing LED-based MoCap system for PMT. With FlashCap, we collect a millisecond-resolution human motion dataset, FlashMotion, comprising the event, RGB, LiDAR, and IMU modalities, and demonstrate its high quality through rigorous validation. To evaluate the merits of FlashMotion, we perform two tasks: precise motion timing and high-temporal-resolution HPE. For these tasks, we propose ResPose, a simple yet effective baseline that learns residual poses based on events and RGBs. Experimental results show that ResPose reduces pose estimation errors by ~40% and achieves millisecond-level timing accuracy, enabling new research opportunities. The dataset and code will be shared with the community.

cs.CV↗

SPRITE: From Static Mockups to Engine-Ready Game UI

Game UI implementation requires translating stylized mockups into interactive engine entities. However, current "Screenshot-to-Code" tools often struggle with the irregular geometries and deep visual hierarchies typical of game interfaces. To bridge this gap, we introduce SPRITE, a pipeline that transforms static screenshots into editable engine assets. By integrating Vision-Language Models (VLMs) with a structured YAML intermediate representation, SPRITE explicitly captures complex container relationships and non-rectangular layouts. We evaluated SPRITE against a curated Game UI benchmark and conducted expert reviews with professional developers to assess reconstruction fidelity and prototyping efficiency. Our findings demonstrate that SPRITE streamlines development by automating tedious coding and resolving complex nesting. By facilitating rapid in-engine iteration, SPRITE effectively blurs the boundaries between artistic design and technical implementation in game development. Project page: https://baiyunshu.github.io/sprite.github.io/

cs.HC↗

Inexact Bregman Sparse Newton Method for Efficient Optimal Transport

Computing exact Optimal Transport (OT) distances for large-scale datasets is computationally prohibitive. While entropy-regularized alternatives offer speed, they sacrifice precision and frequently suffer from numerical instability in high-accuracy regimes. To address these limitations, we propose the Inexact Bregman Sparse Newton (IBSN) method, which efficiently solves the exact OT problems. Our approach utilizes a Bregman proximal point framework through a sequence of semi-dual subproblems. By solving these subproblems inexactly, we significantly reduce per-iteration complexity while maintaining a theoretical guarantee of convergence to the true optimal plan. To further accelerate the algorithm, we develop a sparse Newton-type solver for the subproblem and employ a Hessian sparsification strategy that drastically lowers memory and time costs without sacrificing accuracy. We provide rigorous theoretical guarantees for the global convergence of the algorithm. Extensive experiments demonstrate that IBSN consistently outperforms state-of-the-art methods in both computational speed and solution precision.

math.OC↗