SearcharxivSearch

arXiv subjects

Wentao Liu

Publications and source records attributed to Wentao Liu.

At least 19 recordsLinked to original sources

Thinking with Cameras: Active Visual Reasoning via Dynamic Viewpoint Control for Surveillance Video Understanding

Large vision-language models (LVLMs) have recently achieved remarkable progress in general-purpose video understanding. However, their application to surveillance videos remains challenging due to the lack of large-scale domain-specific datasets and the limitation of passive observation from fixed viewpoints. In surveillance scenarios, critical visual evidence can be easily missed when targets are distant, small, occluded, or move beyond the current camera view. In this work, we introduce CamVLM, a new framework for Thinking with Cameras, which enables LVLMs to actively acquire visual evidence through dynamic viewpoint control rather than passively analyzing fixed video streams. We first construct CCTV-Anomaly, a large-scale surveillance video understanding dataset containing 14,459 videos across 10 anomaly categories, with detailed captions and event annotations. We further formulate viewpoint control as an active visual perception problem and build CamTrack-53K, an object-centric viewpoint trajectory dataset for learning camera actions. Moreover, we propose a reinforcement learning based viewpoint policy optimization framework, which models camera control as a sequential decision-making process and learns long-horizon observation strategies beyond supervised trajectory imitation. Extensive experiments demonstrate that CamVLM achieves state-of-the-art performance under both passive observation and dynamic viewpoint settings, validating the effectiveness of active camera-based reasoning for surveillance video understanding. Our datasets, model, and code will be available at https://github.com/xiaozhang79/CamVLM .

cs.CV

Escape of quantum information across an analogue black hole horizon

The complete evaporation of black holes, as a natural endpoint of Hawking radiation, gives rise to the black hole information paradox, which fundamentally challenges the principles of unitarity and information conservation in quantum mechanics. Although the AdS/CFT correspondence indicates that information is preserved during black hole evaporation, the precise mechanism by which it is recovered from the Hawking radiation remains an open question. To explore a potential resolution, we investigate information transfer in a spin-chain model featuring an effective horizon realized via position-dependent couplings within an XY spin chain. We derive and demonstrate Page curve-like behavior, and analyze the transmission of quantum resources, such as entanglement and coherence, across the effective horizon. Our results show that quantum resources initially localized within an interior subsystem can be transferred to the exterior via particle radiation through the horizon. This study provides a novel perspective from quantum simulation on how information may escape from black holes, thereby contributing to the further understanding of the black hole information paradox.

gr-qc

Beyond Direct Answering: Aligning Educational LLMs as Socratic Guides via Heuristic Reinforcement Learning

Large language models (LLMs) deployed in educational settings often behave as direct answerers: they disclose target concepts in the opening turn instead of guiding students through progressive inquiry, as Socratic pedagogy prescribes. We present HeuristicEdu, a two-phase pipeline that aligns Qwen2.5-7B toward Socratic tutoring via supervised warm-up and Group Relative Policy Optimization (GRPO). Training uses SocraticEdu, 797 multi-turn Chinese children's science dialogues reconstructed from a live platform, with a heuristic reward over cognitive depth (R_cog), curiosity engagement (R_eng), and directness (R_dir), together with a K_query correction for student-introduced terms. We introduce Scaffolding Effectiveness (SE) and Conversation Depth (CD) to evaluate outcomes beyond surface fluency. On 30 held-out questions, the best GRPO variant improves SE from 30.0% to 63.3% and lowers keyword leakage from 30.0% to 13.3%. Notably, this best variant omits the directness penalty during optimization, suggesting that explicit anti-leakage terms can conflict with gradient-based behavioral alignment. An unaligned Qwen-72B baseline reaches 0% SE and 96.7% leakage, showing that scale alone does not induce Socratic behavior.

cs.CL

Toward Anthropomorphic Dialogue: A Closed-Loop Framework for Human-Like Chat Generation, Evaluation, and Preference Alignment

Human-like private chat requires more than fluent response generation: a system must preserve persona, relationship, memory, bounded knowledge, medium-specific timing, and a coherent multi-turn arc. We present AnthroDial, a closed-loop framework that formulates anthropomorphic dialogue as a joint problem of system architecture, executable evaluation, and diagnostic alignment. It combines (1) a role-conditioned scheduled dialogue runtime with persona and scenario cards, long-term memory, virtual time, and single-draft message decisions; (2) an executable benchmark with an L0 validity gate, five per-turn dimensions, and five dialogue-level dimensions; and (3) a post-training pipeline that filters 16,436 scheduled-decision examples for SFT and applies GRPO with a cognitive-diagnostic, ZPD-aware reward. The reward maintains Kalman-filtered capability estimates for each behavioral dimension, upweights dimensions with larger capability deficits, and uses rollout scores as task-level ZPD matches to focus optimization on learnable weak skills. On a benchmark with 55 personas, 50 scenarios, 50 persona-scenario bindings, and 100 role-conditioned cases per model, we evaluate 16 systems spanning frontier baselines, open models, thinking/no-think variants, and SFT/RL ablations. The strongest non-trained baseline reaches 32.00% strict ACC, while Qwen3.6-27B-SFT+RL reaches 39.00% strict ACC and a 98.5 overall score. In the 9B no-think setting, SFT and RL improve strict ACC from 0.00% to 13.00% and 18.37%. These results show that anthropomorphic dialogue benefits when generation, evaluation, and reward shaping share the same behavioral dimensions.

cs.AI

Asymmetric quantum steering harvested near a Lorentz-violating BTZ black hole

We investigate the harvesting of quantum steering and its directional asymmetry between two Unruh-DeWitt detectors in a Lorentz-violating BTZ black hole spacetime. Since the detectors are located at different radial positions outside the black hole, they experience inequivalent local environments induced by gravitational redshift, causing Alice to undergo stronger effective thermal noise than Bob. Remarkably, we uncover a counterintuitive phenomenon in which the detector subjected to a higher effective temperature exhibits stronger steerability than the other one, revealing a nontrivial inversion of thermal intuition in curved spacetime. Furthermore, quantum steering survives only within a finite window of detector energy gaps and reaches its maximum within an optimal regime. We find that Lorentz violation suppresses steering most strongly near this optimal energy gap, indicating an enhanced sensitivity of maximal correlation extraction to symmetry breaking effects. Our results demonstrate that Lorentz violation acts as a geometric constraint on the quantum information capacity of spacetime, simultaneously restricting both the strength and the directionality of quantum correlations.

gr-qc

Empowering a Single-Frequency GNSS Receiver to Achieve High-Precision Positioning with Relative Observations

Global Navigation Satellite System (GNSS) navigation is widely used to provide absolute, outdoor positioning in field robotics. Advances in Real-Time Kinematic (RTK) technology can achieve centimeter-level accuracy, facilitating autonomous navigation tasks. However, the cost and extra infrastructure used for RTK still hinder the application and more cost-effective solutions are desired. In this letter, we present a novel tightly-coupled state estimation framework that achieves high-precision localization by using low-cost, mass-market single-frequency GNSS receivers with any relative motion sensors (e.g., wheel encoder, camera, LiDAR). We propose a sliding-window factor graph that integrates generic relative motion with global epoch-to-anchor constraints derived from continuous carrier phase tracking. To eliminate the reliance on physical base stations, we introduce a virtual anchor mechanism: upon the initial observation of a satellite, its state is locked as a virtual reference to establish global epoch-to-anchor constraints. By substituting multi-frequency hardware redundancy with single-frequency multi-modal kinematic priors and a robust cycle-slip recovery technique, our approach ensures carrier-phase integrity on cheap receivers. Extensive real-world experiments on heterogeneous low-cost sensor suites validate that our method improves the accuracy of a single-frequency receiver from several meters to decimeter-level precision across diverse environments, providing an accurate, cost-effective and reliable alternative for autonomous navigation.

cs.RO

Comparison Results for a class of Neumann Problems of the $p$-Laplace Equation on Riemannian Manifolds

We consider Neumann boundary value problems for the $p$-Laplace equation on Riemannian manifolds with nonnegative Ricci curvature. Using spherical symmetrization under appropriate constraints, we derive Talenti-type comparison results in Lorentz spaces. We further show that, in contrast to the Robin case, the Neumann setting admits weaker constraints, which yields stronger comparison principles.

math.AP

Elmes*: Automated Construction of Fine-Grained Evaluation Rubrics for Large Language Models in Long-Tail Educational Scenarios

Evaluating large language models (LLMs) for education requires measuring how models teach, not only what they know. Existing benchmarks emphasize domain-general correctness or depend on manually designed rubrics that scale poorly to long-tail pedagogical scenarios. We introduce Elmes*, an end-to-end framework for constructing, refining, and applying fine-grained scenario-specific rubrics. Elmes* combines a declarative multi-agent engine for teacher--student--judge interactions with SceneGen, a self-evolving module that co-optimizes evaluation criteria and test data from expert-defined pedagogical dimensions. Using Elmes*, we build Edu-330, covering 330 scenarios across 11 subjects, 3 grade bands, and 10 task types, with over 1{,}000 second-level indicators. Experiments on Edu-330 and four expert-authored gold-standard scenarios show that educational capability is multidimensional: top-tier LLMs differ mainly in creativity and values integration, knowledge-strong models may fail at Socratic scaffolding, and the education-specialized InnoSpark achieves the best human-evaluated average score. LLM judges preserve human-comparable rankings with much lower scoring variance, but exhibit judge-specific biases such as self-preference. Ablations show that expert-scored few-shot anchoring improves human--LLM alignment, while reasoning enforcement and greedy decoding are model-dependent. Elmes* thus provides scalable diagnostic infrastructure for pedagogically grounded LLM evaluation.

cs.LG

UniCAD: A Unified Benchmark and Universal Model for Multi-Modal Multi-Task CAD

Computer-Aided Design (CAD) underpins modern engineering and manufacturing by enabling the creation of precise, editable 3D models. However, CAD research typically studies tasks in isolation, and multi-modal, multi-task learning for CAD is hindered by the absence of a unified benchmark. To address this gap, we introduce UniCAD, a comprehensive benchmark for multi-modal CAD learning that covers point-to-CAD reconstruction, text/image-to-CAD generation, and CAD question answering across diverse input modalities. Alongside the benchmark, we present UniCAD-MLLM, a universal multi-modal large language model that ingests text, images, sketches, and point clouds and performs these heterogeneous tasks in an end-to-end fashion within a single framework. Extensive experiments on the UniCAD and Fusion360 benchmarks demonstrate that UniCAD-MLLM achieves state-of-the-art performance across all tasks, outperforming existing task-specific and multi-task baselines. We will release the dataset, code, and pretrained models to accelerate future research.

cs.CV

How Far Has AI Come in Liver Fibrosis Staging? A Large-Scale Real-World Dataset and Benchmark

Despite years of methodological progress, how far AI has come in liver fibrosis staging has never been systematically evaluated under the heterogeneous, multi-center conditions that define clinical practice. To address this gap, we introduce LiFS, a large-scale dataset and benchmark derived from the MICCAI 2025 CARE-Liver challenge, comprising 610 patients across multiple centers and scanners with multi-sequence MRI. To the best of our knowledge, LiFS is the first benchmark providing complete gadoxetic acid-enhanced sequences with histopathology-confirmed annotations from diverse real-world scanners. Through systematic evaluation of 9 independently developed methods selected from 96 registered teams against in-cohort radiologist reference results, our findings address how far current AI has progressed toward clinical-level liver fibrosis staging from three complementary perspectives. First, against radiologists, the best AI methods were broadly comparable to the senior radiologist and significantly exceeded the junior radiologist in selected settings, while median AI performance generally approached junior-radiologist levels. Second, from a data perspective, cross-center heterogeneity, label imbalance, and contrast-enhanced sequence variability emerge as the dominant challenges for AI methods. Third, from a technical perspective, methodological design choices, including spatial registration, input dimensionality, multi-modal fusion strategy, and backbone architecture, appear to modulate cross-center robustness, although no single choice alone closes the gap. Overall, LiFS provides a rigorous real-world benchmark for positioning the current state of AI in liver fibrosis staging and for enabling future research on the key challenges that limit clinically reliable deployment.

cs.CV

A Universal Framework for Horizon-Scale Tests of Gravity with Black Hole Shadows

In this Letter, we have developed a numerically efficient framework for evaluating parameters in metric theories of gravity, and applied it to constrain the horizon-scale magnetic field in the Kerr-Bertotti-Robinson (Kerr-BR) spacetime using the latest EHT observations. The method's adaptive ray-tracing strategy achieves near-linear computational efficiency without loss of numerical accuracy. Owing to this efficiency, the framework enables high precision shadow modeling at minimal computational cost and, for the first time, supports statistically robust inference of black hole parameters from horizon-scale observations for arbitrary stationary black holes. The above framework is applied to the recently obtained Kerr-BR black hole, an exact magnetized and rotating solution to the Einstein field equations. We have evaluated the horizon-scale magnetic fields of M87* and Sgr A*, with the latter showing a field strength of $93.3^{+14.7}_{-23.8}G$, consistent with the equipartition estimate of $71G$ from polarized ALMA observations, thereby supporting Einstein's gravity.

gr-qc

Gravitational-Bumblebee perturbations: Exact decoupling and isospectrality

In this paper, we present the exact decoupling of the full metric and bumblebee field perturbations in a Schwarzschild-like background. The coupled system reduces to four decoupled master equations, revealing in each parity sector a Schwarzschild-like gravitational sector and a Lorentz-violating Maxwell-like vector sector. While Lorentz violation modifies the propagation speed of the emergent vector modes, we demonstrate that the gravitational master modes exhibit a ``dynamical immunity'' to the non-minimal Lorentz-violating coupling, and that the odd- and even-parity perturbations remain strictly isospectral. Our work provides a rare example in which Lorentz-violating couplings reshape the field reconstruction while leaving the gravitational ringdown spectrum intact. This mismatch in propagation speeds suggests a possible timing signature of bumblebee vector dynamics in black hole perturbations, offering a theoretical route to testing spontaneous Lorentz symmetry breaking in the era of multi-messenger astronomy.

gr-qc

Decoupled Attention Fusion: Accelerating RAG with Efficient KV Cache Reuse

Retrieval-Augmented Generation (RAG) effectively mitigates hallucinations in Large Language Models (LLMs) but suffers from prohibitive Time-To-First-Token (TTFT) latency in long-context scenarios. Reusing pre-computed document KV caches addresses this but introduces a distribution mismatch, where offline caches lack the inter-document attention patterns required for coherent reasoning. CacheBlend reduces recomputation via selective attention, but suffers severe accuracy degradation at longer contexts. To address these challenges, we propose Decoupled Attention Fusion (DAF), a framework that maintains high accuracy while significantly reducing recomputation overhead. DAF decouples the attention process into three integrated stages: important-token self-attention to restore missing inter-document attention, question-document self-attention for standard inference, and a state fusion that concatenates their outputs to synthesize the final hidden states. By decoupling these operations into dense patterns, DAF is natively compatible with Flash-Attention kernels, maximizing hardware utilization without requiring complex attention masks. Experiments show that DAF delivers up to 2 times speedup over CacheBlend and 5.6 times over full recomputation with vLLM on long-context benchmarks, without sacrificing accuracy.

cs.PF

ATRS: Adaptive Trajectory Re-splitting via a Shared Neural Policy for Parallel Optimization

Parallel trajectory optimization via the Alternating Direction Method of Multipliers (ADMM) has emerged as a scalable approach to long-horizon motion planning. However, existing frameworks typically decompose the problem into parallel subproblems based on a predefined fixed structure. Such structural rigidity often causes optimization stagnation in highly constrained regions, where a few lagging subproblems delay global convergence. A natural remedy is to adaptively re-split these stagnating segments online. Yet, deciding when, where, and how to split exceeds the capability of rule-based heuristics. To this end, we propose ATRS, a novel framework that embeds a shared Deep Reinforcement Learning policy into the parallel ADMM loop. We formulate this adaptive adjustment as a Multi-Agent Shared-Policy Markov Decision Process, where all trajectory segments act as homogeneous agents and share a unified neural policy network. This parameter-sharing architecture endows the system with size invariance, enabling it to handle dynamically changing segment counts during re-splitting and generalize to arbitrary trajectory lengths. Furthermore, our formulation inherently supports zero-shot generalization to unseen environments, as our network relies solely on the internal states of the numerical solver rather than on the geometric features of the environment. To ensure solver stability, a Confidence-Based Election mechanism selects only the most stagnating segment for re-splitting at each step. Extensive simulations demonstrate that ATRS accelerates convergence, reducing the number of iterations by up to 26.0% and the computation time by up to 19.1%. Real-world experiments further confirm its applicability to both large-scale offline global planning and real-time onboard replanning within 35 ms per cycle, with no sim-to-real degradation.

cs.RO

Macroscopic Optical Nonreciprocity: A Black Hole as an Optical Diode

Optical reciprocity--the principle that light retraces the same path when source and detector are interchanged--is a foundational concept in geometric optics. In this Letter, we demonstrate that this ``symmetry-protected'' behavior can be qualitatively overturned in a rotating black hole when spontaneous Lorentz symmetry breaking introduces a nonminimally coupled background structure with a preferred direction. Through numerical ray-tracing simulations, we reveal a striking macroscopic signature: upon optical-path reversal achieved by exchanging the source and the observer, the shadow of the same black hole morphs from a quasi-symmetric rugby-ball shape into a distinct teardrop profile. This high-contrast nonreciprocity effectively turns the black hole into a cosmic-scale optical diode, offering a novel pathway to probe fundamental symmetries using current and next-generation horizon-scale imaging.

gr-qc

Survival of nonclassical correlations in Lorentz-violating spacetime

The breakdown of Lorentz invariance, a potential signature of quantum gravity, offers a window into physics beyond general relativity. We investigate how such a violation, embodied by the Einstein-Bumblebee black hole spacetime, influences the nonlocal quantum correlations. Specifically, we study the quantum steering and Bell nonlocality between modes trapped inside and outside the event horizon of an Einstein-Bumblebee black hole. Our analysis demonstrates that quantum steering for an initially correlated state is confined to a narrow region near the event horizon, with the Lorentz-violating parameter further constraining this domain. Notably, the degree of steering asymmetry is significantly modulated by both the distance from the horizon and the Lorentz-violating parameter, with the two spatially separated regions exhibiting opposite trends. Furthermore, the Bell nonlocality measurable by an external observer strengthens with increasing distance from the black hole. These findings confirm the persistence of nonclassical correlations in a Lorentz-violating gravitational background and and offer a novel perspective on the interplay between quantum information and fundamental spacetime symmetries.

gr-qc

Boundary neuron method for solving partial differential equations

We propose a boundary neuron method with random features (BNM-RF) for solving partial differential equations. The method approximates the unknown boundary function by a shallow network within the boundary integral formulation. With randomly sampled and fixed hidden parameters, the computation reduces to a linear least squares problem for the output coefficients, which avoids gradient based nonconvex optimization. This construction retains the dimensionality reduction of boundary integral equations and the linear solution structure of the random feature method. For elliptic problems, we establish convergence analysis by combining kernel-based method with random feature approximation, and obtain error bounds on both the boundary and the interior solution. Numerical experiments on Laplace and Helmholtz problems, including interior and exterior cases, show that the proposed method achieves competitive accuracy relative to the boundary element method and favorable performance relative to boundary integral neural networks in the tested settings with only few neurons. Overall, the proposed method provides a practical framework for combining boundary integral equations with neural network for problems on complex geometries and unbounded domains.

math.NA

Automating Skill Acquisition through Large-Scale Mining of Open-Source Agentic Repositories: A Framework for Multi-Agent Procedural Knowledge Extraction

The transition from monolithic large language models (LLMs) to modular, skill-equipped agents represents a fundamental architectural shift in artificial intelligence deployment. While general-purpose models demonstrate remarkable breadth in declarative knowledge, their utility in autonomous workflows is frequently constrained by insufficient specialized procedural expertise. This report investigates a systematic framework for automated acquisition of high-quality agent skills through mining of open-source repositories on platforms such as GitHub. We focus on the extraction of visualization and educational capabilities from state-of-the-art systems including TheoremExplainAgent and Code2Video, both utilizing the Manim mathematical animation engine. The framework encompasses repository structural analysis, semantic skill identification through dense retrieval, and translation to the standardized SKILL.md format. We demonstrate that systematic extraction from agentic repositories, combined with rigorous security governance and multi-dimensional evaluation metrics, enables scalable acquisition of procedural knowledge that augments LLM capabilities without requiring model retraining. Our analysis reveals that agent-generated educational content can achieve 40\% gains in knowledge transfer efficiency while maintaining pedagogical quality comparable to human-crafted tutorials.

cs.AI