SearcharxivSearch

arXiv subjects

Xusheng Liu

Publications and source records attributed to Xusheng Liu.

7 recordsLinked to original sources

VideoSEAL: Mitigating Evidence Misalignment in Agentic Long Video Understanding by Decoupling Answer Authority

Long video question answering requires locating sparse, time-scattered visual evidence within highly redundant content. Although current MLLMs perform well on short videos, long videos introduce long-horizon search and verification, which often necessitates multi-turn, agentic interaction. We show that existing LVU agents can exhibit "evidence misalignment": they produce correct answers that are not supported by the retrieved or inspected evidence. To characterize this failure, we introduce two diagnostics (temporal groundedness and semantic groundedness) and use them to reveal two pressures that amplify misalignment: prompt pressure from shared-context saturation at inference time and reward pressure from outcome-only optimization during training. These findings point to a structural root cause: the coupled agent paradigm conflates long-horizon planning with answer authority. We therefore propose the decoupled planner-inspector framework, which separates planning from answer authority and gates final answering on pixel-level verification. Across four long-video benchmarks, our framework improves both answer accuracy and evidence alignment, achieving 55.1% on LVBench and 62.0% on LongVideoBench while producing interpretable search trajectories. Moreover, the decoupled architecture scales consistently with increased search budgets and supports plug-and-play upgrades of the MLLM backbone without retraining the planner. Code and models are available at https://github.com/Echochef/VideoSEAL.

cs.CV

UniSync: Towards Generalizable and High-Fidelity Lip Synchronization for Challenging Scenarios

Lip synchronization aims to generate realistic talking videos that match given audio, which is essential for high-quality video dubbing. However, current methods have fundamental drawbacks: mask-based approaches suffer from local color discrepancies, while mask-free methods struggle with global background texture misalignment. Furthermore, most methods struggle with diverse real-world scenarios such as stylized avatars, face occlusion, and extreme lighting conditions. In this paper, we propose UniSync, a unified framework designed for achieving high-fidelity lip synchronization in diverse scenarios. Specifically, UniSync uses a mask-free pose-anchored training strategy to keep head motion and eliminate synthesis color artifacts, while employing mask-based blending consistent inference to ensure structural precision and smooth blending. Notably, fine-tuning on compact but diverse videos empowers our model with exceptional domain adaptability, handling complex corner cases effectively. We also introduce the RealWorld-LipSync benchmark to evaluate models under real-world demands, which covers diverse application scenarios including both human faces and stylized avatars. Extensive experiments demonstrate that UniSync significantly outperforms state-of-the-art methods, advancing the field towards truly generalizable and production-ready lip synchronization.

cs.CV

MARS2 2025 Challenge on Multimodal Reasoning: Datasets, Methods, Results, Discussion, and Outlook

This paper reviews the MARS2 2025 Challenge on Multimodal Reasoning. We aim to bring together different approaches in multimodal machine learning and LLMs via a large benchmark. We hope it better allows researchers to follow the state-of-the-art in this very dynamic area. Meanwhile, a growing number of testbeds have boosted the evolution of general-purpose large language models. Thus, this year's MARS2 focuses on real-world and specialized scenarios to broaden the multimodal reasoning applications of MLLMs. Our organizing team released two tailored datasets Lens and AdsQA as test sets, which support general reasoning in 12 daily scenarios and domain-specific reasoning in advertisement videos, respectively. We evaluated 40+ baselines that include both generalist MLLMs and task-specific models, and opened up three competition tracks, i.e., Visual Grounding in Real-world Scenarios (VG-RS), Visual Question Answering with Spatial Awareness (VQA-SA), and Visual Reasoning in Creative Advertisement Videos (VR-Ads). Finally, 76 teams from the renowned academic and industrial institutions have registered and 40+ valid submissions (out of 1200+) have been included in our ranking lists. Our datasets, code sets (40+ baselines and 15+ participants' methods), and rankings are publicly available on the MARS2 workshop website and our GitHub organization page https://github.com/mars2workshop/, where our updates and announcements of upcoming events will be continuously provided.

cs.CV

L2-estimates on p-convex Riemannian manifolds

In this paper, we establish various L2-estimates for the exterior differential operator on p-convex Riemannian manifolds in the sense of Harvey and Lawson. As geometric applications, we prove vanishing and finiteness results for the de Rham cohomology groups.

math.DG

The partial positivity of the curvature in Riemannian symmetric spaces

In this paper, we determine the partial positivity(resp., negativity) of the curvature of all irreducible Riemannian symmetric spaces. From the classifications of abstract root systems and maximal subsystems, we can give the calculations for symmetric spaces both in classical types and in exceptional types.

math.DG

Vanishing theorem for irreducible symmetric spaces of noncompact type

We prove the following vanishing theorem. Let M be an irreducible symmetric space of noncompact type whose dimension exceeds 2 and $M\ne SO_0(2,2)/SO(2)\tm SO(2).$ Let E be any vector bundle over M, Then any E-valued $L^2$ harmonic 1-form over M vanishes. In particular we get the vanishing theorem for harmonic maps from irreducible symmetric spaces of noncompact type.

math.DG

Curvature estimates for irreducible symmetric spaces

By making use of the classification of real simple Lie algebra, we get the maximum of the squared length of restricted roots case by case, thus we get the upper bounds of sectional curvature for irreducible Riemannian symmetric spaces of compact type. As an application, we verify Sampson's conjecture in all cases for irreducible Riemannian symmetric spaces of noncompact type.

math.DG