SearcharxivSearch

arXiv subjects

Jiangtao Wu

Publications and source records attributed to Jiangtao Wu.

12 recordsLinked to original sources

HappyWorld-Bench

Evaluating world models requires assessing both the quality of the worlds they generate and their consistency and responsiveness under exploration, interaction, and modification. We introduce HappyWorld-Bench, a comprehensive benchmark that evaluates whether generated worlds remain reliable as agents interact with them. Our design is built on a hierarchical capability framework of six world capabilities (W1-W6), from generative construction to unified world modeling, instantiated across three independent evaluation tracks: video world models, spatial world models, and embodied world models. HappyWorld-Bench comprises 1,138 video prompts, 300 spatial scenes, and 254 embodied test cases. Across all three tracks, we build and operate HappyWorld-Arena to organize human A/B comparisons and derive model-level Elo ratings, which complement newly designed automated metrics that capture behavioral correctness. We evaluate 14 video world models, 9 spatial systems, and 8 embodied candidates under this unified framework. Results reveal remaining reliability gaps across all three tracks: video models exhibit reduced consistency during extended rollouts and revisits, spatial models achieve at best 70.14% placement accuracy and 73.33% edit execution, and embodied models struggle to preserve state across multi-step actions and respond precisely to altered action conditions and physical rules. These findings highlight the need to evaluate world models not only by visual quality, but also by state consistency and the correctness of their responses to actions and interventions.

cs.CV

AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities

While instruction-based video editing has advanced rapidly, real-world videos contain tightly coupled audio and visual signals, and editing one modality often requires coordinated changes in the other. Existing benchmarks primarily evaluate visual transformations on silent clips or isolated audio editing, leaving complex audio-visual editing and cross-modal consistency underexplored. We introduce AVE-Compass, a comprehensive benchmark with 145 curated source videos, 196 audio-visually coupled editing instructions, and 2,688 fine-grained checklist items. It evaluates Instruction Following, Fidelity Preserving, Realism, and Editing Intent through checklist-based MLLM judging and a dedicated realism rubric, complemented by automated cross-modal, video, and audio metrics. Extensive evaluation shows that state-of-the-art models still struggle to execute cross-modal instructions while preserving non-target content. We further propose AVE-Agent, a modular agent framework that decomposes complex instructions into dependent subtasks and iteratively improves editing results through self-reflection and evaluator feedback. AVE-Agent improves instruction execution, Fidelity Preserving, and audio-visual alignment in joint editing while maintaining competitive perceptual quality.

cs.MM

CoVEBench: Can Video Editing Models Handle Complex Instructions?

While recent text-guided video editing models excel at elementary tasks (e.g., style transfer, object insertion), real-world user requests are highly compositional. A single prompt often demands multiple coupled edits, such as modifying subjects, actions, and camera views, while strictly preserving unrelated spatiotemporal content. Existing benchmarks, heavily constrained by isolated edits and coarse global metrics, fail to diagnose how models handle such complex workflows. To address this gap, we introduce CoVEBench, a compositional video editing benchmark comprising 416 curated source videos, 626 multi-point editing instructions, and 9,990 fine-grained checklist items. Covering diverse editing dimensions, CoVEBench evaluates models via MLLM-judged instruction compliance and video fidelity, alongside automated metrics for video quality. Extensive experiments reveal that compositional editing remains a profound challenge: current models frequently omit edits, violate preservation constraints, or introduce artifacts when handling multiple operations simultaneously. CoVEBench provides a challenging, diagnostic testbed to advance video editing toward realistic user workflows.

cs.CV

Where Do Deep-Research Agents Go Wrong? Span-Level Error Localization in Agent Trajectories

Deep-research agents solve tasks through long trajectories of search, tool use, evidence inspection, and answer synthesis. Evaluation based on final answers shows whether an agent succeeds, but not which parts of the trajectory make the answer unreliable. We study span-level error localization for deep-research agents. We collect 2,790 real trajectories from two agent frameworks, three backbone models, and three benchmarks, convert raw logs into semantic spans, and annotate harmful error spans through LLM-assisted expert review. From these annotations, we build TELBench, a 1,000-instance benchmark for identifying error spans among normal exploration, failed searches, tentative hypotheses, and harmless noise. We further propose DRIFT, a claim-centric auditing framework that tracks agent claims, checks their support in trajectory evidence, and marks spans where unsupported or conflicting claims affect the answer path. Experiments across model families and auditing frameworks show that DRIFT improves span-level error localization and first-error accuracy by up to 30 percentage points. Our work provides a process-level view of reliability in deep-research agents.

cs.AI

ViDiC: Video Difference Captioning

Understanding visual differences between dynamic scenes requires the comparative perception of compositional, spatial, and temporal changes--a capability that remains underexplored in existing vision-language systems. While prior work on Image Difference Captioning (IDC) has enabled models to describe semantic changes between static images, these approaches fail to capture motion continuity, event evolution, or editing consistency over time. We introduce the ViDiC (Video Difference Captioning) task and its corresponding ViDiC-1K dataset, designed to evaluate the ability of Multimodal Large Language Models (MLLMs) to provide fine-grained descriptions of similarities and differences between video pairs. ViDiC-1K comprises 1,000 curated video pairs annotated with over 4,000 comparative checklist items, covering seven categories: subject, style, background, cinematography, motion, location, and playback techniques. To ensure reliable evaluation, we propose a dual-checklist framework that measures the accuracy of similarity and difference separately, based on the LLM-as-a-Judge protocol. Experiments on nineteen representative multimodal models reveal a significant performance gap in their comparative description and difference perception abilities. We hope ViDiC-1K can be a challenging benchmark that lays a solid foundation for advancing video understanding, edit awareness, and comparative reasoning in multimodal intelligence.

cs.CV

IF-VidCap: Can Video Caption Models Follow Instructions?

Although Multimodal Large Language Models (MLLMs) have demonstrated proficiency in video captioning, practical applications require captions that follow specific user instructions rather than generating exhaustive, unconstrained descriptions. Current benchmarks, however, primarily assess descriptive comprehensiveness while largely overlooking instruction-following capabilities. To address this gap, we introduce IF-VidCap, a new benchmark for evaluating controllable video captioning, which contains 1,400 high-quality samples. Distinct from existing video captioning or general instruction-following benchmarks, IF-VidCap incorporates a systematic framework that assesses captions on two dimensions: format correctness and content correctness. Our comprehensive evaluation of over 20 prominent models reveals a nuanced landscape: despite the continued dominance of proprietary models, the performance gap is closing, with top-tier open-source solutions now achieving near-parity. Furthermore, we find that models specialized for dense captioning underperform general-purpose MLLMs on complex instructions, indicating that future work should simultaneously advance both descriptive richness and instruction-following fidelity.

cs.CV

Multi-Agent Amodal Completion: Direct Synthesis with Fine-Grained Semantic Guidance

Amodal completion, generating invisible parts of occluded objects, is vital for applications like image editing and AR. Prior methods face challenges with data needs, generalization, or error accumulation in progressive pipelines. We propose a Collaborative Multi-Agent Reasoning Framework based on upfront collaborative reasoning to overcome these issues. Our framework uses multiple agents to collaboratively analyze occlusion relationships and determine necessary boundary expansion, yielding a precise mask for inpainting. Concurrently, an agent generates fine-grained textual descriptions, enabling Fine-Grained Semantic Guidance. This ensures accurate object synthesis and prevents the regeneration of occluders or other unwanted elements, especially within large inpainting areas. Furthermore, our method directly produces layered RGBA outputs guided by visible masks and attention maps from a Diffusion Transformer, eliminating extra segmentation. Extensive evaluations demonstrate our framework achieves state-of-the-art visual quality.

cs.CV

Static and dynamical properties of the spin-5/2 nearly ideal triangular lattice antiferromagnet Ba3MnSb2O9

We study the ground state and spin excitations in Ba3MnSb2O9, an easy-plane S = 5/2 triangular lattice antiferromagnet. By combining single-crystal neutron scattering, electric spin resonance (ESR), and spin wave calculations, we determine the frustrated quasi-two-dimensional spin Hamiltonian parameters describing the material. While the material has a slight monoclinic structural distortion, which could allow for isosceles-triangular exchanges and biaxial anisotropy by symmetry, we observe no deviation from the behavior expected for spin waves in the in-plane 120o state. Even the easy-plane anisotropy is so small that it can only be detected by ESR in our study. In conjunction with the quasi-two-dimensionality, our study establishes that Ba3MnSb2O9 is a nearly ideal triangular lattice antiferromagnet with the quasi-classical spin S = 5/2, which suggests that it has the potential for an experimental study of Z- or Z2-vortex excitations.

cond-mat.str-el

Magnetic field effects on the quantum spin liquid behaviors of NaYbS$_2$

Spin-orbit coupling is an important ingredient to regulate the many-body physics, especially for many spin liquid candidate materials such as rare-earth magnets and Kitaev materials. The rare-earth chalcogenides NaYbCh$_2$ (Ch = O, S, Se) is a congenital frustrating system to exhibit the intrinsic landmark of spin liquid by eliminating both the site disorders between Na$^{+}$ and Yb$^{3+}$ ions with the big ionic size difference and the Dzyaloshinskii-Moriya interaction with the perfect triangular lattice of the Yb$^{3+}$ ions. The temperature versus magnetic-field phase diagram is established by the magnetization, specific heat, and neutron-scattering measurements. Notably, the neutron diffraction spectra and the magnetization curve might provide microscopic evidence for a series of spin configuration for in-plane fields, which include the disordered spin liquid state, 120$^{o}$ antiferromagnet, and one-half magnetization state. Furthermore, the ground state is suggested to be a gapless spin liquid from inelastic neutron scattering, and the magnetic field adjusts the spin orbit coupling. Therefore, the strong spin-orbit coupling in the frustrated quantum magnet substantially enriches low-energy spin physics. This rare-earth family could offer a good platform for exploring the quantum spin liquid ground state and quantum magnetic transitions.

cond-mat.str-el

Regulate the direct-indirect electronic band gap transition by electron-phonon interaction in BaSnO3

The neutron powder diffraction, specific heat, thermal conductivity, and Raman scattering measurements were presented to study the interplays of lattice, phonons and electrons of the Sr-doping Ba1-xSrxSnO3 (x was less than or equal to 0.1). Although Ba1-xSrxSnO3 kept the cubic lattice, the Raman spectra suggested a dynamic distortion at low temperature. The density functional theory was applied to analyze the electronic structures and phonon dispersions of Ba1-xSrxSnO3(x = 0, 0.0125), and the behaviors of electron bands around Fermi levels were discussed. According to the experimental and theoretical results, the Sr-doping played a significant role in tuning the indirect band gap of BaSnO3 and influenced the electron-phonon interaction.

cond-mat.mtrl-sci

Tunable Band Structures of Polycrystalline Graphene by External and Mismatch Strains

Lacking a band gap largely limits the application of graphene in electronic devices. Previous study shows that grain boundaries (GBs) in polycrystalline graphene can dramatically alter the electrical properties of graphene. Here, we investigate the band structure of polycrystalline graphene tuned by externally imposed strains and intrinsic mismatch strains at the GB by density functional theory (DFT) calculations. We found that graphene with symmetrical GBs typically has zero band gap even with large uniaxial and biaxial strain. However, some particular asymmetrical GBs can open a band gap in graphene and their band structures can be substantially tuned by external strains. A maximum band gap about 0.19 eV was observed in matched-armchair GB (5, 5) | (3, 7) with a misorientation of θ=13o when the applied uniaxial strain increases to 9%. Although mismatch strain is inevitable in asymmetrical GBs, it has a small influence on the band gap of polycrystalline graphene.

cond-mat.mtrl-sci

Mechanics and Tunable Bandgap by Straining in Single-Layer Hexagonal Boron-Nitride

Current interest in two-dimensional materials extends from graphene to others systems like single-layer hexagonal boron-nitride (h-BN), for the possibility of making heterogeneous structures to achieve exceptional properties that cannot be realized in graphene.The electrically insulating h-BN and semi-metal graphene may open good opportunities to realize a semiconductor by manipulating the morphology and composition of such heterogeneous structures.Here we report the mechanical properties of h-BN and its band structures tuned by mechanical straining by using the density functional theory calculations.The elastic properties, both the Young's modulus and bending rigidity for h-BN, are isotropic.We reveal that there is a bi-linear dependence of band gap on the applied tensile strains in h-BN. Mechanical strain can tune single-layer h-BN from an insulator to a semiconductor, with a band gap in the 4.7eV to 1.5eV range.

cond-mat.mtrl-sci