Searcharxiv⌕ Search

arXiv subjects

Haojie Zhang

Publications and source records attributed to Haojie Zhang.

At least 19 recordsLinked to original sources

MuSS: A Large-Scale Dataset and Cinematic Narrative Benchmark for Multi-Shot Subject-to-Video Generation

While video foundation models excel at single-shot generation, real-world cinematic storytelling inherently relies on complex multi-shot sequencing. Further progress is constrained by the absence of datasets that address three core challenges: authentic narrative logic, spatiotemporal text-video alignment conflicts, and the "copy-paste" dilemma prevalent in Subject-to-Video (S2V) generation. To bridge this gap, we introduce MuSS, a large-scale, dual-track dataset tailored for multi-shot video and S2V generation. Sourced from over 3,000 movies, MuSS explicitly supports both complex montage transitions and subject-centric narratives. To construct this dataset, we pioneer a progressive captioning pipeline that eliminates contextual conflicts by ensuring local shot-level accuracy before enforcing global narrative coherence. Crucially, we implement a cross-shot matching mechanism to fundamentally eradicate the S2V copy-paste shortcut. Alongside the dataset, we propose the Cinematic Narrative Benchmark, featuring a visual-logic-driven paradigm and a novel Anti-Copy-Paste Variance (ACP-Var) metric to rigorously assess continuous storytelling and 3D structural consistency. Extensive experiments demonstrate that while current baselines struggle with continuous narrative logic or degenerate into trivial 2D sticker generators, our MuSS-augmented model achieves state-of-the-art narrative effectiveness and cross-shot identity preservation.

cs.CV↗

Infinite Worlds with Versatile Interactions

We present LingBot-World 2.0 (also known as LingBot-World-Infinity), an advanced iteration of LingBot-World featuring four distinct upgrades. (1) Our model achieves an unbounded interaction horizon while maintaining consistent output quality, benefiting from a carefully crafted causal pretraining paradigm. (2) Through distilling a real-time variant from the base model, our system guarantees rapid response time, sufficient to drive 720p video streams at 60 fps. (3) Compared to the previous version, this update introduces highly diverse interactive elements, comprising a broader spectrum of actions (e.g., attacking, archery, spell-casting, and shooting) alongside a richer variety of text-driven events. (4) We pioneer the integration of an agentic harness within the domain of world modeling, wherein a pilot agent is tasked with planning and executing character behaviors, while a director agent is responsible for synthesizing novel environmental elements as the scene progresses. Additionally, to facilitate a shared experience, we develop an interface that permits multiple players to simultaneously immerse themselves in this vivid world simulator. We pair our primary 14B model with a lightweight 1.3B counterpart, which supports effortless deployment on a single GPU.

cs.CV↗

Advancing WordArt-Oriented Scene Text Recognition: Datasets and Methods

WordArt (artistic text) features highly customized fonts, textures, and layouts, making WordArt-oriented scene TExt Recognition (WATER) substantially more challenging than general Scene Text Recognition (STR). Existing STR datasets and methods, typically built around regular scene text and fixed-template inputs, struggle to scale to WATER. Thus, we aim to advance this task from both data and model perspectives. On the data side, we construct a 2M synthetic dataset, WATER-S, with the scale improved by hundreds of times compared to existing artistic text data. WATER-S consists of two complementary subsets. One rendered by an upgraded rendering pipeline (SynthWordArt), which provides highly accurate and controllable synthetic WordArt data. The other is generated by combining Qwen3-VL for prompt mining and Z-Image for image synthesis, which improves the coverage of realistic and diverse data. On the model side, we propose WATERec. It adopts an visual encoder supporting arbitrary-shaped inputs and an autoregressive decoder to model complex layouts, structurally breaking the bottleneck of fixed-template STR on WordArt. Experiments show that this architecture outperforms prior STR methods, achieving state-of-the-art performance on irregular texts such as WordArt. Together with WATER-R, carefully reorganized from existing real STR data, our strong baseline with the new synthetic data and model design reaches 90.40% accuracy on WordArt-Bench, surpassing both general-purpose and OCR-specialized vision-language models by a large margin. Code and data are available at https://github.com/YesianRohn/WATER.

cs.CV↗

Multimodal Diffusion Transformer with Memory Bank for Scalable Long-Duration Talking Video Generation

Long-duration talking video synthesis faces enduring challenges in achieving high video quality, portrait consistency, temporal coherence, and computational efficiency. As video length increases, issues such as visual degradation, portrait drift, temporal artifacts, and error accumulation become increasingly problematic, severely affecting the realism and reliability of the results. To address these challenges, we present LetsTalk, a diffusion transformer framework equipped with multimodal guidance and a novel memory bank mechanism, explicitly maintaining contextual continuity and enabling robust, high-quality, and efficient generation of long-duration talking videos. In particular, LetsTalk introduces a noise-regularized memory bank to alleviate error accumulation and sampling artifacts during extended video generation. To further improve efficiency and spatiotemporal modeling, LetsTalk employs a deep compression autoencoder and a spatiotemporal-aware transformer with linear attention for effective multimodal fusion. We systematically analyze three fusion schemes and show that combining deep (Symbiotic Fusion) for portrait features and shallow (Direct Fusion) for audio achieves superior visual realism and precise speech-driven motion, while preserving diversity of movements. Extensive experiments demonstrate that LetsTalk establishes new state-of-the-art in generation quality, producing temporally coherent and realistic talking videos with enhanced diversity and liveliness, and maintains remarkable efficiency with 8x fewer parameters than previous approaches.

cs.CV↗

Video Understanding Reward Modeling: A Robust Benchmark and Performant Reward Models

Multimodal reward models have advanced substantially in text and image domains, yet progress in video understanding reward modeling remains severely limited by the lack of robust evaluation benchmarks and high-quality preference data. To address this, we propose a unified framework spanning benchmark design, data construction, and reward model training. We introduce Video Understanding Reward Bench (VURB), a benchmark featuring 2,100 preference pairs with long chain-of-thought reasoning traces (averaging 1,143 tokens) and majority voting evaluation across general, long, and reasoning-oriented video tasks. We further construct Video Understanding Preference Dataset (VUP-35K) via a fully automated pipeline, providing large-scale high-quality supervision for video reward training. Building on the data, we train VideoDRM and VideoGRM, a discriminative and a generative reward model, both achieving state-of-the-art performance on VURB and VideoRewardBench. Further analysis confirms that VUP-35K enhances both reward performance and model reasoning capability, while VideoDRM and VideoGRM yield significant gains under best-of-$N$ test-time scaling.

cs.CV↗

MSD-Score: Multi-Scale Distributional Scoring for Reference-Free Image Caption Evaluation

Evaluating image captions without references remains challenging because global embedding similarity often misses fine-grained mismatches such as hallucinated objects, missing attributes, or incorrect relations. We propose MSD-Score, a reference-free metric that models image patch and text token embeddings as von Mises-Fisher mixtures on the unit hypersphere. Instead of treating each modality as a single point, MSD-Score formulates image-text matching as a multi-scale distributional scoring problem. Semantic discrepancies are quantified via a weighted bi-directional KL divergence and combined with global similarity in a multi-scale framework for both single- and multi-candidate evaluations. Extensive experiments show that MSD-Score achieves state-of-the-art correlation with human judgments among reference-free metrics. Beyond accuracy, its probabilistic formulation yields transparent and decomposable diagnostics of local grounding errors, providing a deterministic complementary signal to holistic similarity metrics and judge-based evaluators.

cs.CV↗

DiffCap-Bench: A Comprehensive, Challenging, Robust Benchmark for Image Difference Captioning

Image Difference Captioning (IDC) generates natural language descriptions that precisely identify differences between two images, serving as a key benchmark for fine-grained change perception, cross-modal reasoning, and image editing data construction. However, existing benchmarks lack diversity and compositional complexity, and standard lexical-overlap metrics (e.g., BLEU, METEOR) fail to capture semantic consistency or penalize hallucinations, which together prevent a comprehensive and robust evaluation of multimodal large language models (MLLMs) on IDC. To address these gaps, we introduce DiffCap-Bench, a comprehensive IDC benchmark covering ten distinct difference categories to ensure diversity and compositional complexity. Furthermore, we propose an LLM-as-a-Judge evaluation protocol grounded in human-validated Difference Lists, enabling a robust assessment of models' ability to both capture and describe visual changes. Through extensive evaluation of state-of-the-art MLLMs, we reveal significant performance gaps between proprietary and open-source models, highlight the critical importance of reasoning capability, and identify clear limitations in model scaling. Our framework also demonstrates strong alignment with human expert judgments and strong correlation with downstream image editing data construction quality. These findings establish DiffCap-Bench as both a reliable IDC evaluation framework and a practical predictor of downstream utility. The benchmark and code will be made publicly available to support further research.

cs.CV↗

Towards AI Search Paradigm

In this paper, we introduce the AI Search Paradigm, a comprehensive blueprint for next-generation search systems capable of emulating human information processing and decision-making. The paradigm employs a modular architecture of four LLM-powered agents (Master, Planner, Executor and Writer) that dynamically adapt to the full spectrum of information needs, from simple factual queries to complex multi-stage reasoning tasks. These agents collaborate dynamically through coordinated workflows to evaluate query complexity, decompose problems into executable plans, and orchestrate tool usage, task execution, and content synthesis. We systematically present key methodologies for realizing this paradigm, including task planning and tool integration, execution strategies, aligned and robust retrieval-augmented generation, and efficient LLM inference, spanning both algorithmic techniques and infrastructure-level optimizations. By providing an in-depth guide to these foundational components, this work aims to inform the development of trustworthy, adaptive, and scalable AI search systems.

cs.CL↗

Experimental Study of Bremsstrahlung Gamma Ray Emission and Short-Range Correlations in $^{124}$Sn+$^{124}$Sn Collisions at 25 MeV/u

Short-range correlation (SRC) in nuclei refers to nucleons forming temporally correlated pairs in close proximity, giving rise to the high momentum of the nucleons beyond the Fermi surface. It has been reported that bremsstrahlung $γ$ production from neutron-proton process in heavy-ion reactions provides a potential probe to the SRC abundance in nuclei. In this paper, we present in detail the precision measurement of bremsstrahlung $γ$-rays in $\rm ^{124}Sn$+$\rm ^{124}Sn$ reactions at 25 MeV/u using the Compact Spectrometer for Heavy IoN Experiment (CSHINE). A comprehensive experimental and analysis framework is established to ensure the reliability and robustness of the extracted results. Background contributions are evaluated and subtracted using independent methods, and the consistency of the analysis is systematically validated. By comparing the experimental $γ$ spectrum with the Isospin-dependent Boltzmann-Uehling-Uhlenbeck simulations, the high momentum tail (HMT) fraction of $R_{\rm HMT}=(20 \pm 3)\%$ is derived in $^{124}$Sn nuclei. This work provides a detailed and validated experimental framework for extracting SRC information from bremsstrahlung $γ$-ray emission and demonstrates the feasibility of studying nucleon SRCs with high precision in low-energy heavy-ion collisions.

nucl-ex↗

UniVTAC: A Unified Simulation Platform for Visuo-Tactile Manipulation Data Generation, Learning, and Benchmarking

Robotic manipulation has seen rapid progress with vision-language-action (VLA) policies. However, visuo-tactile perception is critical for contact-rich manipulation, as tasks such as insertion are difficult to complete robustly using vision alone. At the same time, acquiring large-scale and reliable tactile data in the physical world remains costly and challenging, and the lack of a unified evaluation platform further limits policy learning and systematic analysis. To address these challenges, we propose UniVTAC, a simulation-based visuo-tactile data synthesis platform that supports three commonly used visuo-tactile sensors and enables scalable and controllable generation of informative contact interactions. Based on this platform, we introduce the UniVTAC Encoder, a visuo-tactile encoder trained on large-scale simulation-synthesized data with designed supervisory signals, providing tactile-centric visuo-tactile representations for downstream manipulation tasks. In addition, we present the UniVTAC Benchmark, which consists of eight representative visuo-tactile manipulation tasks for evaluating tactile-driven policies. Experimental results show that integrating the UniVTAC Encoder improves average success rates by 17.1% on the UniVTAC Benchmark, while real-world robotic experiments further demonstrate a 25% improvement in task success. Our webpage is available at https://univtac.github.io/.

cs.RO↗

Probing the Three-dimension Emission Source and Neutron Skin via $π$-$π$ Correlations in Heavy-Ion Collisions

The Richardson-Lucy algorithm is applied to reconstruct the three-dimensional source function of identical pions from their two-particle correlation functions. The algorithm's performance is first evaluated through simulations with Gaussian-type initial source functions. Its imaging quality and robustness are further demonstrated with experimental data from Au+Au collisions at 1.23 A GeV, collected by the HADES Collaboration. Additionally, using UrQMD simulations of Pb+Pb collisions at 1.5 A GeV, we show that the deblurred source functions exhibit sensitivity to the initial neutron skin thickness of the colliding nuclei. This highlights the potential of the Richardson-Lucy algorithm as a tool for probing the neutron density distribution in heavy nuclei.

nucl-ex↗

Unlocking the initial neutron density distribution from the two-pion HBT correlation function in heavy-ion collisions

Revealing the neutron density distribution in the nucleus is one of the crucial tasks of nuclear physics. Within the framework of the ultrarelativistic quantum molecular dynamic model followed by a correlation afterburner program, we investigate the effects of the initial neutron density distribution on the charged-pion yield ratio $π^{-}/π^{+}$, the two-pion momentum correlation function, and the emission source dimension. It is found that the $π^{-}/π^{+}$ ratio is sensitive to the initial neutron density distribution and the impact parameter, especially for collisions at large impact parameter. However, the charge splitting in the correlation functions between positively $π^{+}π^{+}$ and negatively $π^{-}π^{-}$, as well as the source radii and volumes extracted exhibit a stronger dependence on the initial neutron density distribution, but a weaker dependence on the impact parameter. The present study highlights that $π^{+}π^{+}$ and $π^{-}π^{-}$ correlation functions in heavy-ion collisions could be used to probe the initial neutron density distribution of nuclei.

nucl-th↗

Patch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs

Multimodal large language models (MLLMs) have advanced rapidly in recent years. However, existing approaches for vision tasks often rely on indirect representations, such as generating coordinates as text for detection, which limits performance and prevents dense prediction tasks like segmentation. To overcome these challenges, we introduce Patch-as-Decodable Token (PaDT), a unified paradigm that enables MLLMs to directly generate both textual and diverse visual outputs. Central to PaDT are Visual Reference Tokens (VRTs), derived from visual patch embeddings of query images and interleaved seamlessly with LLM's output textual tokens. A lightweight decoder then transforms LLM's outputs into detection, segmentation, and grounding predictions. Unlike prior methods, PaDT processes VRTs independently at each forward pass and dynamically expands the embedding table, thus improving localization and differentiation among similar objects. We further tailor a training strategy for PaDT by randomly selecting VRTs for supervised fine-tuning and introducing a robust per-token cross-entropy loss. Our empirical studies across four visual perception and understanding tasks suggest PaDT consistently achieving state-of-the-art performance, even compared with significantly larger MLLM models. The code is available at https://github.com/Gorilla-Lab-SCUT/PaDT.

cs.CV↗

DropLoRA: Sparse Low-Rank Adaptation for Parameter-Efficient Fine-Tuning

LoRA-based large model parameter-efficient fine-tuning (PEFT) methods use low-rank de- composition to approximate updates to model parameters. However, compared to full- parameter fine-tuning, low-rank updates often lead to a performance gap in downstream tasks. To address this, we introduce DropLoRA, a novel pruning-based approach that focuses on pruning the rank dimension. Unlike conven- tional methods that attempt to overcome the low-rank bottleneck, DropLoRA innovatively integrates a pruning module between the two low-rank matrices in LoRA to simulate dy- namic subspace learning. This dynamic low- rank subspace learning allows DropLoRA to overcome the limitations of traditional LoRA, which operates within a static subspace. By continuously adapting the learning subspace, DropLoRA significantly boosts performance without incurring additional training or infer- ence costs. Our experimental results demon- strate that DropLoRA consistently outperforms LoRA in fine-tuning the LLaMA series across a wide range of large language model gener- ation tasks, including commonsense reason- ing, mathematical reasoning, code generation, and instruction-following. Our code is avail- able at https://github.com/TayeeChang/DropLoRA.

cs.CL↗

Contrastive Regularization over LoRA for Multimodal Biomedical Image Incremental Learning

Multimodal Biomedical Image Incremental Learning (MBIIL) is essential for handling diverse tasks and modalities in the biomedical domain, as training separate models for each modality or task significantly increases inference costs. Existing incremental learning methods focus on task expansion within a single modality, whereas MBIIL seeks to train a unified model incrementally across modalities. The MBIIL faces two challenges: I) How to preserve previously learned knowledge during incremental updates? II) How to effectively leverage knowledge acquired from existing modalities to support new modalities? To address these challenges, we propose MSLoRA-CR, a method that fine-tunes Modality-Specific LoRA modules while incorporating Contrastive Regularization to enhance intra-modality knowledge sharing and promote inter-modality knowledge differentiation. Our approach builds upon a large vision-language model (LVLM), keeping the pretrained model frozen while incrementally adapting new LoRA modules for each modality or task. Experiments on the incremental learning of biomedical images demonstrate that MSLoRA-CR outperforms both the state-of-the-art (SOTA) approach of training separate models for each modality and the general incremental learning method (incrementally fine-tuning LoRA). Specifically, MSLoRA-CR achieves a 1.88% improvement in overall performance compared to unconstrained incremental learning methods while maintaining computational efficiency. Our code is publicly available at https://github.com/VentusAislant/MSLoRA_CR.

cs.LG↗

Precise Measurement of Short-Range Correlations in Nuclei from Bremsstrahlung Gamma Ray Emission in Low-Energy Heavy-Ion Collisions

Atomic nuclei and dense nucleonic matter in neutron stars exhibit short-range correlations (SRCs), where nucleons form temporally correlated pairs in proximity beyond mean-field approximation. It is essential to make precision measurement of the fraction of SRC since it carries the signature of underlying quark dynamics in nuclear medium. In this letter, we present the first high-precision measurement of neutron-proton bremsstrahlung $γ$-ray emission from the symmetric $\rm ^{124}Sn$+$\rm ^{124}Sn$ reactions at 25 MeV/u. From the observed spectral hardening, the precise SRC fraction in the $\rm ^{124}Sn$ nucleus is extracted to be $(20 \pm 3)\%$. This result provides a novel, direct and unambiguous evidence of SRCs, and demonstrates that low-energy heavy-ion collisions offers a new approach to studying nuclear structure in connection with quark-level dynamics.

nucl-ex↗

MADUV: The 1st INTERSPEECH Mice Autism Detection via Ultrasound Vocalization Challenge

The Mice Autism Detection via Ultrasound Vocalization (MADUV) Challenge introduces the first INTERSPEECH challenge focused on detecting autism spectrum disorder (ASD) in mice through their vocalizations. Participants are tasked with developing models to automatically classify mice as either wild-type or ASD models based on recordings with a high sampling rate. Our baseline system employs a simple CNN-based classification using three different spectrogram features. Results demonstrate the feasibility of automated ASD detection, with the considered audible-range features achieving the best performance (UAR of 0.600 for segment-level and 0.625 for subject-level classification). This challenge bridges speech technology and biomedical research, offering opportunities to advance our understanding of ASD models through machine learning approaches. The findings suggest promising directions for vocalization analysis and highlight the potential value of audible and ultrasound vocalizations in ASD detection.

cs.SD↗

Extract neutron-neutron interaction strength and spatial-temporal dynamics of neutron emission from two-particle correlation function

The neutron-neutron ($nn$) correlation function has been measured in 25 MeV/u $^{124}$Sn+$^{124}$Sn reactions. Using the Lednický-Lyuboshitz approach, the $nn$ scattering length and effective range ($f_{0}^{nn}$, $d_{0}^{nn}$), as well as the reduced space-time size $R^{(0)}$ of the neutron emission source are simultaneously extracted as ($18.9^{+1.3}_{-1.2}$ fm, $1.9^{+1.3}_{-1.0}$ fm) and $4.12 \pm 0.12$ fm, respectively. The measured $nn$ scattering length is consistent with the results obtained in the low-energy scattering $^{2}{\rm H}(π^{-},γ)2n$, indicating heavy-ion collisions can serve as an effective approach for measuring $nn$ interactions and further investigating the charge symmetry breaking of nuclear force. The space-time size extracted from momentum-gated correlation functions exhibits clear dependence on the pair momentum, with $R^{(0)}=2.8 \pm 0.1 $ fm and $4.9 \pm 0.2$ fm being determined for the high and low momentum neutrons, respectively.

nucl-ex↗