Searcharxiv⌕ Search

arXiv subjects

Jun Song

Publications and source records attributed to Jun Song.

At least 73 records · Page 4Linked to original sources

Video SimpleQA: Towards Factuality Evaluation in Large Video Language Models

Recent advancements in Large Video Language Models (LVLMs) have highlighted their potential for multi-modal understanding, yet evaluating their factual grounding in videos remains a critical unsolved challenge. To address this gap, we introduce Video SimpleQA, the first comprehensive benchmark tailored for factuality evaluation in video contexts. Our work differs from existing video benchmarks through the following key features: 1) Knowledge required: demanding integration of external knowledge beyond the video's explicit narrative; 2) Multi-hop fact-seeking question: Each question involves multiple explicit facts and requires strict factual grounding without hypothetical or subjective inferences. We also include per-hop single-fact-based sub-QAs alongside final QAs to enable fine-grained, stepby-step evaluation; 3) Short-form definitive answer: Answers are crafted as unambiguous and definitively correct in a short format with minimal scoring variance; 4) Temporal grounded required: Requiring answers to rely on one or more temporal segments in videos, rather than single frames. We extensively evaluate 33 state-of-the-art LVLMs and summarize key findings as follows: 1) Current LVLMs exhibit notable deficiencies in factual adherence, with the best-performing model o3 merely achieving an F-score of 66.3%; 2) Most LVLMs are overconfident in what they generate, with self-stated confidence exceeding actual accuracy; 3) Retrieval-augmented generation demonstrates consistent improvements at the cost of additional inference time overhead; 4) Multi-hop QA demonstrates substantially degraded performance compared to single-hop sub-QAs, with first-hop object or event recognition emerging as the primary bottleneck. We position Video SimpleQA as the cornerstone benchmark for video factuality assessment, aiming to steer LVLM development toward verifiable grounding in real-world contexts.

cs.CV↗

DeepPHY: Benchmarking Agentic VLMs on Physical Reasoning

Although Vision Language Models (VLMs) exhibit strong perceptual abilities and impressive visual reasoning, they struggle with attention to detail and precise action planning in complex, dynamic environments, leading to subpar performance. Real-world tasks typically require complex interactions, advanced spatial reasoning, long-term planning, and continuous strategy refinement, usually necessitating understanding the physics rules of the target scenario. However, evaluating these capabilities in real-world scenarios is often prohibitively expensive. To bridge this gap, we introduce DeepPHY, a novel benchmark framework designed to systematically evaluate VLMs' understanding and reasoning about fundamental physical principles through a series of challenging simulated environments. DeepPHY integrates multiple physical reasoning environments of varying difficulty levels and incorporates fine-grained evaluation metrics. Our evaluation finds that even state-of-the-art VLMs struggle to translate descriptive physical knowledge into precise, predictive control.

cs.AI↗

Thermophysical and Mechanical Properties Prediction of Rear-earth High-entropy Pyrochlore Based on Deep-learning Potential

High-entropy pyrochlore oxides possess ultra-low thermal conductivity and excellent high-temperature phase stability, making them promising candidate for next-generation thermal barrier coating (TBC) materials. However, reliable predictive models for such complex and disordered systems remain challenging. Ab initio methods, although accurate in describing anharmonic phonon-phonon interactions, struggle to capture the strong inherent phonon-disorder scattering in high-entropy systems. Moreover, the limited simulation cell size, hundreds of atoms, cannot fully represent the configurational complexity of high-entropy phases. On the other hand, classical molecular dynamics (MD) simulations lack accurate and transferable interatomic potentials, particularly in multi-component systems like high-entropy ceramics. In this work, we employed Deep Potential Molecular Dynamics (DPMD) to predict the thermophysical and mechanical properties of rare-earth high-entropy pyrochlore oxide system. The deep-potential (DP) model is trained on a limited dataset from ab initio molecular dynamics (AIMD) calculations, enabling large-scale molecular dynamics simulations with on-the-fly potential evaluations. This model not only achieves high accuracy in reproducing ab initio results but also demonstrates strong generalizability, making it applicable to medium-entropy ceramics containing the same constituent elements. Our study successfully develops a deep potential model for rare-earth pyrochlore systems and demonstrates that the deep-learning-based potential method offers a powerful computational approach for designing high-entropy TBC materials.

cond-mat.mtrl-sci↗

A Novel Discovery of Negative Thermal Expansion in Rare-earth Pyrochlore through Anion Order-Disorder Transition

In this study, we report for the first time the occurrence and investigation of the negative thermal expansion (NTE) effect in rare-earth pyrochlores. It is found that the NTE originates from the migration of oxygen anions from 48f sites to 8b sites, where one-twelfth of the original anions gradually occupy half of the available oxygen vacancies. This initial rapid transition leads to the distortion and rotation of polyhedral units, effectively contracting the lattice and manifesting as macroscopic NTE. The transition is sensitive to external isotropic pressure, where increasing pressure delays the onset of anion migration. This study deepens our understanding of NTE in complex oxides and demonstrates the utility of deep learning potentials for exploring intricate structural behaviors.

cond-mat.mtrl-sci↗

LongDocURL: a Comprehensive Multimodal Long Document Benchmark Integrating Understanding, Reasoning, and Locating

Large vision language models (LVLMs) have improved the document understanding capabilities remarkably, enabling the handling of complex document elements, longer contexts, and a wider range of tasks. However, existing document understanding benchmarks have been limited to handling only a small number of pages and fail to provide a comprehensive analysis of layout elements locating. In this paper, we first define three primary task categories: Long Document Understanding, numerical Reasoning, and cross-element Locating, and then propose a comprehensive benchmark, LongDocURL, integrating above three primary tasks and comprising 20 sub-tasks categorized based on different primary tasks and answer evidences. Furthermore, we develop a semi-automated construction pipeline and collect 2,325 high-quality question-answering pairs, covering more than 33,000 pages of documents, significantly outperforming existing benchmarks. Subsequently, we conduct comprehensive evaluation experiments on both open-source and closed-source models across 26 different configurations, revealing critical performance gaps in this field.

cs.AI↗

"See the World, Discover Knowledge": A Chinese Factuality Evaluation for Large Vision Language Models

The evaluation of factual accuracy in large vision language models (LVLMs) has lagged behind their rapid development, making it challenging to fully reflect these models' knowledge capacity and reliability. In this paper, we introduce the first factuality-based visual question-answering benchmark in Chinese, named ChineseSimpleVQA, aimed at assessing the visual factuality of LVLMs across 8 major topics and 56 subtopics. The key features of this benchmark include a focus on the Chinese language, diverse knowledge types, a multi-hop question construction, high-quality data, static consistency, and easy-to-evaluate through short answers. Moreover, we contribute a rigorous data construction pipeline and decouple the visual factuality into two parts: seeing the world (i.e., object recognition) and discovering knowledge. This decoupling allows us to analyze the capability boundaries and execution mechanisms of LVLMs. Subsequently, we evaluate 34 advanced open-source and closed-source models, revealing critical performance gaps within this field. Our evaluation-friendly code and data have already been open-sourced.

cs.CL↗

Masked Self-distilled Transducer-based Keyword Spotting with Semi-autoregressive Decoding

RNN-T-based keyword spotting (KWS) with autoregressive decoding~(AR) has gained attention due to its streaming architecture and superior performance. However, the simplicity of the prediction network in RNN-T poses an overfitting issue, especially under challenging scenarios, resulting in degraded performance. In this paper, we propose a masked self-distillation (MSD) training strategy that avoids RNN-Ts overly relying on prediction networks to alleviate overfitting. Such training enables masked non-autoregressive (NAR) decoding, which fully masks the RNN-T predictor output during KWS decoding. In addition, we propose a semi-autoregressive (SAR) decoding approach to integrate the advantages of AR and NAR decoding. Our experiments across multiple KWS datasets demonstrate that MSD training effectively alleviates overfitting. The SAR decoding method preserves the superior performance of AR decoding while benefits from the overfitting suppression of NAR decoding, achieving excellent results.

cs.SD↗

Atomistic insight into the effects of solute and pressure on phase transformation in titanium alloys

The phase stability and transformation between hexagonal close-packed (hcp) α-phase and body-centered cubic (bcc) \b{eta}-phase in titanium (Ti) alloys are critical to their mechanical properties and manufacturing processes for engineering applications. However, many factors, both intrinsic and extrinsic (e.g., solute elements and external pressures, respectively), may govern their phase transformations dynamically, which is crucial to the design of new Ti alloys with desirable properties. In this work, we study the effects of various solute elements and external hydrostatic pressures on the solid-state phase transformations in Ti alloys using density functional theory (DFT) and nudged elastic band (NEB) calculations. The results show that both alloying and applied pressure reduce transformation barriers, with Al and Mo being most effective under ambient conditions, while Nb, V, Zr, and Sn show enhanced transformation kinetics under stress. Solute-induced modifications to the local electronic structure and bonding environment, particularly under pressure, contribute to variations in phase stability. We identify a synergistic interaction between solute effects and external stress, which facilitates phase transitions that are unachievable under static conditions. These findings provide atomistic insights into the coupled chemical-mechanical mechanisms underlying phase transformations in Ti alloys with improved phase stability and mechanical performance.

cond-mat.mtrl-sci↗

FILA: Fine-Grained Vision Language Models

Recently, there has been growing interest in the capability of multimodal large language models (MLLMs) to process high-resolution images. A common approach currently involves dynamically cropping the original high-resolution image into smaller sub-images, which are then fed into a vision encoder that was pre-trained on lower-resolution images. However, this cropping approach often truncates objects and connected areas in the original image, causing semantic breaks. To address this limitation, we introduce HyViLM, designed to process images of any resolution while retaining the overall context during encoding. Specifically, we: (i) Design a new visual encoder called Hybrid Encoder that not only encodes individual sub-images but also interacts with detailed global visual features, significantly improving the model's ability to encode high-resolution images. (ii) Propose an optimal feature fusion strategy for the dynamic cropping approach, effectively leveraging information from different layers of the vision encoder. Compared with the state-of-the-art MLLMs under the same setting, our HyViLM outperforms existing MLLMs in nine out of ten tasks. Specifically, HyViLM achieves a 9.6% improvement in performance on the TextVQA task and a 6.9% enhancement on the DocVQA task.

cs.CV↗

FiLA-Video: Spatio-Temporal Compression for Fine-Grained Long Video Understanding

Recent advancements in video understanding within visual large language models (VLLMs) have led to notable progress. However, the complexity of video data and contextual processing limitations still hinder long-video comprehension. A common approach is video feature compression to reduce token input to large language models, yet many methods either fail to prioritize essential features, leading to redundant inter-frame information, or introduce computationally expensive modules.To address these issues, we propose FiLA(Fine-grained Vision Language Model)-Video, a novel framework that leverages a lightweight dynamic-weight multi-frame fusion strategy, which adaptively integrates multiple frames into a single representation while preserving key video information and reducing computational costs. To enhance frame selection for fusion, we introduce a keyframe selection strategy, effectively identifying informative frames from a larger pool for improved summarization. Additionally, we present a simple yet effective long-video training data generation strategy, boosting model performance without extensive manual annotation. Experimental results demonstrate that FiLA-Video achieves superior efficiency and accuracy in long-video comprehension compared to existing methods.

cs.CV↗

GeoSense: Evaluating Identification and Application of Geometric Principles in Multimodal Reasoning

Geometry problem-solving (GPS), a challenging task requiring both visual comprehension and symbolic reasoning, effectively measures the reasoning capabilities of multimodal large language models (MLLMs). Humans exhibit strong reasoning ability in this task through accurate identification and adaptive application of geometric principles within visual contexts. However, existing benchmarks fail to jointly assess both dimensions of the human-like geometric reasoning mechanism in MLLMs, remaining a critical gap in assessing their ability to tackle GPS. To this end, we introduce GeoSense, the first comprehensive bilingual benchmark designed to systematically evaluate the geometric reasoning abilities of MLLMs through the lens of geometric principles. GeoSense features a five-level hierarchical framework of geometric principles spanning plane and solid geometry, an intricately annotated dataset of 1,789 problems, and an innovative evaluation strategy. Through extensive experiments on GeoSense with various open-source and closed-source MLLMs, we observe that Gemini-2.0-pro-flash performs best, achieving an overall score of $65.3$. Our in-depth analysis reveals that the identification and application of geometric principles remain a bottleneck for leading MLLMs, jointly hindering their reasoning abilities. These findings underscore GeoSense's potential to guide future advancements in MLLMs' geometric reasoning capabilities, paving the way for more robust and human-like reasoning in artificial intelligence.

cs.CL↗

Productions of $^3_Λ$H, $^4_Λ$H and $^4_Λ$He in different coalescence channels in Au-Au collisions at $\sqrt{s_{NN}}=3$ GeV

We study the productions of $Λ$-hypernuclei $^3_Λ$H, $^4_Λ$H and $^4_Λ$He in the coalescence mechanism in Au-Au collisions at $\sqrt{s_{NN}}=3$ GeV. Considering the abundance and great importance of baryons and light (hyper-)nuclei on the collision dynamics, we include not only nucleon$+Λ$ coalescence but also nucleus+nucleon($Λ$) coalescence. We present contributions from different coalescence channels for $^3_Λ$H, $^4_Λ$H and $^4_Λ$He in their productions. We predict the production asymmetry between $^4_Λ$H and $^4_Λ$He, characterized by yield ratios $^4_Λ\text{He}/^4_Λ\text{H}$ and $(^4_Λ\text{H}-^4_Λ\text{He})/(^4_Λ\text{H}+^4_Λ\text{He})$, which can shed light on the existence constraints of the possible neutron-$Λ$ bound states $^2_Λn~(nΛ)$ and $^3_Λn~(nnΛ)$.

nucl-th↗

Anisotropic flows of identified hadrons in the equal-velocity quark combination model at RHIC energy

We employ an equal-velocity quark combination model to study anisotropic flows $v_{2}$, $v_{3}$ and $v_{4}$ of identified hadrons at mid-rapidity in heavy-ion collisions at RHIC energies. Under the equal-velocity combination mechanism of constituent quarks at hadronization, we build analytical formulas of anisotropic flows of hadrons in terms of those of quarks just before hadronization. We systematically analyze the contribution of higher order flows of quarks, and show how simple formulas of $v_{2}$, $v_{3}$ and $v_{4}$ of identified hadrons with the desired precision can be obtained by neglecting the small contribution of higher order flows of quarks. We systematically test these simple formulas of hadronic flows by the experimental data of $v_{2}$, $v_{3}$ and $v_{4}$ of identified hadrons $ϕ$, $Λ$, $Ξ^{-}$, $Ω^{-}$, $\barΛ$, $\barΞ^{+}$, $\barΩ^{+}$, $p$ and $\bar{p}$ in Au+Au collisions at $\sqrt{s_{NN}}=$ 19.6, 54.4 and 200 GeV, and we find that the equal-velocity quark combination model can well describe the measured $v_{2}$, $v_{3}$ and $v_{4}$ of identified hadrons in Au+Au collisions at those collision energies. We further study the obtained anisotropic flows of quarks and find two scaling properties\textcolor{red}{{} }which can be qualitatively understood by the hydrodynamic evolution of thermal quark medium produced in relativistic heavy-ion collisions.

hep-ph↗

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer

Vision transformers (ViTs) are widely employed in multimodal large language models (MLLMs) for visual encoding. However, they exhibit inferior performance on tasks regarding fine-grained visual perception. We attribute this to the limitations of ViTs in capturing diverse multi-modal visual levels, such as low-level details. To address this issue, we present LLaVA-UHD v2, an MLLM with advanced perception abilities by introducing a well-designed vision-language projector, the Hierarchical window (Hiwin) transformer. Hiwin transformer enhances MLLM's ability to capture diverse multi-modal visual granularities, by incorporating our constructed high-resolution semantic pyramid. Specifically, Hiwin transformer comprises two key modules: (i) a visual detail injection module, which progressively injects low-level visual details into high-level language-aligned semantics features, thereby forming an inverse semantic pyramid (ISP), and (ii) a hierarchical window attention module, which leverages cross-scale windows to condense multi-level semantics from the ISP. Extensive experiments show that LLaVA-UHD v2 outperforms compared MLLMs on a wide range of benchmarks. Notably, our design achieves an average boost of 3.7% across 14 benchmarks compared with the baseline method, 9.3% on DocVQA for instance. All the data and code will be publicly available to facilitate future research.

cs.CV↗

Laser: Efficient Language-Guided Segmentation in Neural Radiance Fields

In this work, we propose a method that leverages CLIP feature distillation, achieving efficient 3D segmentation through language guidance. Unlike previous methods that rely on multi-scale CLIP features and are limited by processing speed and storage requirements, our approach aims to streamline the workflow by directly and effectively distilling dense CLIP features, thereby achieving precise segmentation of 3D scenes using text. To achieve this, we introduce an adapter module and mitigate the noise issue in the dense CLIP feature distillation process through a self-cross-training strategy. Moreover, to enhance the accuracy of segmentation edges, this work presents a low-rank transient query attention mechanism. To ensure the consistency of segmentation for similar colors under different viewpoints, we convert the segmentation task into a classification task through label volume, which significantly improves the consistency of segmentation in color-similar areas. We also propose a simplified text augmentation strategy to alleviate the issue of ambiguity in the correspondence between CLIP features and text. Extensive experimental results show that our method surpasses current state-of-the-art technologies in both training speed and performance. Our code is available on: https://github.com/xingy038/Laser.git.

cs.CV↗

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback

Reinforcement learning from human feedback (RLHF) has proven effective in enhancing the instruction-following capabilities of large language models; however, it remains underexplored in the cross-modality domain. As the number of modalities increases, aligning all-modality models with human intentions -- such as instruction following -- becomes a pressing challenge. In this work, we make the first attempt to fine-tune all-modality models (i.e. input and output with any modality, also named any-to-any models) using human preference data across all modalities (including text, image, audio, and video), ensuring its behavior aligns with human intentions. This endeavor presents several challenges. First, there is no large-scale all-modality human preference data in existing open-source resources, as most datasets are limited to specific modalities, predominantly text and image. Secondly, the effectiveness of binary preferences in RLHF for post-training alignment in complex all-modality scenarios remains an unexplored area. Finally, there is a lack of a systematic framework to evaluate the capabilities of all-modality models, particularly regarding modality selection and synergy. To address these challenges, we propose the align-anything framework, which includes meticulously annotated 200k all-modality human preference data. Then, we introduce an alignment method that learns from unified language feedback, effectively capturing complex modality-specific human preferences and enhancing the model's instruction-following capabilities. Furthermore, to assess performance improvements in all-modality models after post-training alignment, we construct a challenging all-modality capability evaluation framework -- eval-anything. All data, models, and code frameworks have been open-sourced for the community. For more details, please refer to https://github.com/PKU-Alignment/align-anything.

cs.AI↗

Demystify Mamba in Vision: A Linear Attention Perspective

Mamba is an effective state space model with linear computation complexity. It has recently shown impressive efficiency in dealing with high-resolution inputs across various vision tasks. In this paper, we reveal that the powerful Mamba model shares surprising similarities with linear attention Transformer, which typically underperform conventional Transformer in practice. By exploring the similarities and disparities between the effective Mamba and subpar linear attention Transformer, we provide comprehensive analyses to demystify the key factors behind Mamba's success. Specifically, we reformulate the selective state space model and linear attention within a unified formulation, rephrasing Mamba as a variant of linear attention Transformer with six major distinctions: input gate, forget gate, shortcut, no attention normalization, single-head, and modified block design. For each design, we meticulously analyze its pros and cons, and empirically evaluate its impact on model performance in vision tasks. Interestingly, the results highlight the forget gate and block design as the core contributors to Mamba's success, while the other four designs are less crucial. Based on these findings, we propose a Mamba-Inspired Linear Attention (MILA) model by incorporating the merits of these two key designs into linear attention. The resulting model outperforms various vision Mamba models in both image classification and high-resolution dense prediction tasks, while enjoying parallelizable computation and fast inference speed. Code is available at https://github.com/LeapLabTHU/MLLA.

cs.CV↗

Fine-Grained Embedding Dimension Optimization During Training for Recommender Systems

Huge embedding tables in modern deep learning recommender models (DLRM) require prohibitively large memory during training and inference. This paper proposes FIITED, a system to automatically reduce the memory footprint via FIne-grained In-Training Embedding Dimension pruning. By leveraging the key insight that embedding vectors are not equally important, FIITED adaptively adjusts the dimension of each individual embedding vector during model training, assigning larger dimensions to more important embeddings while adapting to dynamic changes in data. We prioritize embedding dimensions with higher frequencies and gradients as more important. To enable efficient pruning of embeddings and their dimensions during model training, we propose an embedding storage system based on virtually-hashed physically-indexed hash tables. Experiments on two industry models and months of realistic datasets show that FIITED can reduce DLRM embedding size by more than 65% while preserving model quality, outperforming state-of-the-art in-training embedding pruning methods. On public datasets, FIITED can reduce the size of embedding tables by 2.1x to 800x with negligible accuracy drop, while improving model throughput.

cs.IR↗