SearcharxivSearch

arXiv subjects

Wang Xiao

Publications and source records attributed to Wang Xiao.

5 recordsLinked to original sources

From Easy to Hard: The MIR Benchmark for Progressive Interleaved Multi-Image Reasoning

Multi-image Interleaved Reasoning aims to improve Multi-modal Large Language Models (MLLMs) ability to jointly comprehend and reason across multiple images and their associated textual contexts, introducing unique challenges beyond single-image or non-interleaved multi-image tasks. While current multi-image benchmarks overlook interleaved textual contexts and neglect distinct relationships between individual images and their associated texts, enabling models to reason over multi-image interleaved data may significantly enhance their comprehension of complex scenes and better capture cross-modal correlations. To bridge this gap, we introduce a novel benchmark MIR, requiring joint reasoning over multiple images accompanied by interleaved textual contexts to accurately associate image regions with corresponding texts and logically connect information across images. To enhance MLLMs ability to comprehend multi-image interleaved data, we introduce reasoning steps for each instance within the benchmark and propose a stage-wise curriculum learning strategy. This strategy follows an "easy to hard" approach, progressively guiding models from simple to complex scenarios, thereby enhancing their ability to handle challenging tasks. Extensive experiments benchmarking multiple MLLMs demonstrate that our method significantly enhances models reasoning performance on MIR and other established benchmarks. We believe that MIR will encourage further research into multi-image interleaved reasoning, facilitating advancements in MLLMs capability to handle complex inter-modal tasks.

cs.CV

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Humans acquire knowledge through three cognitive stages: perceiving information, comprehending knowledge, and adapting knowledge to solve novel problems. Videos serve as an effective medium for this learning process, facilitating a progression through these cognitive stages. However, existing video benchmarks fail to systematically evaluate the knowledge acquisition capabilities in Large Multimodal Models (LMMs). To address this gap, we introduce Video-MMMU, a multi-modal, multi-disciplinary benchmark designed to assess LMMs' ability to acquire and utilize knowledge from videos. Video-MMMU features a curated collection of 300 expert-level videos and 900 human-annotated questions across six disciplines, evaluating knowledge acquisition through stage-aligned question-answer pairs: Perception, Comprehension, and Adaptation. A proposed knowledge gain metric, {\Delta}knowledge, quantifies improvement in performance after video viewing. Evaluation of LMMs reveals a steep decline in performance as cognitive demands increase and highlights a significant gap between human and model knowledge acquisition, underscoring the need for methods to enhance LMMs' capability to learn and adapt from videos.

cs.CV

Early warning indicators via latent stochastic dynamical systems

Detecting early warning indicators for abrupt dynamical transitions in complex systems or high-dimensional observation data is essential in many real-world applications, such as brain diseases, natural disasters, and engineering reliability. To this end, we develop a novel approach: the directed anisotropic diffusion map that captures the latent evolutionary dynamics in the low-dimensional manifold. Then three effective warning signals (Onsager-Machlup Indicator, Sample Entropy Indicator, and Transition Probability Indicator) are derived through the latent coordinates and the latent stochastic dynamical systems. To validate our framework, we apply this methodology to authentic electroencephalogram (EEG) data. We find that our early warning indicators are capable of detecting the tipping point during state transition. This framework not only bridges the latent dynamics with real-world data but also shows the potential ability for automatic labeling on complex high-dimensional time series.

stat.ML

An eigenvalue problem for self-similar patterns in Hele-Shaw flows

Hele-Shaw problems are prototypes to study the interface dynamics. Linear theory suggests the existence of self-similar patterns in a Hele-Shaw flow. That is, with a specific injection flux the interface shape remains unchanged while its size increases. In this paper, we explore the existence of self-similar patterns in the nonlinear regime and develop a rigorous nonlinear theory characterizing their fundamental features. Using a boundary integral formulation, we pose the question of self-similarity as a generalized nonlinear eigenvalue problem, involving two nonlinear integral operators. The flux constant $C$ is the eigenvalue and the corresponding self-similar pattern $\mathbf{x}$ is the eigenvector. We develop a quasi-Newton method to solve the problem and show the existence of nonlinear shapes with $k$-fold dominated symmetries. The influence of initial guesses on the self-similar patterns is investigated. We are able to obtain a desired self-similar shape once the initial guess is properly chosen. Our results go beyond the predictions of linear theory and establish a bridge between the linear theory and simulations.

math.AP

Fourier neural operator based fluid-structure interaction for predicting the vesicle dynamics

Solving complex fluid-structure interaction (FSI) problems, characterized by nonlinear partial differential equations, is crucial in various scientific and engineering applications. Traditional computational fluid dynamics (CFD) solvers are insufficient to meet the growing requirements for large-scale and long-period simulations. Fortunately, the rapid advancement in neural networks, especially neural operator learning mappings between function spaces, has introduced novel approaches to tackle these challenges via data-driven modeling. In this paper, we propose a Fourier neural operator-based fluid-structure interaction solver (FNO-based FSI solver) for efficient simulation of FSI problems, where the solid solver based on the finite difference method is seamlessly integrated with the Fourier neural operator to predict incompressible flow using the immersed boundary method. We analyze the performance of the FNO-based FSI solver in the following three situations: training data with or without the steady state, training method with one-step label or multi-step labels, and prediction in interpolation or extrapolation. We find that the best performance for interpolation is achieved by training the operator with multi-step labels using steady-state data. Finally, we train the FNO-based FSI solver using this optimal training method and apply it to vesicle dynamics. The results show that the FNO-based FSI solver is capable of capturing the variations in the fluid and the vesicle.

math.DS