SearcharxivSearch

arXiv subjects

Yuanze Li

Publications and source records attributed to Yuanze Li.

16 recordsLinked to original sources

Generalize LMMs to Versatile Visual Modalities via Fabricated Modality Synthesis

Despite the advancements of Large Multimodal Models (LMMs) in RGB vision, their ability to generalize to unseen visual modalities remains a largely unexplored challenge. We argue that different visual modalities are merely distinct samplings of the same physical world. Therefore, effective generalization requires models to possess both modality-agnostic perception of scene semantics and the adaptability to modality-specific characteristics. To achieve this, we propose a training framework, VVM-Tuning, to equip LMMs with these capabilities through modality synthesis and modality contexts. Specifically, we synthesize diverse appearance-varied images from RGB scenes, training the model to disentangle invariant semantics from varying visual appearances, and align these appearances with language for visual concepts decoupled from modalities. We then introduce modality contexts in the prompt and use instruction tuning to assist the model in mapping these appearance variations back to modality-related attributes, enabling zero-shot adaptation to unseen modalities during inference. To facilitate research in this direction, we introduce VVM-Bench, a comprehensive benchmark featuring 6 real and synthetic modalities to evaluate semantic perception and modality understanding. Experiments demonstrate that, via our training on synthetic modalities, 5 tested models exhibit consistent improvements on both real-world and novel synthetic modalities without in-modality training. Source code and data will be publicly available at https://github.com/Hunter-Will/VVM-Tuning.

cs.CV

Topological Surface Charge Detection via Terahertz Time-domain Spectroscopy

The topological magnetoelectric effect (TME) in three-dimensional topological insulators manifests as a quantized surface charge accumulation proportional to an applied magnetic field. Here we demonstrate an optical method using terahertz time-domain spectroscopy (THz-TDS) to detect surface charge accumulation in a chromium-doped (Bi,Sb)$_2$Te$_3$ thin film under oblique incidence, achieving sub-milliradian Faraday rotation precision. Unlike transport probes that require ultralow longitudinal conductivity, this optical technique is robust against finite $\sigma_L$, degrading by less than $0.3\%$ even when $\sigma_L \sim \sigma_T$. We extract the charge accumulation $\eta/B_z$ from the measured Faraday rotation and show results at $45^\circ$ and $60^\circ$ coincide within experimental uncertainty. Extending this to axion insulators, we predict that the TME produces an imaginary Faraday rotation linear in frequency, whose slope directly reflects the single-surface charge density. With improved sample thickness and precision, this optical scheme provides a viable pathway toward direct verification of the TME and four-dimensional quantum Hall effect.

cond-mat.mes-hall

Grokking or Glitching? How Low-Precision Drives Slingshot Loss Spikes

Deep neural networks exhibit periodic loss spikes during unregularized long-term training, a phenomenon known as the "Slingshot Mechanism." Existing work usually attributes this to intrinsic optimization dynamics, but its triggering mechanism remains unclear. This paper proves that this phenomenon is a result of floating-point arithmetic precision limits. As training enters a high-confidence stage, the difference between the correct-class logit and the other logits may exceed the absorption-error threshold. Then during backpropagation, the gradient of the correct class is rounded exactly to zero, while the gradients of the incorrect classes remain nonzero. This breaks the zero-sum constraint of gradients across classes and introduces a systematic drift in the parameter update of the classifier layer. We prove that this drift forms a positive feedback loop with the feature, causing the global classifier mean and the global feature mean to grow exponentially. We call this mechanism Numerical Feature Inflation (NFI). This mechanism explains the rapid norm growth before a Slingshot spike, the subsequent reappearance of gradients, and the resulting loss spike. We further show that NFI is not equivalent to an observed loss spike: in more practical tasks, partial absorption may not produce visible spikes, but it can still break the zero-sum constraint and drive rapid growth of parameter norms. Our results reinterpret Slingshot as a numerical dynamic of finite-precision training, and provide a testable explanation for abnormal parameter growth and logit divergence in late-stage training.

cs.LG

Decoding Scientific Experimental Images: The SPUR Benchmark for Perception, Understanding, and Reasoning

We introduce SPUR, a comprehensive benchmark for scientific experimental image perception, understanding, and reasoning, comprising 4,264 question-answering (QA) pairs derived from 1,084 expert-curated images. SPUR features three key innovations: (1) Panel-Level Fine-Grained Perception: evaluating the visual perception of multimodal large language models (MLLMs) across three dimensions (numerical, morphological, and information localization) on six fine-grained panel types; (2) Cross-Panel Relation Understanding: utilizing complex images with an average of 14.3 panels per sample to evaluate MLLMs' ability to decipher intricate cross-panel relations; (3) Expert-Level Reasoning: assessment of qualitative and quantitative reasoning across five experimental paradigms to determine if models can infer conclusions from evidence as human experts do. Comprehensive evaluation of 20 MLLMs and four multimodal Chain-of-Thought (MCoT) methods reveals that current models fall significantly short of the expert-level requirements for scientific image interpretation, underscoring a critical bottleneck in AI for Science (AI4S) research.

cs.CV

AEGIS: A Holistic Benchmark for Evaluating Forensic Analysis of AI-Generated Academic Images

We introduce AEGIS, A holistic benchmark for Evaluating forensic analysis of AI-Generated academic ImageS. Compared to existing benchmarks, AEGIS features three key advances: (1) Domain-Specific Complexity: covering seven academic categories with 39 fine-grained subtypes, exposing intrinsic forensic difficulty, where even GPT-5.1 reaches 48.80% overall performance and expert models achieve only limited localization accuracy (IoU 30.09%); (2) Diverse Forgery Simulations: modeling four prevalent academic forgery strategies across 25 generative models, with 11 yielding average forensic accuracy below 50%, showing that forensics lag behind generative advances; and (3) Multi-Dimensional Forensic Evaluation: jointly assessing detection, reasoning, and localization, revealing complementary strengths between model families, with multimodal large language models (MLLMs) at 84.74% accuracy in textual artifact recognition and expert detectors peaking at 79.54% accuracy in binary authenticity detection. By evaluating 25 leading MLLMs, nine expert models, and one unified multimodal understanding and generation model, AEGIS serves as a diagnostic testbed exposing fundamental limitations in academic image forensics.

cs.CV

THEMIS: Towards Holistic Evaluation of MLLMs for Scientific Paper Fraud Forensics

We present THEMIS, a novel multi-task benchmark designed to comprehensively evaluate multimodal large language models (MLLMs) on visual fraud reasoning within real-world academic scenarios. Compared to existing benchmarks, THEMIS introduces three major advances. (1) Real-World Scenarios and Complexity: Our benchmark comprises over 4,000 questions spanning seven scenarios, derived from authentic retracted-paper cases and carefully curated multimodal synthetic data. With 60.47% complex-texture images, THEMIS bridges the critical gap between existing benchmarks and the complexity of real-world academic fraud. (2) Fraud-Type Diversity and Granularity: THEMIS systematically covers five challenging fraud types and introduces 16 fine-grained manipulation operations. On average, each sample undergoes multiple stacked manipulation operations, with the diversity and difficulty of these manipulations demanding a high level of visual fraud reasoning from the models. (3) Multi-Dimensional Capability Evaluation: We establish a mapping from fraud types to five core visual fraud reasoning capabilities, thereby enabling an evaluation that reveals the distinct strengths and specific weaknesses of different models across these core capabilities. Experiments on 16 leading MLLMs show that even the best-performing model, GPT-5, achieves an overall performance of only 56.15%, demonstrating that our benchmark presents a stringent test. We expect THEMIS to advance the development of MLLMs for complex, real-world fraud reasoning tasks.

cs.CV

Observation of Superfluidity and Meissner Effect of Composite Bosons in GaAs Quantum Hall System

The quantum Hall effect (QHE) is theoretically understood as a superfluid condensate of composite bosons (CBs) -- bound states of electrons and magnetic flux quanta. While dissipationless transport is consistent with this picture, other signatures of superfluidity, such as the Meissner effect, remain elusive. Here, we present direct experimental evidence for CB superfluidity by probing the system's response to a controlled, time-varying magnetic field in Corbino disk geometries. We simultaneously observe the quantized Laughlin charge pumping and a new, quantized charge accumulation phenomenon, governed by the relation $\Delta Q_{\rm a}/e = \nu\,(\Delta \Phi/\Phi_0)$. This relation signifies that the system actively maintains the fixed electron-to-flux ratio that defines the CBs, neutralizing excess flux by drawing in a precise number of electrons. Crucially, devices with multiple concentric top gates reveal that this charge accumulation is uniformly distributed across the bulk of the QHE fluid, demonstrating that it is a collective, bulk property rather than an edge effect -- a key signature of a superfluid condensate. Furthermore, the presence of a top gate determines the screening mechanism: in a "grand canonical" setting with a gate, low Coulomb energy favors a charge-mediated screening (generalized Meissner effect); without a gate, the system enters a "canonical" regime, exhibiting fixed electron density like type-II superconductors. These observations confirm the CB superfluid nature of the QHE ground state and establish a versatile platform for studying macroscopic quantum coherence and its screening transitions in two dimensions.

cond-mat.mes-hall

Topological Surface Charge Detection via Active Capacitive Compensation: A Pathway to the 4D Quantum Hall Effect

The topological magnetoelectric effect (TME) in three-dimensional topological insulators (TIs), described by $\Delta P = \frac{e^2}{2h} N_{\rm Ch}^{(2)} \Delta B$, serves as a condensed-matter realization of the four-dimensional quantum Hall effect (4D QHE). In dual-gate axion-insulator devices, the TME-induced polarization yields a current $I_{\rm TME} \propto (C_{\rm total}/C_{\rm S})\,Q_{\rm 4D\mathrm{-}QHE}$, where the signal is suppressed by the capacitance ratio $C_{\rm total}/C_{\rm S}$. Here we propose an active compensation scheme that introduces a tunable negative capacitance $C_{\rm comp} \approx -C_{\rm gate}$ into the gate line, effectively canceling the gate dielectric capacitance and driving $C_{\rm total}/C_{\rm S} \to 1$. We validate the method using a quantum anomalous Hall (QAH) device, which shares the same surface-state physics as the axion insulator but permits direct charge measurement via a single gate, recovering over $95\%$ of the quantized charge signal from an initially half-attenuated state. This compensation method provides a robust means of resolving minute TME signals, offering a promising pathway toward direct measurements of the 4D QHE.

cond-mat.mes-hall

FinMMDocR: Benchmarking Financial Multimodal Reasoning with Scenario Awareness, Document Understanding, and Multi-Step Computation

We introduce FinMMDocR, a novel bilingual multimodal benchmark for evaluating multimodal large language models (MLLMs) on real-world financial numerical reasoning. Compared to existing benchmarks, our work delivers three major advancements. (1) Scenario Awareness: 57.9% of 1,200 expert-annotated problems incorporate 12 types of implicit financial scenarios (e.g., Portfolio Management), challenging models to perform expert-level reasoning based on assumptions; (2) Document Understanding: 837 Chinese/English documents spanning 9 types (e.g., Company Research) average 50.8 pages with rich visual elements, significantly surpassing existing benchmarks in both breadth and depth of financial documents; (3) Multi-Step Computation: Problems demand 11-step reasoning on average (5.3 extraction + 5.7 calculation steps), with 65.0% requiring cross-page evidence (2.4 pages average). The best-performing MLLM achieves only 58.0% accuracy, and different retrieval-augmented generation (RAG) methods show significant performance variations on this task. We expect FinMMDocR to drive improvements in MLLMs and reasoning-enhanced methods on complex multimodal reasoning tasks in real-world scenarios.

cs.CV

Investigating the relationship between the Weyl semimetal phase and the three-dimensional quantum Hall phase in ZrTe$_5$

The material ZrTe$_5$ exhibits distinct topological phases, including a Weyl semimetal phase, characterized by a chiral anomaly and in-plane Hall effect, and a three-dimensional quantum Hall phase. The relationship between these phases remains poorly understood. This work systematically explores their connection in ZrTe$_5$ through rotatable, pressure-dependent measurements. At ambient pressure, both phases are observed; the WSM phase requires strong electronic polarization, while the 3D QH phase appears when the characteristic resistivity peak temperature $T_p$ is approximately 90 K. Under applied pressure, the polarization diminishes, weakening the WSM phase and its associated nontrivial Hall signals. Concurrently, $T_p$ rises dramatically from 2 K at ambient pressure to 70 K at 2.2 GPa, approaching the expected regime for the 3D QH phase. These findings clarify the conditions underlying the WSM and 3D QH phases and suggest that exploring the 3D QH phase at even higher pressures is a promising direction for future research.

cond-mat.other

Observation of Quantized Charge Accumulation in a Quantum Anomalous Hall System

The quantum anomalous Hall effect in magnetically doped topological insulators exhibits a quantized Hall conductance $\sigma_{xy} = e^2/h$ arising from the two-dimensional surface states. While conventional transport probes confirm this quantization, they remain insensitive to the field-induced surface charge accumulation as a direct manifestation of $\sigma_{xy}$. Here, we experimentally validate an out-of-plane capacitive method that directly detects this quantized charge accumulation in a quantum anomalous Hall system. Using Corbino and simple disk devices, we measure charge accumulation proportional to field variation $\Delta B$, with dissipation characterized by longitudinal conductance $\sigma_{xx}$ and frequency $f$. A quantitative dissipation model extracts the intrinsic quantized charge density $\eta_0 = (e^2/h)\Delta B$, which is confirmed through $f$- and $\sigma_{xx}$-dependent measurements. Under ultra-low dissipation ($\sigma_{xx} \approx 10^{-9}$ S), we directly resolve the fully quantized charge accumulation. This methodology establishes a direct charge-accumulation probe and provides a pathway toward detecting the topological magnetoelectric effect, a condensed matter manifestation of the four-dimensional quantum Hall effect.

cond-mat.mes-hall

Triad: Empowering LMM-based Anomaly Detection with Vision Expert-guided Visual Tokenizer and Manufacturing Process

Although recent methods have tried to introduce large multimodal models (LMMs) into industrial anomaly detection (IAD), their generalization in the IAD field is far inferior to that for general purposes. We summarize the main reasons for this gap into two aspects. On one hand, general-purpose LMMs lack cognition of defects in the visual modality, thereby failing to sufficiently focus on defect areas. Therefore, we propose to modify the AnyRes structure of the LLaVA model, providing the potential anomalous areas identified by existing IAD models to the LMMs. On the other hand, existing methods mainly focus on identifying defects by learning defect patterns or comparing with normal samples, yet they fall short of understanding the causes of these defects. Considering that the generation of defects is closely related to the manufacturing process, we propose a manufacturing-driven IAD paradigm. An instruction-tuning dataset for IAD (InstructIAD) and a data organization approach for Chain-of-Thought with manufacturing (CoT-M) are designed to leverage the manufacturing process for IAD. Based on the above two modifications, we present Triad, a novel LMM-based method incorporating an expert-guided region-of-interest tokenizer and manufacturing process for industrial anomaly detection. Extensive experiments show that our Triad not only demonstrates competitive performance against current LMMs but also achieves further improved accuracy when equipped with manufacturing processes. Source code, training data, and pre-trained models will be publicly available at https://github.com/tzjtatata/Triad.

cs.CV

Myriad: Large Multimodal Model by Applying Vision Experts for Industrial Anomaly Detection

Due to the training configuration, traditional industrial anomaly detection (IAD) methods have to train a specific model for each deployment scenario, which is insufficient to meet the requirements of modern design and manufacturing. On the contrary, large multimodal models~(LMMs) have shown eminent generalization ability on various vision tasks, and their perception and comprehension capabilities imply the potential of applying LMMs on IAD tasks. However, we observe that even though the LMMs have abundant knowledge about industrial anomaly detection in the textual domain, the LMMs are unable to leverage the knowledge due to the modality gap between textual and visual domains. To stimulate the relevant knowledge in LMMs and adapt the LMMs towards anomaly detection tasks, we introduce existing IAD methods as vision experts and present a novel large multimodal model applying vision experts for industrial anomaly detection~(abbreviated to {Myriad}). Specifically, we utilize the anomaly map generated by the vision experts as guidance for LMMs, such that the vision model is guided to pay more attention to anomalous regions. Then, the visual features are modulated via an adapter to fit the anomaly detection tasks, which are fed into the language model together with the vision expert guidance and human instructions to generate the final outputs. Extensive experiments are applied on MVTec-AD, VisA, and PCB Bank benchmarks demonstrate that our proposed method not only performs favorably against state-of-the-art methods, but also inherits the flexibility and instruction-following ability of LMMs in the field of IAD. Source code and pre-trained models are publicly available at \url{https://github.com/tzjtatata/Myriad}.

cs.CV

Unprejudiced Training Auxiliary Tasks Makes Primary Better: A Multi-Task Learning Perspective

Human beings can leverage knowledge from relative tasks to improve learning on a primary task. Similarly, multi-task learning methods suggest using auxiliary tasks to enhance a neural network's performance on a specific primary task. However, previous methods often select auxiliary tasks carefully but treat them as secondary during training. The weights assigned to auxiliary losses are typically smaller than the primary loss weight, leading to insufficient training on auxiliary tasks and ultimately failing to support the main task effectively. To address this issue, we propose an uncertainty-based impartial learning method that ensures balanced training across all tasks. Additionally, we consider both gradients and uncertainty information during backpropagation to further improve performance on the primary task. Extensive experiments show that our method achieves performance comparable to or better than state-of-the-art approaches. Moreover, our weighting strategy is effective and robust in enhancing the performance of the primary task regardless the noise auxiliary tasks' pseudo labels.

cs.CV

Class Balance Matters to Active Class-Incremental Learning

Few-Shot Class-Incremental Learning has shown remarkable efficacy in efficient learning new concepts with limited annotations. Nevertheless, the heuristic few-shot annotations may not always cover the most informative samples, which largely restricts the capability of incremental learner. We aim to start from a pool of large-scale unlabeled data and then annotate the most informative samples for incremental learning. Based on this premise, this paper introduces the Active Class-Incremental Learning (ACIL). The objective of ACIL is to select the most informative samples from the unlabeled pool to effectively train an incremental learner, aiming to maximize the performance of the resulting model. Note that vanilla active learning algorithms suffer from class-imbalanced distribution among annotated samples, which restricts the ability of incremental learning. To achieve both class balance and informativeness in chosen samples, we propose Class-Balanced Selection (CBS) strategy. Specifically, we first cluster the features of all unlabeled images into multiple groups. Then for each cluster, we employ greedy selection strategy to ensure that the Gaussian distribution of the sampled features closely matches the Gaussian distribution of all unlabeled features within the cluster. Our CBS can be plugged and played into those CIL methods which are based on pretrained models with prompts tunning technique. Extensive experiments under ACIL protocol across five diverse datasets demonstrate that CBS outperforms both random selection and other SOTA active learning approaches. Code is publicly available at https://github.com/1170300714/CBS.

cs.CV

On Steering Multi-Annotations per Sample for Multi-Task Learning

The study of multi-task learning has drawn great attention from the community. Despite the remarkable progress, the challenge of optimally learning different tasks simultaneously remains to be explored. Previous works attempt to modify the gradients from different tasks. Yet these methods give a subjective assumption of the relationship between tasks, and the modified gradient may be less accurate. In this paper, we introduce Stochastic Task Allocation~(STA), a mechanism that addresses this issue by a task allocation approach, in which each sample is randomly allocated a subset of tasks. For further progress, we propose Interleaved Stochastic Task Allocation~(ISTA) to iteratively allocate all tasks to each example during several consecutive iterations. We evaluate STA and ISTA on various datasets and applications: NYUv2, Cityscapes, and COCO for scene understanding and instance segmentation. Our experiments show both STA and ISTA outperform current state-of-the-art methods. The code will be available.

cs.CV