SearcharxivSearch

arXiv subjects

Chunyi Li

Publications and source records attributed to Chunyi Li.

At least 19 recordsLinked to original sources

SafeAtlas-VL: Beyond Binary Multimodal Safety with Large-Scale Data and Guard Models

Multimodal safety moderation requires distinguishing risks arising from visual content, user intent, and assistant behavior. Existing safeguards, however, are typically trained for a single judgment target and reduce safety assessment to a binary decision. Consequently, risk becomes difficult to compare across a multimodal interaction, and ambiguous cases are obscured. We introduce SafeAtlas-VL, a dataset of 1.5M training instances that places image-, request-, and response-level judgments on a five-level ordered scale. We curate a broad collection of safety-relevant data from both real-world and synthetic sources and apply a disagreement-aware annotation procedure. The resulting dataset spans 15 harm categories and 55 fine-grained subcategories, covering a broad range of multimodal safety scenarios. We also construct SafeAtlas-Bench, a held-out set of 5,000 instances for evaluating five-level predictions and continuous risk scores. Upon this dataset, we train the SafeAtlas Guard series of models via target-conditioned tuning for multimodal safety detection. Our models not only perform five-way classification of safety levels but also map safety to continuous scores through a soft cumulative ordinal head. Experimental results demonstrate that guard models trained on our dataset exhibit strong generalization: even without using the training sets of other benchmarks, they achieve competitive performance on the corresponding test sets. Notably, our 8B model attains the overall best performance, outperforming the previous SOTA by approximately 4% in F1 score. Code, data, and models are released to support further research. Warning: this paper contains example data that may be offensive, harmful, graphic, or disturbing.

cs.AI

Syzygies of Polarized Abelian Surfaces: A Reider-Type Criterion

Let $(X,L)$ be a polarized complex abelian surface with $L^2=2d$. We establish a Reider-type criterion for Property $N_p$. If $d\geq7$, then $L$ satisfies Property $N_0$ if and only if there is no elliptic curve $E\subseteq X$ with $L\cdot E\leq2$, with one explicitly described exception. If $p\geq1$ and $d\geq(p+2)^2+1$, then $L$ satisfies Property $N_p$ if and only if there is no elliptic curve $E\subseteq X$ with $L\cdot E\leq p+2$. The numerical bounds on $d$ are optimal for $p=0,1$. These results improve upon earlier work of K\"{u}ronya--Lozovanu, Ito, and Rojas. We also construct a polarized abelian surface whose basepoint-freeness threshold is irrational.

math.AG

Stability conditions and moduli spaces on projective families

We extend the construction of stability conditions on projective schemes over a field to projective families over an arbitrary base, and prove that they admit proper relative moduli spaces of semistable objects. We also prove a number of complementary results: the existence of mass-Hom bounds for these stability conditions, as conjectured by Halpern-Leistner and Robotis; a comparison with tilt-stability on surfaces and threefolds; a construction of stability conditions on the supported derived category of total spaces of certain vector bundles, including all local Calabi--Yau varieties; and a simple new proof of Bondal and Orlov's reconstruction theorem.

math.AG

$c$-axis strain tuning of superconductivity and symmetric elastoresistivity in CsV$_3$Sb$_5$

The kagome metal CsV$_{3}$Sb$_{5}$ hosts an intriguing interplay between charge-density-wave (CDW) order and superconductivity that is highly sensitive to lattice distortions. However, determining the specific roles of the in-plane ($A_{1g,1}$) and out-of-plane ($A_{1g,2}$) symmetric strain channels has been hindered by their intrinsic mixing in conventional piezo-based experiments. Here, we combine in-plane uniaxial strain with direct $c$-axis compression to independently access and disentangle these symmetry-resolved responses in CsV$_{3}$Sb$_{5}$. We reveal that $c$-axis compression drives a massive, linear enhancement of the superconducting transition temperature ($T_c$) alongside a suppression of $T_{\rm CDW}$. The tuning efficiency of this out-of-plane deformation acts with an opposite sign and far exceeds that of in-plane strain, demonstrating that $c$-axis lattice control dictates the phase competition. Furthermore, by isolating the pure elastoresistivity coefficients, we find that the out-of-plane cross-coupling coefficient ($m_{13}$) is comparable in magnitude but opposite in sign to the in-plane response ($m_{11}+m_{12}$). Unlike the sharply peaked in-plane response, $m_{13}$ exhibits a distinct, order-parameter-like onset across the CDW transition. Our results establish that out-of-plane lattice control plays a dominant role in tuning the intertwined states in CsV$_{3}$Sb$_{5}$ and provide a general pathway for resolving strain-coupled electronic responses in layered quantum materials.

cond-mat.supr-con

Towards Characterizing Scientific Image Utility and Upgradability

Scientific images function as critical evidence in research communication, yet their integrity faces unprecedented threats from AI-generated content that introduces subtle but consequential errors. Existing evaluation paradigms prove inadequate: perceptual quality metrics poorly correlate with scientific validity, while language models lack domain-specific verification capabilities. To address this gap, we propose the \textbf{S}cientific \textbf{I}mage \textbf{U}tility and \textbf{U}pgradability \textbf{A}ssessment (\textbf{SIU$^2$A}) framework, which introduces two complementary dimensions for scientific image evaluation. \textbf{Utility} encompasses \textit{error detection} (identifying scientific inaccuracies) and \textit{correction feasibility} (assessing whether errors can be reliably repaired). \textbf{Upgradability} measures the quality of correction. We categorize scientific image corruption into four fundamental types: Detail Distortion, Incompleteness, False Content, and Entity Confusion. Based on this taxonomy, we construct SIU$^2$A-Benchmark, a dataset with expert annotations for error identification and repair. The framework implements a two-stage evaluation protocol: the \textit{Utility} stage evaluates error detection capability and repair instruction generation, while the \textit{Upgradability} stage assesses whether corrections faithfully restore scientific validity without compromising existing accurate information. Experiments reveal that current multimodal systems exhibit significant limitations in both scientific error assessment and faithful correction, exposing a fundamental gap between visual perception and scientific usability.

cs.CV

GeoR-Bench: Evaluating Geoscience Visual Reasoning

Geoscience intelligence is expected to understand, reason about, and predict earth system changes to support human decision-making in critical domains such as disaster response, climate adaptation and environmental protection. Although current research has shown promising progress on specific geoscience tasks, such as remote sensing interpretation, geographic question-answering, existing benchmarks remain largely task-specific which failing to capture the open-ended real world geoscience problems. As a result, it remains unclear how far current AI systems are from achieving genuine geoscience intelligence. To address this gap, we present \textbf{GeoR-Bench}, a \underline{Bench}mark for evaluating \underline{Geo}science visual \underline{R}easoning through reasoning informed visual editing tasks. GeoR-Bench contains 440 curated samples spanning 6 geoscience categories and 24 task types, covering earth observation imagery and structured scientific representations such as maps and diagrams. We evaluate outputs along three dimensions, including reasoning, consistency, and quality. Benchmark results of 21 closed- and open-source multimodal models reveal that geoscience reasoning remains a critical bottleneck. The highest-performing model achieves 42.7\% overall strict accuracy, while the best open-source models only get 10.3\%. Notably, the visual consistency and image quality of the outputs frequently surpass their scientific accuracy. Ultimately, these findings indicate that current models generate superficially plausible results but fail to capture underlying earth science processes.

cs.CV

MirrorBench: Evaluating Self-centric Intelligence in MLLMs by Introducing a Mirror

Recent progress in Multimodal Large Language Models (MLLMs) has demonstrated remarkable advances in perception and reasoning, suggesting their potential for embodied intelligence. While recent studies have evaluated embodied MLLMs in interactive settings, current benchmarks mainly target capabilities to perceive, understand, and interact with external objects, lacking a systematic evaluation of self-centric intelligence. To address this, we introduce MirrorBench, a simulation-based benchmark inspired by the classical Mirror Self-Recognition (MSR) test in psychology. MirrorBench extends this paradigm to embodied MLLMs through a tiered framework of progressively challenging tasks, assessing agents from basic visual perception to high-level self-representation. Experiments on leading MLLMs show that even at the lowest level, their performance remains substantially inferior to human performance, revealing fundamental limitations in self-referential understanding. Our study bridges psychological paradigms and embodied intelligence, offering a principled framework for evaluating the emergence of general intelligence in large models. Project page: https://fflahm.github.io/mirror-bench-page/.

cs.AI

SIQA: Toward Reliable Scientific Image Quality Assessment

Scientific images fundamentally differ from natural and AI-generated images in that they encode structured domain knowledge rather than merely depict visual scenes. Assessing their quality therefore requires evaluating not only perceptual fidelity but also scientific correctness and logical completeness. However, existing image quality assessment (IQA) paradigms primarily focus on perceptual distortions or image-text alignment, implicitly assuming that depicted content is factually valid. This assumption breaks down in scientific contexts, where visually plausible figures may still contain conceptual errors or incomplete reasoning. To address this gap, we introduce Scientific Image Quality Assessment (SIQA), a framework that models scientific image quality along two complementary dimensions: Knowledge (Scientific Validity and Scientific Completeness) and Perception (Cognitive Clarity and Disciplinary Conformity). To operationalize this formulation, we design two evaluation protocols: SIQA-U (Understanding), which measures semantic comprehension of scientific content through multiple-choice tasks, and SIQA-S (Scoring), which evaluates alignment with expert quality judgments. We further construct the SIQA Challenge, consisting of an expert-annotated benchmark and a large-scale training set. Experiments across representative multimodal large language models (MLLMs) reveal a consistent discrepancy between scoring alignment and scientific understanding. While models can achieve strong agreement with expert ratings under SIQA-S, their performance on SIQA-U remains substantially lower. Fine-tuning improves both metrics, yet gains in scoring consistently outpace improvements in understanding. These results suggest that rating consistency alone may not reliably reflect scientific comprehension, underscoring the necessity of multidimensional evaluation for scientific image quality assessment.

cs.CV

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond

The success of large language models (LLMs) in scientific domains has heightened safety concerns, prompting numerous benchmarks to evaluate their scientific safety. Existing benchmarks often suffer from limited risk coverage and a reliance on subjective evaluation. To address these problems, we introduce SafeSci, a comprehensive framework for safety evaluation and enhancement in scientific contexts. SafeSci comprises SafeSciBench, a multi-disciplinary benchmark with 0.25M samples, and SafeSciTrain, a large-scale dataset containing 1.5M samples for safety enhancement. SafeSciBench distinguishes between safety knowledge and risk to cover extensive scopes and employs objective metrics such as deterministically answerable questions to mitigate evaluation bias. We evaluate 24 advanced LLMs, revealing critical vulnerabilities in current models. We also observe that LLMs exhibit varying degrees of excessive refusal behaviors on safety-related issues. For safety enhancement, we demonstrate that fine-tuning on SafeSciTrain significantly enhances the safety alignment of models. Finally, we argue that knowledge is a double-edged sword, and determining the safety of a scientific question should depend on specific context, rather than universally categorizing it as safe or unsafe. Our work provides both a diagnostic tool and a practical resource for building safer scientific AI systems.

cs.LG

Doping evolution of spin excitations in La$_{3-x}$Sr$_{x}$Ni$_2$O$_7$/SrLaAlO$_4$ superconducting thin films

Ambient-pressure superconductivity in compressively strained bilayer nickelate films provides a unique platform to test pairing scenarios, yet the evolution of magnetism with carrier doping remains largely unexplored. Here, we utilize Ni $L_3$-edge resonant inelastic x-ray scattering to systematically track the evolution of spin and electronic excitations in coherently strained La$_{3-x}$Sr$_x$Ni$_2$O$_7$/SrLaAlO$_4$ thin films, spanning the superconducting ($x \le 0.21$) and overdoped non-superconducting ($x = 0.38$) regimes. We reveal that dispersive spin excitations, characterized by double-stripe correlations and nearly doping-independent exchange scales, persist robustly throughout the entire superconducting dome. In stark contrast, upon entering the overdoped non-superconducting state, this coherent magnetic framework undergoes an abrupt collapse, melting into a heavily damped, low-spectral-weight continuum. We show that this magnetic breakdown is fundamentally driven by a selective doping-induced orbital reconstruction. While the invariant $\sim\!1.0$~eV intra-atomic $dd$ peak confirms an intact local octahedral crystal field, the concurrent quenching of the $\sim\!0.4$~eV and $\sim\!1.6$~eV features signifies a severe degradation of the apical-oxygen-mediated $d_{z^2}$--$p_z$--$d_{z^2}$ singlet sector and bilayer charge-transfer coherence. The synchronized demise of coherent spin excitations and macroscopic pairing establishes a direct, doping-controlled link, underscoring that maintaining the localized $d_{z^2}$ magnetic framework and robust apical-oxygen coupling is the fundamental prerequisite for high-$T_c$ superconductivity in bilayer nickelates.

cond-mat.supr-con

STAR : Bridging Statistical and Agentic Reasoning for Large Model Performance Prediction

As comprehensive large model evaluation becomes prohibitively expensive, predicting model performance from limited observations has become essential. However, existing statistical methods struggle with pattern shifts, data sparsity, and lack of explanation, while pure LLM methods remain unreliable. We propose STAR, a framework that bridges data-driven STatistical expectations with knowledge-driven Agentic Reasoning. STAR leverages specialized retrievers to gather external knowledge and embeds semantic features into Constrained Probabilistic Matrix Factorization (CPMF) to generate statistical expectations with uncertainty. A reasoning module guided by Expectation Violation Theory (EVT) then refines predictions through intra-family analysis, cross-model comparison, and credibility-aware aggregation, producing adjustments with traceable explanations. Extensive experiments show that STAR consistently outperforms all baselines on both score-based and rank-based metrics, delivering a 14.46% gain in total score over the strongest statistical method under extreme sparsity, with only 1--2 observed scores per test model.

cs.AI

Free-GVC: Towards Training-Free Extreme Generative Video Compression with Temporal Coherence

Building on recent advances in video generation, generative video compression has emerged as a new paradigm for achieving visually pleasing reconstructions. However, existing methods exhibit limited exploitation of temporal correlations, causing noticeable flicker and degraded temporal coherence at ultra-low bitrates. In this paper, we propose Free-GVC, a training-free generative video compression framework that reformulates video coding as latent trajectory compression guided by a video diffusion prior. Our method operates at the group-of-pictures (GOP) level, encoding video segments into a compact latent space and progressively compressing them along the diffusion trajectory. To ensure perceptually consistent reconstruction across GOPs, we introduce an Adaptive Quality Control module that dynamically constructs an online rate-perception surrogate model to predict the optimal diffusion step for each GOP. In addition, an Inter-GOP Alignment module establishes frame overlap and performs latent fusion between adjacent groups, thereby mitigating flicker and enhancing temporal coherence. Experiments show that Free-GVC achieves an average of 93.29% BD-Rate reduction in DISTS over the latest neural codec DCVC-RT, and a user study further confirms its superior perceptual quality and temporal coherence at ultra-low bitrates.

cs.CV

Automated Safety Benchmarking: A Multi-agent Pipeline for LVLMs

Large vision-language models (LVLMs) exhibit remarkable capabilities in cross-modal tasks but face significant safety challenges, which undermine their reliability in real-world applications. Efforts have been made to build LVLM safety evaluation benchmarks to uncover their vulnerability. However, existing benchmarks are hindered by their labor-intensive construction process, static complexity, and limited discriminative power. Thus, they may fail to keep pace with rapidly evolving models and emerging risks. To address these limitations, we propose VLSafetyBencher, the first automated system for LVLM safety benchmarking. VLSafetyBencher introduces four collaborative agents: Data Preprocessing, Generation, Augmentation, and Selection agents to construct and select high-quality samples. Experiments validates that VLSafetyBencher can construct high-quality safety benchmarks within one week at a minimal cost. The generated benchmark effectively distinguish safety, with a safety rate disparity of 70% between the most and least safe models.

cs.CL

Stability conditions on products of curves and Hilbert schemes of surfaces

We prove that stability conditions on the derived category of a product of curves of positive genus are uniquely determined by their central charge and the phase of skyscraper sheaves. As an application, we construct stability conditions on Hilbert schemes of points on certain surfaces, including some K3 surfaces of Kummer type.

math.AG

Using GUI Agent for Electronic Design Automation

Graphical User Interface (GUI) agents adopt an end-to-end paradigm that maps a screenshot to an action sequence, thereby automating repetitive tasks in virtual environments. However, existing GUI agents are evaluated almost exclusively on commodity software such as Microsoft Word and Excel. Professional Computer-Aided Design (CAD) suites promise an order-of-magnitude higher economic return, yet remain the weakest performance domain for existing agents and are still far from replacing expert Electronic-Design-Automation (EDA) engineers. We therefore present the first systematic study that deploys GUI agents for EDA workflows. Our contributions are: (1) a large-scale dataset named GUI-EDA, including 5 CAD tools and 5 physical domains, comprising 2,000+ high-quality screenshot-answer-action pairs recorded by EDA scientists and engineers during real-world component design; (2) a comprehensive benchmark that evaluates 30+ mainstream GUI agents, demonstrating that EDA tasks constitute a major, unsolved challenge; and (3) an EDA-specialized metric named EDAgent, equipped with a reflection mechanism that achieves reliable performance on industrial CAD software and, for the first time, outperforms Ph.D. students majored in Electrical Engineering. This work extends GUI agents from generic office automation to specialized, high-value engineering domains and offers a new avenue for advancing EDA productivity. The dataset will be released at: https://github.com/aiben-ch/GUI-EDA.

cs.CV

Embodied Image Compression

Image Compression for Machines (ICM) has emerged as a pivotal research direction in the field of visual data compression. However, with the rapid evolution of machine intelligence, the target of compression has shifted from task-specific virtual models to Embodied agents operating in real-world environments. To address the communication constraints of Embodied AI in multi-agent systems and ensure real-time task execution, this paper introduces, for the first time, the scientific problem of Embodied Image Compression. We establish a standardized benchmark, EmbodiedComp, to facilitate systematic evaluation under ultra-low bitrate conditions in a closed-loop setting. Through extensive empirical studies in both simulated and real-world settings, we demonstrate that existing Vision-Language-Action models (VLAs) fail to reliably perform even simple manipulation tasks when compressed below the Embodied bitrate threshold. We anticipate that EmbodiedComp will catalyze the development of domain-specific compression tailored for Embodied agents , thereby accelerating the Embodied AI deployment in the Real-world.

cs.CV

Q-REAL: Towards Realism and Plausibility Evaluation for AI-Generated Content

Quality assessment of AI-generated content is crucial for evaluating model capability and guiding model optimization. However, most existing quality assessment datasets and models provide only a single quality score, which is too coarse to offer targeted guidance for improving generative models. In current applications of AI-generated images, realism and plausibility are two critical dimensions, and with the emergence of unified generation-understanding models, fine-grained evaluation along these dimensions becomes especially effective for improving generative performance. Therefore, we introduce Q-Real, a novel dataset for fine-grained evaluation of realism and plausibility in AI-generated images. Q-Real consists of 3,088 images generated by popular text-to-image models. For each image, we annotate the locations of major entities and provide a set of judgment questions and attribution descriptions for these along the dimensions of realism and plausibility. Considering that recent advances in multi-modal large language models (MLLMs) enable fine-grained evaluation of AI-generated images, we construct Q-Real Bench to evaluate them on two tasks: judgment and grounding with reasoning. Finally, to enhance MLLM capabilities, we design a fine-tuning framework and conduct experiments on multiple MLLMs using our dataset. Experimental results demonstrate the high quality and significance of our dataset and the comprehensiveness of the benchmark. Dataset and code will be released upon publication.

cs.CV