SearcharxivSearch

arXiv subjects

Zhen Bi

Publications and source records attributed to Zhen Bi.

At least 19 recordsLinked to original sources

Memory Is Not Always Needed: Characterizing Conditional Memory in Scientific Reasoning

Scientific reasoning requires language models to retrieve specialized knowledge and incorporate it reliably into multi-step computation. Conditional memory provides an explicit lookup pathway that complements dense neural representations, but its usefulness is inherently input- and computation-dependent: retrieved information may repair missing scientific associations, yet it may also introduce distracting shortcuts or interfere with reasoning that the base model can already perform correctly. In this work, we systematically investigate when, where, and to what extent conditional memory should participate in scientific reasoning. We characterize the scientific knowledge boundary and controlled interventions on memory-enabled knowledge-circuit nodes. Based on these analyses, we propose a Knowledge Boundary-Aware Router that uses task-specific input proxies available before generation to determine whether memory is activated, which layer-stage nodes receive memory signals, and how strongly these signals contribute. Experiments on biological and chemical reasoning benchmarks, covering two backbone families and six task types, show that memory effects vary substantially across inputs, tasks, and injection locations. Compared with static and activation-rate-matched random routing, our approach more consistently preserves beneficial memory contributions while suppressing memory-induced regressions, establishing selective memory allocation as an important principle for reliable scientific reasoning.

cs.AI

Mixed-State Symmetry-Protected Topology and Strong-to-Weak Spontaneous Symmetry Breaking in a Superconducting Qubit Array

We experimentally investigate how symmetry-protected topological order in a one-dimensional cluster state is transformed by measurement and decoherence in a five-qubit superconducting array. We first characterize the state's nonlocal string order and show that controlled dephasing selectively suppresses one symmetry sector while leaving the other robust, consistent with average symmetry-protected topological order. We then measure one sublattice in a tunable basis and show that the remaining qubits are driven between a long-range-entangled GHZ state and a paramagnetic state. When the measurement record is discarded, the conventional long-range correlator vanishes while a nonlinear fidelity correlator remains finite, providing a finite-size signature of strong-to-weak spontaneous symmetry breaking. These experiments demonstrate how conditioning, averaging, and decoherence reveal distinct manifestations of order encoded in the same underlying cluster state, and establish a superconducting-circuit setting for probing mixed-state symmetry and topology.

quant-ph

VARM-Bench: Benchmarking Verifiable Structured Reasoning in Chinese Abusive Speech Moderation

The widespread circulation of abusive online content has increased the need for reliable moderation of Chinese social-media text. Existing Chinese benchmarks support label classification, fine-grained toxicity categorization, and target-aware extraction, but do not provide a unified representation for deterministically verifying the stated basis of a moderation decision. We introduce VARM-Bench, a benchmark for field-anchored chain-of-thought rationales in Chinese abusive-speech moderation. Each instance contains a concise natural-language rationale with explicit anchors for six decisions: target, target type, target explicitness, author stance, harmfulness label, and fine-grained category. Our deterministic protocol evaluates field correctness, target alignment, output validity, complete-record agreement, and hidden record errors conditioned on correct final decisions, without relying on an LLM judge. Under a common structured-output protocol, we evaluate language models across multiple model families using zero-shot prompting, taxonomy guidance, and structured CoT supervision, and analyze lexical-cue sensitivity and field-level errors. Results show that strong label-level performance can conceal substantial errors in complete moderation records. VARM-Bench provides an auditable and reproducible benchmark for evaluating verifiable moderation rationales in Chinese abusive-speech moderation.

cs.AI

WaveFilter: Enhancing the Long-Context Capability of Diffusion LLMs via Wavelet-Guided KV Cache Filtering

Diffusion Large Language Models (DLMs) have demonstrated significant advantages across various tasks. However, constrained by their multi-step iterative inference mechanism, their computational overhead and inference latency in long-context tasks have become core bottlenecks restricting their large-scale deployment. When processing long sequences, existing Key-Value (KV) caching mechanisms often face a dilemma where generation quality degrades drastically, where the core challenge lies in precisely and efficiently filtering critical tokens within ultra-long contexts. Inspired by the human reading process, we propose \textbf{WaveFilter}, a universal and training-free caching framework. This framework innovatively introduces the wavelet transform for decomposition of long sequences to achieve precise identification of key tokens, based on which a sparse KV Cache is constructed to compute the final contextual representation. Experimental results demonstrate that WaveFilter, as a plug-and-play generic framework, significantly enhances the performance of existing mainstream KV Cache methods in complex long-context tasks.

cs.CL

ConsisGuard: Aligning Safety Deliberation with Policy Enforcement in LLM Guardrails

Reasoning-based LLM guardrails improve safety moderation by generating explicit rationales before issuing final decisions. However, their rationales do not always lead to faithful enforcement: a model may recognize a harmful intent in its reasoning but still predict a safe label, or issue an unsafe decision without policy-grounded justification. We identify this safety-critical failure mode as the deliberation-to-enforcement gap. Unlike general chain-of-thought faithfulness, guardrail reliability requires policy execution consistency: the generated reasoning should be grounded in the safety policy, and the final decision should be entailed by that reasoning. We propose ConsisGuard, a consistency-aware framework for reasoning-based LLM guardrails. ConsisGuard performs Policy-to-Decision Trajectory Distillation and Functional Coupling Alignment, aligning the internal coupling between safety deliberation and decision enforcement. Experiments on prompt and response harmfulness detection benchmarks show that ConsisGuard improves detection performance while reducing policy execution failures. These results suggest that reliable reasoning-based guardrails require accurate faithful execution of safety policies.

cs.CL

Make LLM Learn to Synthesize from Streaming Experiences through Feedback

Large language models (LLMs) have been widely adopted for synthetic data generation, significantly reducing annotation costs. However, most existing studies treat synthesis as a set of isolated tasks and overlook a more fundamental question: whether a model can learn to synthesize by accumulating experience from past tasks and transferring it to future ones. In this work, we introduce StreamSynth, a new setting in which synthesis tasks arrive sequentially and experience from historical tasks provides informative signals for future synthesis. To address this setting, we propose SynLearner, a general framework that enables synthesis models to acquire reusable synthesis experience over a task stream. Instead of generating data independently for each task, SynLearner encourages the model to explore diverse synthesis patterns, learn from feedback, and balance sample quality with set-level diversity as tasks evolve. Extensive experiments across multiple benchmarks show that SynLearner effectively leverages experience from earlier tasks to improve synthesis performance on later ones, exhibiting consistent cross-task transferability. These findings provide evidence for the feasibility of StreamSynth and highlight synthetic data generation as an experience-driven process that can benefit from task streams.

cs.AI

Robust and Generalizable Safety Steering for Text-to-Image Diffusion Transformers

Diffusion Transformers have become a powerful backbone for text-to-image generation, but their layered and cross-modal generation process makes safety control fundamentally different from prompt-level filtering or output-level detection. Harmful semantics may be weakly expressed in text representations, progressively bound to visual latents, and finally entangled with rendering dynamics. As a result, safety steering at a fixed layer can be unstable, and a steering mechanism learned from known risks may not transfer reliably to a shifted target risk domain. We propose SafeDIG, a safety steering framework that formulates DiT safety adaptation as position-aware sparse feature transfer. SafeDIG first constructs Sparse Autoencoders over functionally distinct DiT intervention positions and uses robustness-aware pre-training routing to prioritize intervention sites that are expected to remain stable under source-target risk shift. It then separates transferable safety features from domain-specific activation geometry by freezing the SAE encoder as a reusable sparse safety dictionary and adapting only the decoder to the target-domain activation manifold. During inference, SafeDIG combines Blend and Repel operations to steer unsafe activations toward transferred safety manifolds or away from harmful sparse directions. Experiments on FLUX.1 Dev and Stable Diffusion 3.5 Large show that SafeDIG consistently reduces target-domain and overall unsafe generation rates while preserving source-domain safety and image quality.

cs.AI

Dissipative Preparation of Correlated Quantum States in Dipolar Rydberg Arrays

Preparing correlated quantum states is essential for emerging technologies, but remains challenging in many-body systems. Here we propose a dissipative protocol that engineers nonreciprocal, energy-selective transitions to steer dipolar quantum systems toward desired many-body states. This is realized by introducing two types of controllable dissipative auxiliary atoms that act as nonreciprocal excitation and de-excitation channels, respectively, enabling a directional walk in Hilbert space. This approach enables stabilization of states across the many-body spectrum, not limited to the ground state and requiring no \textit{a priori} knowledge of the Hamiltonian. Our approach is designed for neutral atoms in dipolar Rydberg arrays, but applies broadly to setups with similar capabilities, providing a flexible and scalable framework for state preparation in programmable platforms.

quant-ph

Strong-to-Weak Spontaneous Symmetry Breaking in a $(2+1)$D Transverse-Field Ising Model under Decoherence

Decoherence in many-body quantum systems can give rise to intrinsically mixed-state phases and phase transitions beyond the pure-state paradigm. Here we study the $(2+1)$D transverse-field Ising model subject to a strongly $\mathbb{Z}_2$-symmetric decoherence channel, with a focus on strong-to-weak spontaneous symmetry breaking (SWSSB). This problem is challenging because the relevant transitions occur in the strong-decoherence regime, beyond the reach of perturbative expansions around the pure-state limit, while conventional quantum Monte Carlo (QMC) methods are hampered by the need to access nonlinear observables and by the sign problem. We overcome these difficulties by developing a QMC algorithm that efficiently evaluates nonlinear R\'enyi-2 correlators in higher dimensions, complemented by an effective field-theoretic approach. We show that the decohered state realizes a rich mixed-state phase diagram governed by an effective 2D Ashkin-Teller theory. This theory enables analytical predictions for the mixed-state phases and the universality classes of the phase boundaries, all of which are confirmed by large-scale QMC simulations.

quant-ph

Matrix Product States for Modulated Symmetries: SPT, LSM, and Beyond

Matrix product states (MPS) provide a powerful framework for characterizing one-dimensional symmetry-protected topological (SPT) phases of matter and for formulating Lieb-Schultz-Mattis (LSM)-type constraints. Here we generalize the MPS formalism to translationally invariant systems with general modulated symmetries. We show that the standard symmetry "push-through" condition for conventional global symmetry must be revised to account for symmetry modulation, and we derive the appropriate generalized condition. Using this generalized push-through structure, we classify one-dimensional SPT phases with modulated symmetries and formulate LSM-type constraints within the same MPS-based framework.

cond-mat.str-el

Generalized symmetry-protected topological phases in mixed states from gauging dualities

Decoherence in realistic quantum platforms motivates a mixed-state notion of topological phases of matter, including average symmetry-protected topological (ASPT) phases. Alongside this progress, generalized symmetries--notably noninvertible and dipole symmetries--have become powerful organizing principles for exotic quantum phases, yet their implications for mixed states remain less explored. In this work, we bridge these directions through a gauging correspondence between mixed-state phases with generalized symmetries and mixed-state phases with ordinary group symmetries, recasting the classification of noninvertible and dipole ASPT phases into familiar classifications of symmetry breaking and ASPT phases with dual symmetries. Using this approach, we classify and construct a subclass of ASPT phases with non-invertible and dipole symmetries in $(1+1)d$, including phases that are intrinsic to mixed states, and characterize them via string order parameters and protected edge modes.

cond-mat.str-el

SkillNet: Create, Evaluate, and Connect AI Skills

Current AI agents can flexibly invoke tools and execute complex tasks, yet their long-term advancement is hindered by the lack of systematic accumulation and transfer of skills. Without a unified mechanism for skill consolidation, agents frequently ``reinvent the wheel'', rediscovering solutions in isolated contexts without leveraging prior strategies. To address this challenge, we introduce SkillNet, an open infrastructure for creating, evaluating, and organizing AI skills at scale. SkillNet structures skills within a unified ontology that supports creating skills from heterogeneous sources, establishing rich relational connections, and performing multi-dimensional evaluation across Safety, Completeness, Executability, Maintainability, and Cost-awareness. Our infrastructure integrates a repository of over 600,000 skills, an interactive platform, and a versatile Python toolkit. Experiments on ALFWorld, WebShop, and ScienceWorld show 40% higher average rewards and 30% fewer execution steps across multiple backbone models. Furthermore, SkillNet-Gym benchmarks skill retrieval, utilization, and composition, while SkillNet-Fabric enables task-specific skill routing through lightweight Wikis. By formalizing skills as evolving, composable assets, SkillNet provides a robust foundation for agents to move from transient experience to durable mastery.

cs.AI

RooflineBench: A Benchmarking Framework for On-Device LLMs via Roofline Analysis

The transition toward localized intelligence through Small Language Models (SLMs) has intensified the need for rigorous performance characterization on resource-constrained edge hardware. However, objectively measuring the theoretical performance ceilings of diverse architectures across heterogeneous platforms remains a formidable challenge. In this work, we propose a systematic framework based on the Roofline model that unifies architectural primitives and hardware constraints through the lens of operational intensity (OI). By defining an inference-potential region, we introduce the Relative Inference Potential as a novel metric to compare efficiency differences between Large Language Models (LLMs) on the same hardware substrate. Extensive empirical analysis across diverse compute tiers reveals that variations in performance and OI are significantly influenced by sequence length. We further identify a critical regression in OI as model depth increases. Additionally, our findings highlight an efficiency trap induced by hardware heterogeneity and demonstrate how structural refinements, such as Multi-head Latent Attention (MLA), can effectively unlock latent inference potential across various hardware substrates. These insights provide actionable directions for hardware-software co-design to align neural structures with physical constraints in on-device intelligence. The released code is available in the Appendix C.

cs.LG

Beyond Superficial Unlearning: Sharpness-Aware Robust Erasure of Hallucinations in Multimodal LLMs

Multimodal LLMs are powerful but prone to object hallucinations, which describe non-existent entities and harm reliability. While recent unlearning methods attempt to mitigate this, we identify a critical flaw: structural fragility. We empirically demonstrate that standard erasure achieves only superficial suppression, trapping the model in sharp minima where hallucinations catastrophically resurge after lightweight relearning. To ensure geometric stability, we propose SARE, which casts unlearning as a targeted min-max optimization problem and uses a Targeted-SAM mechanism to explicitly flatten the loss landscape around hallucinated concepts. By suppressing hallucinations under simulated worst-case parameter perturbations, our framework ensures robust removal stable against weight shifts. Extensive experiments demonstrate that SARE significantly outperforms baselines in erasure efficacy while preserving general generation quality. Crucially, it maintains persistent hallucination suppression against relearning and parameter updates, validating the effectiveness of geometric stabilization.

cs.LG

Your One-Stop Solution for AI-Generated Video Detection

Recent advances in generative modeling can create remarkably realistic synthetic videos, making it increasingly difficult for humans to distinguish them from real ones and necessitating reliable detection methods. However, two key limitations hinder the development of this field. \textbf{From the dataset perspective}, existing datasets are often limited in scale and constructed using outdated or narrowly scoped generative models, making it difficult to capture the diversity and rapid evolution of modern generative techniques. Moreover, the dataset construction process frequently prioritizes quantity over quality, neglecting essential aspects such as semantic diversity, scenario coverage, and technological representativeness. \textbf{From the benchmark perspective}, current benchmarks largely remain at the stage of dataset creation, leaving many fundamental issues and in-depth analysis yet to be systematically explored. Addressing this gap, we propose AIGVDBench, a benchmark designed to be comprehensive and representative, covering \textbf{31} state-of-the-art generation models and over \textbf{440,000} videos. By executing more than \textbf{1,500} evaluations on \textbf{33} existing detectors belonging to four distinct categories. This work presents \textbf{8 in-depth analyses} from multiple perspectives and identifies \textbf{4 novel findings} that offer valuable insights for future research. We hope this work provides a solid foundation for advancing the field of AI-generated video detection. Our benchmark is open-sourced at https://github.com/LongMa-2025/AIGVDBench.

cs.CV

Logical Structure as Knowledge: Enhancing LLM Reasoning via Structured Logical Knowledge Density Estimation

The reasoning capabilities of Large Language Models (LLMs) are increasingly attributed to training data quality rather than mere parameter scaling. However, existing data-centric paradigms often equate quality with factuality or diversity and ignore the internal logical complexity of training samples. In this work, we propose that natural language harbors Structured Logical Knowledge manifested through entailment relationships and logical topologies. To quantify this, we introduce Structured Logical Knowledge Density (SLKD), a novel metric that measures logical information content by decomposing natural language into executable predicates and logical primitives. Our analysis reveals a significant logical disparity in current datasets where sparse logical signals predominate. Consequently, we propose a density aware re-cognizing optimization strategy that prioritizes high-density logical samples to enhance with the LLM's reasoning ability. Extensive experiments demonstrate that our approach enhances reasoning performance and generalization without increasing total data volume. These results, further validated within a reinforcement learning framework, suggest that elevating logical density is more critical than expanding data scale for realizing the full cognitive potential of LLMs. The released code is available in the Appendix C.

cs.AI

Exploring the nature of the emergent gauge field in composite-fermion metals: A large-scale microscopic study

Field theories of the composite-fermion (CF) metal model it as a Fermi sea of composite fermions coupled to an emergent gauge field. Within a random phase approximation, these theories predict that the Landau damping of the gauge field resulting from its coupling to the low-energy, long-wavelength CF particle-hole excitations modifies the electrons' density-density correlation function related to the static structure factor $S(q)$ at wave vector $q$. This produces a non-analytic correction $\propto q^{3}\ln q$ to $S(q)$ (with the magnetic length $\ell_{B}=1$). Thanks to the recently developed quaternion formulation for Jain-Kamilla projection of CF wave functions, the evaluation of $S(q)$ from the accurate microscopic theory of composite fermions has now become possible for systems containing as many as $N=900$ CFs, which enables a reliable determination of the small-$q$ behavior of $S(q)$. We study CF metals corresponding to electrons at Landau level filling factors $\nu=1/2$ and $1/4$, and for completeness, also of bosons at $\nu=1$ and $1/3$. In the $q\rightarrow0$ limit, our microscopic calculation reveals a $q^{3}$ term in $S(q)$ of the CF metals rather than $q^{3} \ln q$. This behavior is well-predicted by a model of a non-interacting Fermi sea of dipolar CFs, which also obtains its coefficient accurately.

cond-mat.str-el

Classification of Average Crystalline Topological Superconductors through a Generalized Real-Space Construction

We investigate a novel class of topological superconducting phases protected by exact fermion-parity symmetry and average crystalline symmetries. These phases belong to the broader class of average crystalline symmetry-protected topological (ACSPT) states and include numerous examples of intrinsic ACSPTs -- topological phases that arise only in the presence of disorder or decoherence. Unlike conventional symmetry-protected topological (SPT) phases, which require exact symmetry protection, average SPT (ASPT) phases remain robust as long as the symmetry is restored on average across disorder realizations or mixed-state ensembles. To classify these phases, we extend the real-space block state construction framework to account for average crystalline symmetries. In this generalized setting, lower-dimensional cells are decorated with ASPT phases, and the obstruction-free conditions are reformulated to incorporate the constraints imposed by average symmetry at block intersections. This provides a physically transparent and systematic method for classifying ASPTs with spatial symmetries that are only preserved statistically. We further validate our classification using a generalized spectral sequence analysis, which serves as an independent consistency check. Our results demonstrate that many crystalline topological superconductors remain well defined under realistic imperfections, and they uncover a rich landscape of intrinsically average-symmetry-protected phases that have no analog in clean systems.

cond-mat.str-el