SearcharxivSearch

arXiv subjects

Bowen Yan

Publications and source records attributed to Bowen Yan.

At least 19 recordsLinked to original sources

Encoding Circuit Satisfiability in Rydberg Atom Arrays

Rydberg atom arrays natively encode the maximum-weight independent set (MWIS) problem through the blockade mechanism, so the Boolean circuit satisfiability problem (Circuit-SAT) can be brought onto the platform once it is reduced to MWIS. The conventional encoding of Circuit-SAT in the Rydberg atom array proceeds through conjunctive normal form (CNF) and incurs a substantial atom overhead. We introduce CAMERA (Circuit-SAT Atom-efficient MWIS Encoding for Rydberg Arrays), a method that provides MWIS encodings of Circuit-SAT instances on the king subgraph geometry of the array. CAMERA represents each logic gate as a compact weighted gadget and assembles the gadgets with a placement and routing compiler inspired by very large scale integration (VLSI) design. On random multi-gate benchmarks, the direct encoding route lowers the atom cost relative to the CNF route by an average factor of $22.4 \pm 1.8$. To demonstrate that the encoding extends from individual weighted gadgets to multi-gate arithmetic blocks, we compile a full adder and a multiplier, verifying each against its complete truth table by exact classical ground state calculations. We further showcase solving a representative Circuit-SAT instance end-to-end, from gate level compilation through a closed-system tensor-network simulation of a hardware-compatible annealing protocol on the encoded 30-atom instance to readout of a satisfying assignment. These results establish a complete encoding and simulation workflow as a proof of principle, and a concrete route toward solving a broader family of combinatorial problems on Rydberg atom arrays.

quant-ph

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation

As the foundational architecture of modern machine learning, Transformers have driven remarkable progress across diverse AI domains. Despite their transformative impact, a persistent challenge across various Transformers is Attention Sink (AS), in which a disproportionate amount of attention is focused on a small subset of specific yet uninformative tokens. AS complicates interpretability, significantly affecting the training and inference dynamics, and exacerbates issues such as hallucinations. In recent years, substantial research has been dedicated to understanding and harnessing AS. However, a comprehensive survey that systematically consolidates AS-related research and offers guidance for future advancements remains lacking. To address this gap, we present the first survey on AS, structured around three key dimensions that define the current research landscape: Fundamental Utilization, Mechanistic Interpretation, and Strategic Mitigation. Our work makes a pivotal contribution by highlighting the key concepts and main trends in the field, guiding researchers through the evolution of AS-related studies. We envision this survey as a valuable resource, empowering researchers to effectively manage AS within the current Transformer paradigm, while simultaneously inspiring innovative advancements for the next generation of Transformers. The paper list of this work is available at https://github.com/ZunhaiSu/Awesome-Attention-Sink.

cs.LG

Data Assessment for Embodied Intelligence

In embodied intelligence, datasets play a pivotal role, serving as both a knowledge repository and a conduit for information transfer. The two most critical attributes of a dataset are the amount of information it provides and how easily this information can be learned by models. However, the multimodal nature of embodied data makes evaluating these properties particularly challenging. Prior work has largely focused on diversity, typically counting tasks and scenes or evaluating isolated modalities, which fails to provide a comprehensive picture of dataset diversity. On the other hand, the learnability of datasets has received little attention and is usually assessed post-hoc through model training, an expensive, time-consuming process that also lacks interpretability, offering little guidance on how to improve a dataset. In this work, we address both challenges by introducing two principled, data-driven tools. First, we construct a unified multimodal representation for each data sample and, based on it, propose diversity entropy, a continuous measure that characterizes the amount of information contained in a dataset. Second, we introduce the first interpretable, data-driven algorithm to efficiently quantify dataset learnability without training, enabling researchers to assess a dataset's learnability immediately upon its release. We validate our algorithm on both simulated and real-world embodied datasets, demonstrating that it yields faithful, actionable insights that enable researchers to jointly improve diversity and learnability. We hope this work provides a foundation for designing higher-quality datasets that advance the development of embodied intelligence.

cs.RO

Efficient Reasoning for LLMs through Speculative Chain-of-Thought

Large reasoning language models such as OpenAI-o1 and Deepseek-R1 have recently attracted widespread attention due to their impressive task-solving abilities. However, the enormous model size and the generation of lengthy thought chains introduce significant reasoning costs and response latency. Existing methods for efficient reasoning mainly focus on reducing the number of model parameters or shortening the chain-of-thought length. In this paper, we introduce Speculative Chain-of-Thought (SCoT), which reduces reasoning latency from another perspective by accelerated average reasoning speed through large and small model collaboration. SCoT conducts thought-level drafting using a lightweight draft model. Then it selects the best CoT draft and corrects the error cases with the target model. The proposed thinking behavior alignment improves the efficiency of drafting and the draft selection strategy maintains the prediction accuracy of the target model for complex tasks. Experimental results on GSM8K, MATH, GaoKao, CollegeMath and Olympiad datasets show that SCoT reduces reasoning latency by 48\%$\sim$66\% and 21\%$\sim$49\% for Deepseek-R1-Distill-Qwen-32B and Deepseek-R1-Distill-Llama-70B while achieving near-target-model-level performance. Our code is available at https://github.com/Jikai0Wang/Speculative_CoT.

cs.CL

Gradient Co-occurrence Analysis for Detecting Unsafe Prompts in Large Language Models

Unsafe prompts pose significant safety risks to large language models (LLMs). Existing methods for detecting unsafe prompts rely on data-driven fine-tuning to train guardrail models, necessitating significant data and computational resources. In contrast, recent few-shot gradient-based methods emerge, requiring only few safe and unsafe reference prompts. A gradient-based approach identifies unsafe prompts by analyzing consistent patterns of the gradients of safety-critical parameters in LLMs. Although effective, its restriction to directional similarity (cosine similarity) introduces ``directional bias'', limiting its capability to identify unsafe prompts. To overcome this limitation, we introduce GradCoo, a novel gradient co-occurrence analysis method that expands the scope of safety-critical parameter identification to include unsigned gradient similarity, thereby reducing the impact of ``directional bias'' and enhancing the accuracy of unsafe prompt detection. Comprehensive experiments on the widely-used benchmark datasets ToxicChat and XStest demonstrate that our proposed method can achieve state-of-the-art (SOTA) performance compared to existing methods. Moreover, we confirm the generalizability of GradCoo in detecting unsafe prompts across a range of LLM base models with various sizes and origins.

cs.CL

Floquet Codes from Coupled Spin Chains

We propose a novel construction of the Floquet 3D toric code and Floquet $X$-cube code through the coupling of spin chains. This approach not only recovers the coupling layer construction on foliated lattices in three dimensions but also avoids the complexity of coupling layers in higher dimensions, offering a more localized and easily generalizable framework. Our method extends the Floquet 3D toric code to a broader class of lattices, aligning with its topological phase properties. Furthermore, we generalize the Floquet $X$-cube model to arbitrary manifolds, provided the lattice is locally cubic, consistent with its Fractonic phases. We also introduce a unified error-correction paradigm for Floquet codes by defining a subgroup, the Steady Stabilizer Group (SSG), of the Instantaneous Stabilizer Group (ISG), emphasizing that not all terms in the ISG contribute to error correction, but only those terms that can be referred to at least twice before being removed from the ISG. We show that correctable Floquet codes naturally require the SSG to form a classical error-correcting code, and we present a simple 2-step Bacon-Shor Floquet code as an example, where SSG forms instantaneous repetition codes. Finally, our construction intrinsically supports the extension to $n$-dimensional Floquet $(n,1)$ toric codes and generalized $n$-dimensional Floquet $X$-cube codes.

quant-ph

FIHA: Autonomous Hallucination Evaluation in Vision-Language Models with Davidson Scene Graphs

The rapid development of Large Vision-Language Models (LVLMs) often comes with widespread hallucination issues, making cost-effective and comprehensive assessments increasingly vital. Current approaches mainly rely on costly annotations and are not comprehensive -- in terms of evaluating all aspects such as relations, attributes, and dependencies between aspects. Therefore, we introduce the FIHA (autonomous Fine-graIned Hallucination evAluation evaluation in LVLMs), which could access hallucination LVLMs in the LLM-free and annotation-free way and model the dependency between different types of hallucinations. FIHA can generate Q&A pairs on any image dataset at minimal cost, enabling hallucination assessment from both image and caption. Based on this approach, we introduce a benchmark called FIHA-v1, which consists of diverse questions on various images from MSCOCO and Foggy. Furthermore, we use the Davidson Scene Graph (DSG) to organize the structure among Q&A pairs, in which we can increase the reliability of the evaluation. We evaluate representative models using FIHA-v1, highlighting their limitations and challenges. We released our code and data.

cs.CV

Numerical simulations of attachment-line boundary layer in hypersonic flow, Part I: roughness-induced subcritical transitions

The attachment-line boundary layer is critical in hypersonic flows because of its significant impact on heat transfer and aerodynamic performance. In this study, high-fidelity numerical simulations are conducted to analyze the subcritical roughness-induced laminar-turbulent transition at the leading-edge attachment-line boundary layer of a blunt swept body under hypersonic conditions. This simulation represents a significant advancement by successfully reproducing the complete leading-edge contamination process induced by surface roughness elements in a realistic configuration, thereby providing previously unattainable insights. Two roughness elements of different heights are examined. For the lower-height roughness element, additional unsteady perturbations are required to trigger a transition in the wake, suggesting that the flow field around the roughness element acts as a disturbance amplifier for upstream perturbations. Conversely, a higher roughness element can independently induce the transition. A low-frequency absolute instability is detected behind the roughness, leading to the formation of streaks. The secondary instabilities of these streaks are identified as the direct cause of the final transition.

physics.flu-dyn

Numerical simulations of attachment-line boundary layer in hypersonic flow, Part II: the features of three-dimensional turbulent boundary layer

In this study,we investigate the characteristics of three-dimensional turbulent boundary layers influenced by transverse flow and pressure gradients. Our findings reveal that even without assuming an infinite sweep, a fully developed turbulent boundary layer over the present swept blunt body maintains spanwise homogeneity, consistent with infinite sweep assumptions.We critically examine the law-of-the and temperature-velocity relationships, typically applied two-dimensional turbulent boundary layers, in three-dimensional contexts. Results show that with transverse velocity and pressure gradient, streamwise velocity adheres to classical velocity transformation relationships and the predictive accuracy of classical temperaturevelocity relationships diminishes because of pressure gradient. We show that near-wall streak structures persist and correspond with energetic structures in the outer region, though three-dimensional effects redistribute energy to align more with the external flow direction. Analysis of shear Reynolds stress and mean flow shear directions reveals in near-wall regions with low transverse flow velocity, but significant deviations at higher transverse velocities. Introduction of transverse pressure gradients together with the transverse velocities alter the velocity profile and mean flow shear directions, with shear Reynolds stress experiencing similar changes but with a lag increasing with transverse. Consistent directional alignment in outer regions suggests a partitioned relationship between shear Reynolds stress and mean flow shear: nonlinear in the inner region and approximately linear in the outer region.

physics.flu-dyn

Representing arbitrary ground states of toric code by a restricted Boltzmann machine

We systematically analyze the representability of toric code ground states by Restricted Boltzmann Machine with only local connections between hidden and visible neurons. This analysis is pivotal for evaluating the model's capability to represent diverse ground states, thus enhancing our understanding of its strengths and weaknesses. Subsequently, we modify the Restricted Boltzmann Machine to accommodate arbitrary ground states by introducing essential non-local connections efficiently. The new model is not only analytically solvable but also demonstrates efficient and accurate performance when solved using machine learning techniques. Then we generalize our the model from $Z_2$ to $Z_n$ toric code and discuss future directions.

cond-mat.dis-nn

Demonstration Augmentation for Zero-shot In-context Learning

Large Language Models (LLMs) have demonstrated an impressive capability known as In-context Learning (ICL), which enables them to acquire knowledge from textual demonstrations without the need for parameter updates. However, many studies have highlighted that the model's performance is sensitive to the choice of demonstrations, presenting a significant challenge for practical applications where we lack prior knowledge of user queries. Consequently, we need to construct an extensive demonstration pool and incorporate external databases to assist the model, leading to considerable time and financial costs. In light of this, some recent research has shifted focus towards zero-shot ICL, aiming to reduce the model's reliance on external information by leveraging their inherent generative capabilities. Despite the effectiveness of these approaches, the content generated by the model may be unreliable, and the generation process is time-consuming. To address these issues, we propose Demonstration Augmentation for In-context Learning (DAIL), which employs the model's previously predicted historical samples as demonstrations for subsequent ones. DAIL brings no additional inference cost and does not rely on the model's generative capabilities. Our experiments reveal that DAIL can significantly improve the model's performance over direct zero-shot inference and can even outperform few-shot ICL without any external information.

cs.CL

Rethinking Negative Instances for Generative Named Entity Recognition

Large Language Models (LLMs) have demonstrated impressive capabilities for generalizing in unseen tasks. In the Named Entity Recognition (NER) task, recent advancements have seen the remarkable improvement of LLMs in a broad range of entity domains via instruction tuning, by adopting entity-centric schema. In this work, we explore the potential enhancement of the existing methods by incorporating negative instances into training. Our experiments reveal that negative instances contribute to remarkable improvements by (1) introducing contextual information, and (2) clearly delineating label boundaries. Furthermore, we introduce an efficient longest common subsequence (LCS) matching algorithm, which is tailored to transform unstructured predictions into structured entities. By integrating these components, we present GNER, a Generative NER system that shows improved zero-shot performance across unseen entity domains. Our comprehensive evaluation illustrates our system's superiority, surpassing state-of-the-art (SoTA) methods by 9 $F_1$ score in zero-shot evaluation.

cs.CL

CMD: a framework for Context-aware Model self-Detoxification

Text detoxification aims to minimize the risk of language models producing toxic content. Existing detoxification methods of directly constraining the model output or further training the model on the non-toxic corpus fail to achieve a decent balance between detoxification effectiveness and generation quality. This issue stems from the neglect of constrain imposed by the context since language models are designed to generate output that closely matches the context while detoxification methods endeavor to ensure the safety of the output even if it semantically deviates from the context. In view of this, we introduce a Context-aware Model self-Detoxification~(CMD) framework that pays attention to both the context and the detoxification process, i.e., first detoxifying the context and then making the language model generate along the safe context. Specifically, CMD framework involves two phases: utilizing language models to synthesize data and applying these data for training. We also introduce a toxic contrastive loss that encourages the model generation away from the negative toxic samples. Experiments on various LLMs have verified the effectiveness of our MSD framework, which can yield the best performance compared to baselines.

cs.CL

Generalized Kitaev Spin Liquid model and Emergent Twist Defect

The Kitaev spin liquid model on honeycomb lattice offers an intriguing feature that encapsulates both Abelian and non-Abelian anyons. Recent studies suggest that the comprehensive phase diagram of possible generalized Kitaev model largely depends on the specific details of the discrete lattice, which somewhat deviates from the traditional understanding of "topological" phases. In this paper, we propose an adapted version of the Kitaev spin liquid model on arbitrary planar lattices. Our revised model recovers the toric code model under certain parameter selections within the Hamiltonian terms. Our research indicates that changes in parameters can initiate the emergence of holes, domain walls, or twist defects. Notably, the twist defect, which presents as a lattice dislocation defect, exhibits non-Abelian braiding statistics upon tuning the coefficients of the Hamiltonian on a standard translationally invariant lattice. Additionally, we illustrate that the creation, movement, and fusion of these defects can be accomplished through natural time evolution by linearly interpolating the static Hamiltonian. These defects demonstrate the Ising anyon fusion rule as anticipated. Our findings hint at possible implementation in actual physical materials owing to a more realistically achievable two-body interaction.

cond-mat.str-el

Ribbon operators in the generalized Kitaev quantum double model based on Hopf algebras

Kitaev's quantum double model is a family of exactly solvable lattice models that realize two dimensional topological phases of matter. Originally it is based on finite groups, and is later generalized to semi-simple Hopf algebras. We rigorously define and study ribbon operators in the generalized Kitaev quantum double model. These ribbon operators are important tools to understand quasi-particle excitations. It turns out that there are some subtleties in defining the operators in contrast to what one would naively think. In particular, one has to distinguish two classes of ribbons which we call locally clockwise and locally counterclockwise ribbons. Moreover, this issue already exists in the original model based on finite non-Abelian groups. We show how certain properties would fail even in the original model if we do not distinguish these two classes of ribbons. Perhaps not surprisingly, under the new definitions ribbon operators satisfy all properties that are expected. For instance, they create quasi-particle excitations only at the end of the ribbon, and the types of the quasi-particles correspond to irreducible representations of the Drinfeld double of the input Hopf algebra. However, the proofs of these properties are much more complicated than those in the case of finite groups. This is partly due to the complications in dealing with general Hopf algebras rather than just group algebras.

cond-mat.str-el

Quantum circuits for toric code and X-cube fracton model

We propose a systematic and efficient quantum circuit composed solely of Clifford gates for simulating the ground state of the surface code model. This approach yields the ground state of the toric code in $\lceil 2L+2+log_{2}(d)+\frac{L}{2d} \rceil$ time steps, where $L$ refers to the system size and $d$ represents the maximum distance to constrain the application of the CNOT gates. Our algorithm reformulates the problem into a purely geometric one, facilitating its extension to attain the ground state of certain 3D topological phases, such as the 3D toric model in $3L+8$ steps and the X-cube fracton model in $12L+11$ steps. Furthermore, we introduce a gluing method involving measurements, enabling our technique to attain the ground state of the 2D toric code on an arbitrary planar lattice and paving the way to more intricate 3D topological phases.

cond-mat.str-el

The Superior Knowledge Proximity Measure for Patent Mapping

Network maps of patent classes have been widely used to analyze the coherence and diversification of technology or knowledge positions of inventors, firms, industries, regions, and so on. To create such networks, a measure is required to associate different classes of patents in the patent database and often indicates knowledge proximity (or distance). Prior studies have used a variety of knowledge proximity measures based on different perspectives and association rules. It is unclear how to consistently assess and compare them, and which ones are superior for constructing a generally useful total patent class network. Such uncertainty has limited the generality and applications of the previously reported maps. Herein, we use a statistical method to identify the superior proximity measure from a comprehensive set of typical measures, by evaluating and comparing their explanatory powers on the historical expansions of the patent portfolios of individual inventors and organizations across different patent classes. Based on the complete United States granted patent database from 1976 to 2017, our analysis identifies a reference-based Jaccard index as the statistically superior measure, for explaining the historical diversifications and predicting future movement directions of both individual inventors and organizations across technology domains.

cs.DL

Multicores-periphery structure in networks

Many real-world networks exhibit a multicores-periphery structure, with densely connected vertices in multiple cores surrounded by a general periphery of sparsely connected vertices. Identification of the multicores-periphery structure can provide a new lens to understand the structures and functions of various real-world networks. This paper defines the multicores-periphery structure and introduces an algorithm to identify the optimal partition of multiple cores and the periphery in general networks. We demonstrate the performance of our algorithm by applying it to a well-known social network and a patent technology network, which are best characterized by the multicores-periphery structure. The analyses also reveal the differences between our multicores-periphery detection algorithm and two state-of-the-art algorithms for detecting the single core-periphery structure and community structure.

cs.SI