Searcharxiv⌕ Search

arXiv subjects

Wei Han

Publications and source records attributed to Wei Han.

At least 55 records · Page 3Linked to original sources

Observation of quantized vortex in an atomic Bose-Einstein condensate at Dirac point with emergent spin-orbit coupling

When two or more energy bands become degenerate at a singular point in the momentum space, such singularity, or ``Dirac points", gives rise to intriguing quantum phenomena as well as unusual material properties. Systems at the Dirac points can possess topological charges and their unique properties can be probed by various methods, such as transport measurement, interferometry and momentum spectroscopy. While the topology of Dirac point in the momentum space is well studied theoretically, observation of topological defects in a many-body quantum systems at Dirac point remain an elusive goal. Based on atomic Bose-Einstein condensate in a graphene-like optical honeycomb lattice, we directly observe emergence of quantized vortices at the Dirac point. The phase diagram of lattice bosons at the Dirac point is revealed. Our work provides a new way of generating vortices in a quantum gas, and the method is generic and can be applied to different types of optical lattices with topological singularity, especially twisted bilayer optical lattices.

cond-mat.quant-gas↗

From Long to Short: LLMs Excel at Trimming Own Reasoning Chains

O1/R1 style large reasoning models (LRMs) signal a substantial leap forward over conventional instruction-following LLMs. By applying test-time scaling to generate extended reasoning paths, they establish many SOTAs across a wide range of complex reasoning tasks. However, recent studies show that LRMs are prone to suffer from overthinking -- the tendency to overcomplicate simple problems, leading to excessive strategy switching and long, convoluted reasoning traces that hinder their interpretability. To mitigate this issue, we conduct a systematic investigation into the reasoning efficiency of a broad set of LRMs and uncover a common dilemma: the difficulty in balancing multiple generation objectives such as correctness and brevity. Based on this discovery, we propose a test-time scaling method, EDIT (Efficient Dynamic Inference Trimming), which efficiently guides LRMs to identify the shortest correct reasoning paths at test time. EDIT employs constraint-guided generation while jointly tracking length and answer distributions under varying constraints, allowing it to select responses that strike an optimal balance between conciseness and correctness. Extensive experiments across diverse models and datasets show that EDIT substantially enhance the reasoning efficiency, producing compact yet informative outputs that improve readability and user experience.

cs.AI↗

Spiking Variational Graph Representation Inference for Video Summarization

With the rise of short video content, efficient video summarization techniques for extracting key information have become crucial. However, existing methods struggle to capture the global temporal dependencies and maintain the semantic coherence of video content. Additionally, these methods are also influenced by noise during multi-channel feature fusion. We propose a Spiking Variational Graph (SpiVG) Network, which enhances information density and reduces computational complexity. First, we design a keyframe extractor based on Spiking Neural Networks (SNN), leveraging the event-driven computation mechanism of SNNs to learn keyframe features autonomously. To enable fine-grained and adaptable reasoning across video frames, we introduce a Dynamic Aggregation Graph Reasoner, which decouples contextual object consistency from semantic perspective coherence. We present a Variational Inference Reconstruction Module to address uncertainty and noise arising during multi-channel feature fusion. In this module, we employ Evidence Lower Bound Optimization (ELBO) to capture the latent structure of multi-channel feature distributions, using posterior distribution regularization to reduce overfitting. Experimental results show that SpiVG surpasses existing methods across multiple datasets such as SumMe, TVSum, VideoXum, and QFVS. Our codes and pre-trained models are available at https://github.com/liwrui/SpiVG.

cs.CV↗

Non-Equilibrium Probing of Topological Supersolids in Spin-Orbit-Coupled Dipolar Condensates

A chiral supersolid is a quantum phase that simultaneously exhibits crystalline order, superfluidity, and topological spin texture, with spontaneously broken translational, U(1) gauge, and chiral symmetries. Here, we demonstrate a chiral supersolid with tunable non-equilibrium dynamics in a spin-orbit coupled dipolar Bose-Einstein condensate. By adjusting dipolar interaction and spin-orbit coupling, we uncover two distinct quantum phase transitions: (i) a first-order transition from a single skyrmion superfluid to a triangular meron supersolid, and (ii) a second-order transition from this superfluid to a square skyrmion supersolid. These phases are characterized by their lattice symmetries, nonclassical rotational inertia, and spin textures. Under parity-time symmetric dissipation, we predict phase-dependent damping of the current oscillations, directly linked to the superfluid fraction. The predicted chiral supersolid phase can be experimentally observed in ultracold magnetic atoms with spin-orbit coupling. Our results establish dipolar quantum gases as a platform for designing topological matter with spintronic functionality.

cond-mat.quant-gas↗

NeuralDB: Scaling Knowledge Editing in LLMs to 100,000 Facts with Neural KV Database

Efficiently editing knowledge stored in large language models (LLMs) enables model updates without large-scale training. One possible solution is Locate-and-Edit (L\&E), allowing simultaneous modifications of a massive number of facts. However, such editing may compromise the general abilities of LLMs and even result in forgetting edited facts when scaling up to thousands of edits. In this paper, we model existing linear L\&E methods as querying a Key-Value (KV) database. From this perspective, we then propose NeuralDB, an editing framework that explicitly represents the edited facts as a neural KV database equipped with a non-linear gated retrieval module, % In particular, our gated module only operates when inference involves the edited facts, effectively preserving the general abilities of LLMs. Comprehensive experiments involving the editing of 10,000 facts were conducted on the ZsRE and CounterFacts datasets, using GPT2-XL, GPT-J (6B) and Llama-3 (8B). The results demonstrate that NeuralDB not only excels in editing efficacy, generalization, specificity, fluency, and consistency, but also preserves overall performance across six representative text understanding and generation tasks. Further experiments indicate that NeuralDB maintains its effectiveness even when scaled to 100,000 facts (\textbf{50x} more than in prior work).

cs.CL↗

Distinguishing dual lattice by strong-pulse matter-wave diffraction

Dual lattices such as honeycomb and hexagonal lattices typically obey Babinet's principle in optics, which states that the expected interference patterns of two complementary diffracting objects are identical and indistinguishable, except for their overall intensity. Here, we study Kapitza--Dirac diffraction of Bose--Einstein condensates in optical lattices and find that matter waves in dual lattices obey Babinet's principle only under the condition of weak-pulse Raman--Nath regimes. In contrast, the Kapitza--Dirac matter-wave diffraction in the strong-pulse Raman--Nath regime (corresponding to the phase wrapping method we developed to generate sub-wavelength phase structures in Sci. Rep. 10, 5870 (2020)) can break Babinet's principle and clearly resolve the distinct interference patterns of the dual honeycomb and hexagonal lattices. This method offers exceptional precision in characterizing lattice configurations and advance the study of symmetry-related phenomena, overcoming the limitations of real-space imaging.

cond-mat.quant-gas↗

Uncertainty-Aware Safety-Critical Decision and Control for Autonomous Vehicles at Unsignalized Intersections

Reinforcement learning (RL) has demonstrated potential in autonomous driving (AD) decision tasks. However, applying RL to urban AD, particularly in intersection scenarios, still faces significant challenges. The lack of safety constraints makes RL vulnerable to risks. Additionally, cognitive limitations and environmental randomness can lead to unreliable decisions in safety-critical scenarios. Therefore, it is essential to quantify confidence in RL decisions to improve safety. This paper proposes an Uncertainty-aware Safety-Critical Decision and Control (USDC) framework, which generates a risk-averse policy by constructing a risk-aware ensemble distributional RL, while estimating uncertainty to quantify the policy's reliability. Subsequently, a high-order control barrier function (HOCBF) is employed as a safety filter to minimize intervention policy while dynamically enhancing constraints based on uncertainty. The ensemble critics evaluate both HOCBF and RL policies, embedding uncertainty to achieve dynamic switching between safe and flexible strategies, thereby balancing safety and efficiency. Simulation tests on unsignalized intersections in multiple tasks indicate that USDC can improve safety while maintaining traffic efficiency compared to baselines.

cs.RO↗

MoNetV2: Enhanced Motion Network for Freehand 3D Ultrasound Reconstruction

Three-dimensional (3D) ultrasound (US) aims to provide sonographers with the spatial relationships of anatomical structures, playing a crucial role in clinical diagnosis. Recently, deep-learning-based freehand 3D US has made significant advancements. It reconstructs volumes by estimating transformations between images without external tracking. However, image-only reconstruction poses difficulties in reducing cumulative drift and further improving reconstruction accuracy, particularly in scenarios involving complex motion trajectories. In this context, we propose an enhanced motion network (MoNetV2) to enhance the accuracy and generalizability of reconstruction under diverse scanning velocities and tactics. First, we propose a sensor-based temporal and multi-branch structure that fuses image and motion information from a velocity perspective to improve image-only reconstruction accuracy. Second, we devise an online multi-level consistency constraint that exploits the inherent consistency of scans to handle various scanning velocities and tactics. This constraint exploits both scan-level velocity consistency, path-level appearance consistency, and patch-level motion consistency to supervise inter-frame transformation estimation. Third, we distill an online multi-modal self-supervised strategy that leverages the correlation between network estimation and motion information to further reduce cumulative errors. Extensive experiments clearly demonstrate that MoNetV2 surpasses existing methods in both reconstruction quality and generalizability performance across three large datasets.

eess.IV↗

Topologically nontrivial and trivial flat bands via weak and strong interlayer coupling in twisted bilayer honeycomb optical lattices for ultracold atoms

In recent years, flat electronic bands in twisted bilayer graphene (TBG) have attracted significant attention due to their intriguing topological properties, extremely slow electron velocities, and enhanced density of states. Extending twisted bilayer systems to new configurations is highly desirable, as it offers promising opportunities to explore flat bands beyond TBG. Here, we study both topological and trivial flat bands in a twisted bilayer honeycomb lattice for ultracold atoms and present the evolution of the flat bands with different interlayer coupling strength (ICS). Our results demonstrate that an isolated topological flat band can emerge at the Dirac point energy for a specific value of weak ICS, referred to as the ``critical coupling". This occurs over a wide range of twist angles, surpassing the limits of the magic angle in TBG systems. When the ICS is slightly increased beyond the critical coupling value, the topological flat band exhibits degenerate band crossings with both the upper and lower adjacent bands at the high-symmetry $Γ_s$ point. As the ICS is further increased into the strong coupling regime, trivial flat bands arise around Dirac point energy. Meanwhile, more trivial flat bands appear, extending from the lowest to higher energy bands, and remain flat as the ICS increases. The topological properties of the flat bands are studied through the winding pattern of the Wilson loop spectrum. Our research provides deeper insights into the formation of flat bands in ultracold atoms with highly controllable twisted bilayer optical lattices, and may contribute to the discovery of new strongly correlated states of matter.

cond-mat.quant-gas↗

PREMISE: Matching-based Prediction for Accurate Review Recommendation

We present PREMISE (PREdict with Matching ScorEs), a new architecture for the matching-based learning in the multimodal fields for the multimodal review helpfulness (MRHP) task. Distinct to previous fusion-based methods which obtains multimodal representations via cross-modal attention for downstream tasks, PREMISE computes the multi-scale and multi-field representations, filters duplicated semantics, and then obtained a set of matching scores as feature vectors for the downstream recommendation task. This new architecture significantly boosts the performance for such multimodal tasks whose context matching content are highly correlated to the targets of that task, compared to the state-of-the-art fusion-based methods. Experimental results on two publicly available datasets show that PREMISE achieves promising performance with less computational cost.

cs.CL↗

APSeg: Auto-Prompt Model with Acquired and Injected Knowledge for Nuclear Instance Segmentation and Classification

Nuclear instance segmentation and classification provide critical quantitative foundations for digital pathology diagnosis. With the advent of the foundational Segment Anything Model (SAM), the accuracy and efficiency of nuclear segmentation have improved significantly. However, SAM imposes a strong reliance on precise prompts, and its class-agnostic design renders its classification results entirely dependent on the provided prompts. Therefore, we focus on generating prompts with more accurate localization and classification and propose \textbf{APSeg}, \textbf{A}uto-\textbf{P}rompt model with acquired and injected knowledge for nuclear instance \textbf{Seg}mentation and classification. APSeg incorporates two knowledge-aware modules: (1) Distribution-Guided Proposal Offset Module (\textbf{DG-POM}), which learns distribution knowledge through density map guided, and (2) Category Knowledge Semantic Injection Module (\textbf{CK-SIM}), which injects morphological knowledge derived from category descriptions. We conducted extensive experiments on the PanNuke and CoNSeP datasets, demonstrating the effectiveness of our approach. The code will be released upon acceptance.

eess.IV↗

Risk-Aware Reinforcement Learning for Autonomous Driving: Improving Safety When Driving through Intersection

Applying reinforcement learning to autonomous driving has garnered widespread attention. However, classical reinforcement learning methods optimize policies by maximizing expected rewards but lack sufficient safety considerations, often putting agents in hazardous situations. This paper proposes a risk-aware reinforcement learning approach for autonomous driving to improve the safety performance when crossing the intersection. Safe critics are constructed to evaluate driving risk and work in conjunction with the reward critic to update the actor. Based on this, a Lagrangian relaxation method and cyclic gradient iteration are combined to project actions into a feasible safe region. Furthermore, a Multi-hop and Multi-layer perception (MLP) mixed Attention Mechanism (MMAM) is incorporated into the actor-critic network, enabling the policy to adapt to dynamic traffic and overcome permutation sensitivity challenges. This allows the policy to focus more effectively on surrounding potential risks while enhancing the identification of passing opportunities. Simulation tests are conducted on different tasks at unsignalized intersections. The results show that the proposed approach effectively reduces collision rates and improves crossing efficiency in comparison to baseline algorithms. Additionally, our ablation experiments demonstrate the benefits of incorporating risk-awareness and MMAM into RL.

cs.RO↗

ER-RAG: Enhance RAG with ER-Based Unified Modeling of Heterogeneous Data Sources

Large language models (LLMs) excel in question-answering (QA) tasks, and retrieval-augmented generation (RAG) enhances their precision by incorporating external evidence from diverse sources like web pages, databases, and knowledge graphs. However, current RAG methods rely on agent-specific strategies for individual data sources, posing challenges low-resource or black-box environments and complicates operations when evidence is fragmented across sources. To address these limitations, we propose ER-RAG, a framework that unifies evidence integration across heterogeneous data sources using the Entity-Relationship (ER) model. ER-RAG standardizes entity retrieval and relationship querying through ER-based APIs with GET and JOIN operations. It employs a two-stage generation process: first, a preference optimization module selects optimal sources; second, another module constructs API chains based on source schemas. This unified approach allows efficient fine-tuning and seamless integration across diverse data sources. ER-RAG demonstrated its effectiveness by winning all three tracks of the 2024 KDDCup CRAG Challenge, achieving performance on par with commercial RAG pipelines using an 8B LLM backbone. It outperformed hybrid competitors by 3.1% in LLM score and accelerated retrieval by 5.5X.

cs.IR↗

Chiral Raman coupling for spin-orbit coupling in ultracold atomic gases

Spin-orbit coupling (SOC) in ultracold atoms is engineered by light-atom interaction, such as two-photon Raman transitions between two Zeeman spin states. In this work, we propose and experimentally realize chiral Raman coupling to generate SOC in ultracold atomic gases, which exhibits high quantization axis direction-dependence. Chiral Raman coupling for SOC is created by chiral light-atom interaction, in which a circularly polarized electromagnetic field generated by two Raman lasers interacts with two Zeeman spin states $δm_{F}=\pm 1$ (chiral transition). We present a simple scheme of chiral one-dimension (1D) Raman coupling by employing two Raman lasers at an intersecting angle 90$^{\circ}$ with the proper polarization configuration. In this case, Raman coupling for SOC exist in one direction of the magnetic quantization axis and disappears in the opposite direction. Then we extend this scheme into a chiral 2D optical square Raman lattice configuration to generate the 1D SOC. There are two orthogonal 1D SOC, which exists in the positive and negative directions of the magnetic quantization axis respectively. This case is compared with 2D SOC based on the nonchiral 2D optical Raman lattice scheme for studying the topological energy band. This work broadens the horizon for understanding chiral physics and simulating topological quantum systems.

cond-mat.quant-gas↗

Efficient Prompt Compression with Evaluator Heads for Long-Context Transformer Inference

Although applications involving long-context inputs are crucial for the effective utilization of large language models (LLMs), they also result in increased computational costs and reduced performance. To address this challenge, we propose an efficient, training-free prompt compression method that retains key information within compressed prompts. We identify specific attention heads in transformer-based LLMs, which we designate as evaluator heads, that are capable of selecting tokens in long inputs that are most significant for inference. Building on this discovery, we develop EHPC, an Evaluator Head-based Prompt Compression method, which enables LLMs to rapidly "skim through" input prompts by leveraging only the first few layers with evaluator heads during the pre-filling stage, subsequently passing only the important tokens to the model for inference. EHPC achieves state-of-the-art results across two mainstream benchmarks: prompt compression and long-context inference acceleration. Consequently, it effectively reduces the complexity and costs associated with commercial API calls. We further demonstrate that EHPC attains competitive results compared to key-value cache-based acceleration methods, thereby highlighting its potential to enhance the efficiency of LLMs for long-context tasks.

cs.CL↗

A Survey on LLM-based Multi-Agent System: Recent Advances and New Frontiers in Application

LLM-based Multi-Agent Systems ( LLM-MAS ) have become a research hotspot since the rise of large language models (LLMs). However, with the continuous influx of new related works, the existing reviews struggle to capture them comprehensively. This paper presents a comprehensive survey of these studies. We first discuss the definition of LLM-MAS, a framework encompassing much of previous work. We provide an overview of the various applications of LLM-MAS in (i) solving complex tasks, (ii) simulating specific scenarios, and (iii) evaluating generative agents. Building on previous studies, we also highlight several challenges and propose future directions for research in this field.

cs.CL↗

Extremely Large Anisotropy of Effective Gilbert Damping in Half-Metallic CrO2

Half-metals are a class of quantum materials with 100% spin-polarization at the Fermi level and have attracted a lot of attention for future spintronic device applications. CrO2 is one of the most promising half-metal candidates, for which the electrical and magnetic properties have been intensively studied in the last several decades. Here, we report the observation of a giant anisotropy (~1600%) of effective Gilbert damping in the single crystalline half metallic (100)-CrO2 thin films, which is significantly larger than the values observed on conventional ferromagnetic Fe and CoFe thin films. Furthermore, the effective Gilbert damping exhibits opposite temperature-dependent behaviors below 50 K with magnetic field along [010] direction and near [001] direction. These experimental results suggest the strong spin-orbit coupling anisotropy of the half-metallic CrO2 and might pave the way for future magnonic computing applications.

cond-mat.mtrl-sci↗

Riemann-based Multi-scale Attention Reasoning Network for Text-3D Retrieval

Due to the challenges in acquiring paired Text-3D data and the inherent irregularity of 3D data structures, combined representation learning of 3D point clouds and text remains unexplored. In this paper, we propose a novel Riemann-based Multi-scale Attention Reasoning Network (RMARN) for text-3D retrieval. Specifically, the extracted text and point cloud features are refined by their respective Adaptive Feature Refiner (AFR). Furthermore, we introduce the innovative Riemann Local Similarity (RLS) module and the Global Pooling Similarity (GPS) module. However, as 3D point cloud data and text data often possess complex geometric structures in high-dimensional space, the proposed RLS employs a novel Riemann Attention Mechanism to reflect the intrinsic geometric relationships of the data. Without explicitly defining the manifold, RMARN learns the manifold parameters to better represent the distances between text-point cloud samples. To address the challenges of lacking paired text-3D data, we have created the large-scale Text-3D Retrieval dataset T3DR-HIT, which comprises over 3,380 pairs of text and point cloud data. T3DR-HIT contains coarse-grained indoor 3D scenes and fine-grained Chinese artifact scenes, consisting of 1,380 and over 2,000 text-3D pairs, respectively. Experiments on our custom datasets demonstrate the superior performance of the proposed method. Our code and proposed datasets are available at \url{https://github.com/liwrui/RMARN}.

cs.CV↗