SearcharxivSearch

arXiv subjects

Qingze Wang

Publications and source records attributed to Qingze Wang.

7 recordsLinked to original sources

When Agents Go Rogue: Activation-Based Detection of Malicious Behaviors in Multi-Agent Systems

While enabling effective collaboration on complex tasks, LLM-based Multi-Agent Systems (MAS) face critical security challenges due to vulnerabilities at the agent and interaction levels. Most existing MAS security defenses are built upon two core assumptions: semantically-explicit malicious attacks and explicit graph-based modeling of the MAS topology and agent-level interactions. In practice, real-world attacks are becoming more semantically stealthy, while MAS execution is typically asynchronous without the temporal alignment assumed by graph-based propagation models. To address these limitations, we propose AcMAS, an activation-based framework for malicious-behavior detection in MAS. By analyzing internal reasoning states in the activation space of local agents, AcMAS detects even stealthy attacks in a synchronization-robust fashion, without relying on explicit interaction graphs. Moreover, our activation analysis provides critical signals to guide AcMAS in restoring the functionality of compromised agents, rather than the disruptive agent isolation commonly used by the state-of-the-art methods. Comprehensive evaluation demonstrates that AcMAS significantly outperforms graph-based baselines against stealthy attacks, by +0.22 F1 in synchronous settings (0.94 vs. 0.72) and by +0.55 F1 in asynchronous settings (0.93 vs. 0.38), with generalization across diverse open-source LLM backbones, attack intensity, and MAS scale.

cs.CR

HoloBrain-0 Technical Report

In this work, we introduce HoloBrain-0, a comprehensive Vision-Language-Action (VLA) framework that bridges the gap between foundation model research and reliable real-world robot deployment. The core of our system is a novel VLA architecture that explicitly incorporates robot embodiment priors, including multi-view camera parameters and kinematic descriptions (URDF), to enhance 3D spatial reasoning and support diverse embodiments. We validate this design through a scalable ``pre-train then post-train" paradigm, achieving state-of-the-art results on simulation benchmarks such as RoboTwin 2.0, LIBERO, and GenieSim, as well as strong results on challenging long-horizon real-world manipulation tasks. Notably, our efficient 0.2B-parameter variant rivals significantly larger baselines, enabling low-latency on-device deployment. To further accelerate research and practical adoption, we fully open-source the entire HoloBrain ecosystem, which includes: (1) powerful pre-trained VLA foundations; (2) post-trained checkpoints for multiple simulation suites and real-world tasks; and (3) RoboOrchard, a full-stack VLA infrastructure for data curation, model training and deployment. Together with standardized data collection protocols, this release provides the community with a complete, reproducible path toward high-performance robotic manipulation.

cs.RO

SWE-IF: Aligning Code Evaluation with Human Preference

Large Language Models (LLMs) have catalyzed vibe coding, where users leverage LLMs to generate and iteratively refine code through natural language interactions until it passes their vibe check. Vibe check reflects human preference and goes beyond functionality: the solution should feel right, read cleanly, preserve intent, and remain correct. However, current code evaluation remains anchored to pass@k and captures only functional correctness, overlooking non-functional instructions that users routinely apply. In this paper, we hypothesize that instruction following is the missing piece underlying vibe check besides functional correctness. To quantify models' code instruction-following capabilities with measurable signals, we present VeriCode, a taxonomy of 30 verifiable code instructions together with deterministic verifiers. We use the taxonomy to augment established evaluation suites, resulting in SWE-IF, a testbed to assess both instruction following and functional correctness. Evaluating 31 LLMs, we show that even the strongest models struggle to comply with multiple instructions and exhibit functional regression. Most importantly, a composite score of functional correctness and instruction following correlates best with human preference, with instruction following emerging as the primary differentiator among LLMs. Our code, data, and taxonomy are available at https://github.com/maszhongming/SWE-IF.

cs.CL

Re-Invoke: Tool Invocation Rewriting for Zero-Shot Tool Retrieval

Recent advances in large language models (LLMs) have enabled autonomous agents with complex reasoning and task-fulfillment capabilities using a wide range of tools. However, effectively identifying the most relevant tools for a given task becomes a key bottleneck as the toolset size grows, hindering reliable tool utilization. To address this, we introduce Re-Invoke, an unsupervised tool retrieval method designed to scale effectively to large toolsets without training. Specifically, we first generate a diverse set of synthetic queries that comprehensively cover different aspects of the query space associated with each tool document during the tool indexing phase. Second, we leverage LLM's query understanding capabilities to extract key tool-related context and underlying intents from user queries during the inference phase. Finally, we employ a novel multi-view similarity ranking strategy based on intents to pinpoint the most relevant tools for each query. Our evaluation demonstrates that Re-Invoke significantly outperforms state-of-the-art alternatives in both single-tool and multi-tool scenarios, all within a fully unsupervised setting. Notably, on the ToolE datasets, we achieve a 20% relative improvement in nDCG@5 for single-tool retrieval and a 39% improvement for multi-tool retrieval.

cs.CL

Spin Texture and Mirror Chern number in Hg-Based Chalcogenides

The unique feature of surface states in topological insulators is the so-called "spin-momentum locking", which means that electron spin is oriented along a fixed direction for a given momentum and forms a texture in the momentum space. In this work, we study spin textures of two typical topological insulators in Hg-Based Chalcogenides, namely HgTe and HgS, based on both the first principles calculation and the eight band Kane model. We find opposite helicities of spin textures between these two materials, originating from the opposite signs of spin-orbit couplings. Furthermore, we reveal that different mirror Chern numbers between HgTe and HgS characterize different topological natures of the systems with opposite spin textures and guarantee the existence of gapless interface states.

cond-mat.mes-hall

Quantum Anomalous Hall Effect in Magnetically Doped InAs/GaSb Quantum Wells

The quantum anomalous Hall effect has recently been observed experimentally in thin films of Cr doped (Bi,Sb)$_2$Te$_3$ at a low temperature ($\sim$ 30mK). In this work, we propose realizing the quantum anomalous Hall effect in more conventional diluted magnetic semiconductors with doped InAs/GaSb type II quantum wells. Based on a four band model, we find an enhancement of the Curie temperature of ferromagnetism due to band edge singularities in the inverted regime of InAs/GaSb quantum wells. Below the Curie temperature, the quantum anomalous Hall effect is confirmed by the direct calculation of Hall conductance. The parameter regime for the quantum anomalous Hall phase is identified based on the eight-band Kane model. The high sample quality and strong exchange coupling make magnetically doped InAs/GaSb quantum wells good candidates for realizing the quantum anomalous Hall insulator at a high temperature.

cond-mat.mes-hall

Hierarchical structure formation in layered superconducting systems with multi-scale inter-vortex interactions

We demonstrate formation of hierarchical structures in two-dimensional systems with multiple length scales in the inter-particle interaction. These include states such as clusters of clusters, concentric rings, clusters inside a ring, and stripes in a cluster. We propose to realize such systems in vortex matter (where a vortex is mapped onto a particle with multi-scale interactions) in layered superconducting systems with varying inter-layer thicknesses and different layer materials.

cond-mat.supr-con