SearcharxivSearch

arXiv subjects

Wanxia Cao

Publications and source records attributed to Wanxia Cao.

6 recordsLinked to original sources

GSAR: Goal-State-Anchor Rewards for Mobile GUI Agents with Self-Evolving Data Synthesis

Vision-Language Models (VLMs) based GUI agents stand to benefit significantly from online reinforcement learning (RL). However, their training is bottlenecked by two fundamental issues: current data synthesis methods for GUI Agents rely on specific environments and struggle to generate diverse data, while existing evaluators either suffer from limited scalability or provide inaccurate and unreliable reward signals. To overcome these challenges, we introduce GSAR (Goal-State-Anchor Reward), a RL reward framework that supports scalable task generation and delivers reliable reward signals for stable and efficient policy optimization. Our approach features self-evolving data synthesis, which produces multiple environments through task execution and generates diverse tasks and goal states. Complementing this, a state-anchor mechanism automatically annotates task-relevant UI elements in successful goal states as reference anchors. During RL training, these reference anchors provide accurate, scalable reward signals that substantially enhance efficiency. Extensive evaluations demonstrate that our framework achieves over 90% accuracy on offline trajectory verification and performs closest to rule-based methods. Furthermore, agents trained using our reward framework exhibit strong performance on both AndroidWorld and our constructed benchmark, establishing a scalable approach for GUI agent training.

cs.AI

Xiaomi-GUI-0 Technical Report

Graphical user interface (GUI) agents build on vision-language models to complete user tasks end-to-end in real applications through interface actions such as tapping, swiping, text entry, and navigation. However, existing GUI agents are trained and evaluated largely on offline trajectories, simulated environments, and standardized benchmarks. These differ substantially from real applications in interface layout, interaction logic, and abnormal-state distribution, and cannot faithfully characterize execution stability in real-world use, where account states, permission dialogs, payment authentication, and risk control continually reshape the state distribution and open a persistent gap between benchmark scores and real usability. To close this gap, we propose Xiaomi-GUI-0, a native multimodal GUI agent for real mobile environments, trained and evaluated within a real-device closed loop. At its core is a real-device-dominant hybrid infrastructure, where physical devices are the primary execution environment and sandboxes provide auxiliary support, so that data collection, training, rollout, and evaluation share an execution distribution close to real deployment. We construct multi-source training data spanning high-frequency head tasks, high-generalization data for long-tail intents, and capability-enhancement data for reflection and memory, and introduce an error-driven data flywheel that turns failure trajectories into corrected actions, reflective explanations, and recovery demonstrations. The model is trained through a progressive three-stage pipeline of supervised fine-tuning, step-level reinforcement learning, and agentic reinforcement learning. Evaluated on public benchmarks and our in-house RealMobile, Xiaomi-GUI-0 achieves 72.0% success on RealMobile and 78.9% on AndroidWorld, while substantially improving execution stability and abnormal-state recognition in real-world tasks.

cs.AI

Attention Basin: Why Contextual Position Matters in Large Language Models

The performance of Large Language Models (LLMs) is significantly sensitive to the contextual position of information in the input. To investigate the mechanism behind this positional bias, our extensive experiments reveal a consistent phenomenon we term the attention basin: when presented with a sequence of structured items (e.g., retrieved documents or few-shot examples), models systematically assign higher attention to the items at the beginning and end of the sequence, while neglecting those in the middle. Crucially, our analysis further reveals that allocating higher attention to critical information is key to enhancing model performance. Based on these insights, we introduce Attention-Driven Reranking (AttnRank), a two-stage framework that (i) estimates a model's intrinsic positional attention preferences using a small calibration set, and (ii) reorders retrieved documents or few-shot examples to align the most salient content with these high-attention positions. AttnRank is a model-agnostic, training-free, and plug-and-play method with minimal computational overhead. Experiments on multi-hop QA and few-shot in-context learning tasks demonstrate that AttnRank achieves substantial improvements across 10 large language models of varying architectures and scales, without modifying model parameters or training procedures.

cs.CL

Non-reciprocity and quantum correlations of light transport in hot atoms via reservoir engineering

The breaking of reciprocity is a topic of great interest in fundamental physics and optical information processing applications. We demonstrate non-reciprocal light transport in a quantum system of hot atoms by engineering the dissipative atomic reservoir. Our scheme is based on the phase-sensitive light transport in a multi-channel photon-atom interaction configuration, where the phase of collective atomic excitations is tunable through external driving fields. Remarkably, we observe inter-channel quantum correlations which originate from interactions with the judiciously engineered reservoir. The non-reciprocal transport in a quantum optical atomic system constitutes a new paradigm for atom-based, non-reciprocal optics, and offers opportunities for quantum simulations with coupled optical channels.

quant-ph

Quantum correlations near the exceptional point

Recent advances in non-Hermitian physical systems have led to numerous novel optical phenomena and applications. However, most realizations are limited to classical systems and quantum fluctuations of light is unexplored. For the first time, we report the observation of quantum correlations between light channels in an anti-symmetric optical system made of flying atoms. Two distant optical channels coupled dissipatively, display gain, phase sensitivity and quantum correlations with each other, even under linear atom-light interaction within each channel. We found that quantum correlations emerge in the phase unbroken regime and disappears after crossing the exceptional point. Our microscopic model considering quantum noise evolution produces results in good qualitative agreement with experimental observations. This work opens up a new direction of experimental quantum nonlinear optics using non-Hermitian systems, and demonstrates the viability of nonlinear coupling with linear systems by using atomic motion as feedback.

quant-ph

Anti-Parity-Time Symmetric Optics via Flying Atoms

The recently-developed notion of 'parity-time (PT) symmetry' in optical systems with a controlled gain-loss interplay has spawned an intriguing way of achieving optical behaviors that are presently unattainable with standard arrangements. In most experimental studies so far, however, the implementations rely highly on the advances of nanotechnologies and sophisticated fabrication techniques to synthesize solid-state materials. Here, we report the first experimental demonstration of optical anti-PT symmetry, a counterpart of conventional PT symmetry, in a warm atomic-vapor cell. By exploiting rapid coherence transport via flying atoms, our scheme illustrates essential features of anti-PT symmetry with an unprecedented precision on phase-transition threshold, and substantially reduces experimental complexity and cost. This result represents a significant advance in non-Hermitian optics by bridging a firm connection with the field of atomic, molecular and optical physics, where novel phenomena and applications in quantum and nonlinear optics aided by (anti-)PT symmetry could be anticipated.

physics.optics