arXiv · 2607.00502
A Task-State Representation for Long-Horizon Mobile GUI Agents
Abstract
While long-horizon mobile GUI agents typically rely on thought-action-observation loops, they struggle to separate persistent task states from transient screen observations. As execution histories grow, this entanglement imposes a severe context burden, causing agents to forget initial requirements, hallucinate progress, or repeatedly interact with stale interfaces. To address this, we introduce Task-State Representation (TSR), a training-free framework that explicitly decouples task state from sensory input. Acting as a lightweight external wrapper, TSR maintains three structured components: a global instruction summary, a dynamic progress tracker for subgoals, and a transition-aware action verifier. By continuously updating through pre- and post-action visual comparisons, TSR effectively guides the agent's reasoning without requiring architectural modifications. Experiments across four mobile GUI benchmarks validate TSR's effectiveness, yielding up to a 12 absolute point increase in success rate on complex cross-application and memory-intensive tasks.
Explore related subjects
Keep this discovery
Yujie Zheng, Zikang Liu, Xin Zhao, Ji-Rong Wen. 2026-07-01. A Task-State Representation for Long-Horizon Mobile GUI Agents. https://arxiv.org/abs/2607.00502
Cite the original work for its findings. Save a collection to share your selection of sources.