PACE: Policy-Native Adaptive Decision Timing for Long-Horizon Reasoning
Long-horizon reasoning requires deciding not only what actions to take, but how many to execute open-loop before replanning. This number, the execution depth, balances replanning cost against compounding execution errors. Most current systems either fix the execution depth as a hand-tuned scalar or adjust it at inference time with heuristic rules decoupled from the policy; we argue both can be suboptimal. In this work, we treat the execution depth as a learnable, history-conditioned variable of the policy itself, and propose PACE, a model-native VLM policy that jointly predicts what to execute and for how long under a hard decision budget. We evaluate PACE on four long-horizon environments spanning fully and partially observable settings and visual and textual modalities: Sliding Puzzle, Sokoban, ALFWorld, and ScienceWorld. Across all settings, PACE improves success rate by 3.1 to 15.6 percentage points, while using fewer decisions on average, achieving empirical success-decision Pareto dominance.