arXiv · 2601.19499
Progress-Certified Reversible Simplex Supervision of Goal-Reaching Reinforcement Learning
Abstract
Task completion is difficult to certify when state aggregation, model mismatch, and disturbances invalidate nominal RL transitions. We present a reversible Simplex framework combining a frozen finite-state policy with goal-level robust-adaptive recovery and explicit accounting for post-action progress debt. Outward-rounded reachability verifies disturbed transitions, bounds debt from rejected learned actions, and constructs robust recovery and re-entry sets. If an independent checker accepts both certificates and the recovery decrement exceeds the verified debt, every completed switching cycle decreases storage, yielding finite switching and finite-sample goal entry. In a restricted end-to-end instance, the checker resolves all retained obligations, certifies $\kappa_L=0.482$, and establishes $\varepsilon_{\mathrm{sw}}\ge 0.060$. Matched numerical episodes yield 20/20 goal entries and 0/320 debt-test violations with debt gating, versus 18/20 and 742/796 without it. Twenty-four hardware trials evaluate 50 ms supervision above a 1 kHz actuator stack; formal certification remains limited to the checked model.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Mehdi Heydari Shahna, Jouni Mattila. 2026-01-27. Progress-Certified Reversible Simplex Supervision of Goal-Reaching Reinforcement Learning. https://arxiv.org/abs/2601.19499
Cite the original work for its findings. Save a collection to share your selection of sources.