arXiv · 2606.29699
Early Warning Signals for OpenVLA Failure under Visual Distribution Shift
Abstract
Visual shifts can cause a vision-language-action policy to fail after initially plausible behavior. We ask whether OpenVLA's internal activations contain signals associated with the steps before failure. We freeze the policy, record one MLP activation per LIBERO-10 step, and fit two linear monitors. Occlusion reduces task success from $57\%$ to $17\%$. Within failed matched-reset trajectories, a layer-16 logistic probe attains AUROC $0.972$ and AUPRC $0.352$, whereas action disagreement attains AUROC $0.496$. Without refitting, the occlusion-trained probe reaches AUROC $0.689$ on failed camera-jitter episodes. In a calibration check, however, the same layer-16 monitor averages 3.32 warning onsets per clean episode. This contrast shows that strong retrospective discrimination does not imply operationally quiet warning behavior. Because fitting and evaluation share tasks, resets, and seed, these results establish retrospective separability rather than prediction on independent episodes.
Explore related subjects
Keep this discovery
Dipesh Tharu Mahato, Rachel Ren. 2026-06-29. Early Warning Signals for OpenVLA Failure under Visual Distribution Shift. https://arxiv.org/abs/2606.29699
Cite the original work for its findings. Save a collection to share your selection of sources.