arXiv · 2609.13011
Comfort by Construction: Adaptive, Comfort-Bounded Action Spaces for Learned Driving Policies
Abstract
Data-driven driving simulators command accelerations and steering rates from a fixed grid without constraining the realized accelerations and jerks. As a result, reinforcement-learning policies inflate safety metrics through abrupt, last-second maneuvers that lie far outside the range of human driving and would be unacceptable to occupants of a real vehicle, so the metrics measure simulator permissiveness rather than policy quality. Enforcing comfort bounds naively is not enough: lateral limits shrink quadratically with speed, so clamping a static grid saturates it and destroys fine-grained control ("grid collapse"). We propose an adaptive action parameterization that rediscretizes the grid at every step to span exactly the per-step feasible control set, via closed-form inversion of the lateral-jerk constraint. We further present PufferDrive-Editor, a browser-based tool to audit realized kinematics and author kinematically challenging scenes. On the Waymo Open Motion Dataset and a hand-authored slalom, our adaptive model holds comfort violations below 1% while outperforming clipped-grid and direct-jerk baselines in navigability.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Anna Rothenhäusler, Daniel Jost, Raghu Rajan, Faris Janjos, Oliver Scheel, Andreas Look, Joschka Boedecker. 2026-09-11. Comfort by Construction: Adaptive, Comfort-Bounded Action Spaces for Learned Driving Policies. https://arxiv.org/abs/2609.13011
Cite the original work for its findings. Save a collection to share your selection of sources.