arXiv · 2605.10585
Controllability in preference-conditioned multi-objective reinforcement learning
Abstract
Multi-objective reinforcement learning (MORL) allows a user to express preference over outcomes in terms of the relative importance of the objectives, but standard metrics cannot capture whether changes in preference reliably change the agent's behavior in the intended way, a property termed controllability. As a result, preference-conditioned agents can score well on standard MORL metrics while being insensitive to the preference input. If the ability to control agents cannot be reliably assessed, the symbolic interface that MORL provides between user intent and agent behavior is broken. Mainstream MORL metrics alone fail to measure the controllability of preference-conditioned agents, motivating a complementary metric specifically designed to that end. We hope the results spur discussion in the community on existing evaluation protocols to consolidate advances in preference adaptation in MORL to larger and more complex problems.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Pau de las Heras Molins, Beyazit Yalcinkaya, Lasse Peters, David Fridovich-Keil, Georgios Bakirtzis. 2026-05-11. Controllability in preference-conditioned multi-objective reinforcement learning. https://arxiv.org/abs/2605.10585
Cite the original work for its findings. Save a collection to share your selection of sources.