TY - RPRT TI - Preference-Based Reward Learning under Partial Observability with Inexact Dynamics AU - Reza Zolnouri AU - Semih Cayci PY - 2026 UR - https://arxiv.org/abs/2606.30271 ID - 2606.30271 ER -