SearcharxivSearch

arXiv subjects

Yumi Omori

Publications and source records attributed to Yumi Omori.

2 recordsLinked to original sources

Minimal Ingredients for Reward Assignment from Expert Demonstrations

Reward assignment from scarce demonstrations is a key challenge in both offline and online imitation learning. A common and intuitive strategy assigns rewards according to how closely learner trajectories match expert demonstrations. Although this principle underlies many existing methods, the core ingredients that drive performance remain systematically underexplored. We therefore ask: what is the minimal structure that reward assignment must encode to achieve effective downstream RL performance across settings? We approach this question along two design axes: proximity approximation and temporal alignment. Across 32 benchmarks spanning offline and online settings, and with three downstream RL algorithms, our empirical findings suggest: (1) In offline regimes, proximity alone captures the reward structure necessary for effective offline RL, while (2) lightweight temporal correspondence provides consistent gains that are modest offline but essential online or in the presence of multiple demonstrations. We further complement our offline results with a lightweight theory characterizing when simple proximity approximation suffices. Overall, these findings advocate algorithmic minimalism in reward design before introducing complex schemes in both offline and online imitation learning.

cs.LG

Should We Ever Prefer Decision Transformer for Offline Reinforcement Learning?

In recent years, extensive work has explored the application of the Transformer architecture to reinforcement learning problems. Among these, Decision Transformer (DT) has gained particular attention in the context of offline reinforcement learning due to its ability to frame return-conditioned policy learning as a sequence modeling task. Most recently, Bhargava et al. (2024) provided a systematic comparison of DT with more conventional MLP-based offline RL algorithms, including Behavior Cloning (BC) and Conservative Q-Learning (CQL), and claimed that DT exhibits superior performance in sparse-reward and low-quality data settings. In this paper, through experimentation on robotic manipulation tasks (Robomimic) and locomotion benchmarks (D4RL), we show that MLP-based Filtered Behavior Cloning (FBC) achieves competitive or superior performance compared to DT in sparse-reward environments. FBC simply filters out low-performing trajectories from the dataset and then performs ordinary behavior cloning on the filtered dataset. FBC is not only very straightforward, but it also requires less training data and is computationally more efficient. The results therefore suggest that DT is not preferable for sparse-reward environments. From prior work, arguably, DT is also not preferable for dense-reward environments. Thus, we pose the question: Is DT ever preferable?

cs.AI