arXiv · 2505.12109
SAINT: Attention-Based Policies for Discrete Combinatorial Action Spaces
Abstract
The combinatorial structure of many real-world action spaces leads to exponential growth in the number of possible actions, limiting the effectiveness of conventional reinforcement learning algorithms. Recent approaches for combinatorial action spaces impose factorized or sequential structures over sub-actions, failing to capture complex joint behavior. We introduce the Sub-Action Interaction Network using Transformers (SAINT), a novel policy architecture that represents multi-component actions as unordered sets and models their dependencies via self-attention conditioned on the global state. SAINT is permutation-invariant, sample-efficient, and compatible with standard policy optimization algorithms. In 18 distinct combinatorial environments across three task domains, including environments with $1.35 \times 10^{18}$ possible actions, SAINT consistently outperforms strong baselines.
Explore related subjects
Keep this discovery
Matthew Landers, Taylor W. Killian, Thomas Hartvigsen, Afsaneh Doryab. 2025-05-17. SAINT: Attention-Based Policies for Discrete Combinatorial Action Spaces. https://arxiv.org/abs/2505.12109
Cite the original work for its findings. Save a collection to share your selection of sources.