TY - RPRT TI - Restless Bandits with Individual Penalty Constraints: Near-Optimal Indices and Deep Reinforcement Learning AU - Nida Zamir AU - I-Hong Hou PY - 2026 UR - https://arxiv.org/abs/2604.04101 ID - 2604.04101 ER -