TY - RPRT TI - Regret lower bounds and extended Upper Confidence Bounds policies in stochastic multi-armed bandit problem AU - Antoine Salomon AU - Jean-Yves Audibert AU - Issam El Alaoui PY - 2011 UR - https://arxiv.org/abs/1112.3827 ID - 1112.3827 ER -