TY - RPRT TI - A convex programming approach for discrete-time Markov decision processes under the expected total reward criterion AU - F. Dufour AU - Alexandre Genadot PY - 2019 UR - https://arxiv.org/abs/1903.08853 ID - 1903.08853 ER -