SearcharxivSearch

arXiv subjects

Ziad Kobeissi

Publications and source records attributed to Ziad Kobeissi.

6 recordsLinked to original sources

Fast and Robust Convergence Rate for TD(0) with Linear Function Approximation, Universal Learning Steps and I.I.D. Samples

In this paper, we study the finite-time behavior of the TD(0) temporal-difference method with linear function approximation (LFA). We consider on-policy independent and identically distributed (i.i.d.) samples, a constant learning step, and the Polyak-Juditsky averaging method. We establish a new convergence rate, for the Mean-Square Error (MSE) on the approximated function, that is (i) fast in the sense that it admits an optimal dependency in the number of iterations k (i.e., of order 1/k), (ii) robust to ill-conditioning: it only depends on an initial error and modelindependent constants and (iii) sharp up to a multiplicative constant lower than 11. In particular, it does not depend on the smallest eigenvalue of the uncentered covariance matrix of the linear parametrization, unlike all pre-existing O(1/k) rates in the TD(0) literature. We also introduce PCTD(0), a variant of TD(0), which benefits from better convergence properties under an additional assumption of strong mixing on the Markov Chain.

stat.ML

Mean-field games for harvesting problems: Uniqueness, long-time behaviour and weak KAM theory

The goal of this paper is to study a Mean Field Game (MFG) system stemming from the harvesting of resources. Modelling the latter through a reaction-diffusion equation and the harvesters as competing rational agents, we are led to a non-local (in time and space) MFG system that consists of three equations, the study of which is quite delicate. The main focus of this paper is on the derivation of analytical results (e.g existence, uniqueness) and of long time behaviour (here, convergence to the ergodic system). We provide some explicit solutions to this ergodic system.

math.AP

Temporal Difference Learning with Continuous Time and State in the Stochastic Setting

We consider the problem of continuous-time policy evaluation. This consists in learning through observations the value function associated with an uncontrolled continuous-time stochastic dynamic and a reward function. We propose two original variants of the well-known TD(0) method using vanishing time steps. One is model-free and the other is model-based. For both methods, we prove theoretical convergence rates that we subsequently verify through numerical simulations. Alternatively, those methods can be interpreted as novel reinforcement learning approaches for approximating solutions of linear PDEs (partial differential equations) or linear BSDEs (backward stochastic differential equations).

cs.LG

The tragedy of the commons: A Mean-Field Game approach to the reversal of travelling waves

The goal of this paper is to investigate an instance of the tragedy of the commons in spatially distributed harvesting games. The model we choose is that of a fishes' population that is governed by a parabolic bistable equation and that fishermen harvest. We assume that, when no fisherman is present, the fishes' population is invading (mathematically, there is an invading travelling front). Is it possible that fishermen, when acting selfishly, each in his or her own best interest, might lead to a reversal of the travelling wave and, consequently, to an extinction of the global population? To answer this question, we model the behaviour of individual fishermen using a Mean Field Game approach, and we show that the answer is yes. We then show that, at least in some cases, if the fishermen coordinated instead of acting selfishly, each of them could make more benefit, while still guaranteeing the survival of the population. Our study is illustrated by several numerical simulations.

math.AP

A Non-asymptotic Analysis of Non-parametric Temporal-Difference Learning

Temporal-difference learning is a popular algorithm for policy evaluation. In this paper, we study the convergence of the regularized non-parametric TD(0) algorithm, in both the independent and Markovian observation settings. In particular, when TD is performed in a universal reproducing kernel Hilbert space (RKHS), we prove convergence of the averaged iterates to the optimal value function, even when it does not belong to the RKHS. We provide explicit convergence rates that depend on a source condition relating the regularity of the optimal value function to the RKHS. We illustrate this convergence numerically on a simple continuous-state Markov reward process.

math.OC

On the implementation of a primal-dual algorithm for second order time-dependent mean field games with local couplings

We study a numerical approximation of a time-dependent Mean Field Game (MFG) system with local couplings. The discretization we consider stems from a variational approach described in [Briceno-Arias, Kalise, and Silva, SIAM J. Control Optim., 2017] for the stationary problem and leads to the finite difference scheme introduced by Achdou and Capuzzo-Dolcetta in [SIAM J. Numer. Anal., 48(3):1136-1162, 2010]. In order to solve the finite dimensional variational problems, in [Briceno-Arias, Kalise, and Silva, SIAM J. Control Optim., 2017] the authors implement the primal-dual algorithm introduced by Chambolle and Pock in [J. Math. Imaging Vision, 40(1):120-145, 2011], whose core consists in iteratively solving linear systems and applying a proximity operator. We apply that method to time-dependent MFG and, for large viscosity parameters, we improve the linear system solution by replacing the direct approach used in [Briceno-Arias, Kalise, and Silva, SIAM J. Control Optim., 2017] by suitable preconditioned iterative algorithms.

math.OC