SearcharxivSearch

arXiv subjects

Omar Mrani-Zentar

Publications and source records attributed to Omar Mrani-Zentar.

4 recordsLinked to original sources

Kernel Metrics and Learning for Borel MDPs: Identifiability and Adaptive Control

We consider a Markov decision process with standard Borel spaces and an unknown transition kernel under the average cost criterion. We do not impose any parametrization on the set of possible kernels. To facilitate our analysis, we first develop implication relations between several topologies on kernels defined by pointwise, continuous, or uniform weak convergence; we then review robustness properties on the space of kernels, and finally we establish compactness conditions on the space of kernels. Building on this regularity analysis, we then present two data-driven identifiability results; the first one being Bayesian and the second one empirical. Our conditions for the Bayesian setting are significantly more relaxed compared with prior work which considered either finite or parametric models, though we do not obtain a rate of convergence. Our analysis is asymptotic and builds on measurability in terms of the tail $σ$-field of the available information. Identifiability results are then used to design near-optimal adaptive control policies which alternate between periods of exploration, where the controller acts according to a policy which is conducive to the identification of the true kernel, and periods of exploitation where the controller's information on the kernel is utilized. We will establish that such policies are near optimal. In summary, our contribution is with regard to the general standard Borel setup where there is no apriori parametric representation.

math.OC

Approximations and Learning for Decentralized Stochastic Control and Near Optimal Finite Window Policies

Decentralized stochastic control problems are difficult to study due to information structure dependent subtleties, which prevent many classical methods in stochastic control from being applicable. In this paper we consider such problems with general standard Borel spaces under two related information structures. (a) the one-step delayed information sharing pattern (OSDISP) where agents share their information with one-step delay, and (b) the $K$-step periodic information sharing pattern (KSPISP), where information is shared periodically. It is known that OSDISP and KSPISP problems admit a centralized reduction where the agents view the problem from the perspective of a centralized controller that uses the common information to prescribe function valued actions (local policies) which map each agent's private information to an optimal action in the original problem. We provide rigorous approximation results and performance bounds for the KSPISP and OSDISP problems, which results from replacing the full common information by a finite sliding window of information and we establish near optimality of such policies. The latter depends on a predictor stability condition in expected total variation. As a further contribution, we show that under the information structures provided, corresponding Q-learning algorithms (in quantized or finite memory forms) converge asymptotically to near optimal solutions. While restrictive and hypothetical conditions have been presented in the literature, our contributions are thus to provide, to our knowledge, the first explicit conditions and rigorous approximation and learning results for such decentralized problems with general spaces.

math.OC

Centralized Reduction of Decentralized Stochastic Control Models and their weak-Feller Regularity

Decentralized stochastic control problems involving general state/measurement/action spaces are intrinsically difficult to study because of the inapplicability of standard tools from centralized (single-agent) stochastic control. In this paper, we address some of these challenges for decentralized stochastic control with standard Borel spaces under two different but tightly related information structures: the one-step delayed information sharing pattern (OSDISP), and the $K$-step periodic information sharing pattern (KSPISP). We will show that the one-step delayed and $K$-step periodic problems can be reduced to a centralized Markov Decision Process (MDP), generalizing prior results which considered finite, linear, or static models, by addressing several measurability and topological questions. We then provide sufficient conditions for the transition kernels of both centralized reductions to be weak-Feller. The existence and separated nature of optimal policies under both information structures are then established. The weak Feller regularity also facilitates rigorous approximation and learning theoretic results, as shown in the paper.

math.OC

Non-Sequential Decentralized Stochastic Control Revisited: Causality and Static Reducibility

In decentralized stochastic control (or stochastic team theory) and game theory, if there is a pre-defined order in a system in which agents act, the system is called \textit{sequential}, otherwise it is non-sequential. Much of the literature on stochastic control theory, such as studies on the existence analysis, approximation methods, and on dynamic programming or other analytical or learning theoretic methods, have focused on sequential systems. Many complex practical systems, however, are non-sequential where the order of agents acting is random, and dependent on the realization of solution paths and prior actions taken. The study of such systems is particularly challenging as tools applicable for sequential models are not directly applicable. In this paper, we will first revisit the notion of Causality (a definition due to Witsenhausen and which has been refined by Andersland and Tekenetzis), and provide an alternative representation using imaginary agents. We show that Causality is equivalent to Causal Implementability (and Dead-Lock Freeness), thus, generalizing previous results. We show that Causality, under an absolute continuity condition, allows for an equivalent static model whose reduction is policy-independent. Since the static reduction method for sequential control problems (via change of measures or other techniques), has been shown to be very effective in arriving at existence, structural, approximation and learning theoretic results, our analysis facilitates much of the stochastic analysis available for sequential systems to also be applicable for a class of non-sequential systems.

math.OC