Searcharxiv⌕ Search

arXiv subjects

Tom Erez

Publications and source records attributed to Tom Erez.

24 records · Page 2Linked to original sources

Emergence of Locomotion Behaviours in Rich Environments

The reinforcement learning paradigm allows, in principle, for complex behaviours to be learned directly from simple reward signals. In practice, however, it is common to carefully hand-design the reward function to encourage a particular solution, or to derive it from demonstration data. In this paper explore how a rich environment can help to promote the learning of complex behavior. Specifically, we train agents in diverse environmental contexts, and find that this encourages the emergence of robust behaviours that perform well across a suite of tasks. We demonstrate this principle for locomotion -- behaviours that are known for their sensitivity to the choice of reward. We train several simulated bodies on a diverse set of challenging terrains and obstacles, using a simple reward function based on forward progress. Using a novel scalable variant of policy gradient reinforcement learning, our agents learn to run, jump, crouch and turn as required by the environment without explicit reward-based guidance. A visual depiction of highlights of the learned behavior can be viewed following https://youtu.be/hx_bgoTF7bs .

cs.AI↗

Data-efficient Deep Reinforcement Learning for Dexterous Manipulation

Deep learning and reinforcement learning methods have recently been used to solve a variety of problems in continuous control domains. An obvious application of these techniques is dexterous manipulation tasks in robotics which are difficult to solve using traditional control theory or hand-engineered approaches. One example of such a task is to grasp an object and precisely stack it on another. Solving this difficult and practically relevant problem in the real world is an important long-term goal for the field of robotics. Here we take a step towards this goal by examining the problem in simulation and providing models and techniques aimed at solving it. We introduce two extensions to the Deep Deterministic Policy Gradient algorithm (DDPG), a model-free Q-learning based method, which make it significantly more data-efficient and scalable. Our results show that by making extensive use of off-policy data and replay, it is possible to find control policies that robustly grasp objects and stack them. Further, our results hint that it may soon be feasible to train successful stacking policies by collecting interactions on real robots.

cs.LG↗

Learning Continuous Control Policies by Stochastic Value Gradients

We present a unified framework for learning continuous control policies using backpropagation. It supports stochastic control by treating stochasticity in the Bellman equation as a deterministic function of exogenous noise. The product is a spectrum of general policy gradient algorithms that range from model-free methods with value functions to model-based methods without value functions. We use learned models but only require observations from the environment in- stead of observations from model-predicted trajectories, minimizing the impact of compounded model errors. We apply these algorithms first to a toy stochastic control problem and then to several physics-based control problems in simulation. One of these variants, SVG(1), shows the effectiveness of learning models, value functions, and policies simultaneously in continuous domains.

cs.LG↗

A Scalable Method for Solving High-Dimensional Continuous POMDPs Using Local Approximation

Partially-Observable Markov Decision Processes (POMDPs) are typically solved by finding an approximate global solution to a corresponding belief-MDP. In this paper, we offer a new planning algorithm for POMDPs with continuous state, action and observation spaces. Since such domains have an inherent notion of locality, we can find an approximate solution using local optimization methods. We parameterize the belief distribution as a Gaussian mixture, and use the Extended Kalman Filter (EKF) to approximate the belief update. Since the EKF is a first-order filter, we can marginalize over the observations analytically. By using feedback control and state estimation during policy execution, we recover a behavior that is effectively conditioned on incoming observations despite the unconditioned planning. Local optimization provides no guarantees of global optimality, but it allows us to tackle domains that are at least an order of magnitude larger than the current state-of-the-art. We demonstrate the scalability of our algorithm by considering a simulated hand-eye coordination domain with 16 continuous state dimensions and 6 continuous action dimensions.

cs.AI↗

Social Anti-Percolation and Negative Word of Mouth

Many new products fail, despite preliminary market surveys having determined considerable potential market share. This effect is too systematic to be attributed to bad luck. We suggest an explanation by presenting a new percolation theory model for product propagation, where agents interact over a social network. In our model, agents who do not adopt the product spread negative word of mouth to their neighbors, and so their neighborhood becomes less susceptible to the product. The result is a dramatic increase in the percolation threshold. When the effect of negative word of mouth is strong enough, it is shown to block any product from spreading to a significant fraction of the network. So, rather then being rejected by a large fraction of the agents, the product gets blocked by the rejection of a negligible fraction of the potential market. The rest of the potential buyers do not adopt the product because they are never exposed to it: the negative word of mouth spread by initial rejectors suffocates the diffusion by negatively affecting the immediate neighborhood of the propagation front.

cond-mat.other↗

Statistical Economics on Multi-Variable Layered Networks

We propose a Statistical-Mechanics inspired framework for modeling economic systems. Each agent composing the economic system is characterized by a few variables of distinct nature (e.g. saving ratio, expectations, etc.). The agents interact locally by their individual variables: for example, people working in the same office may influence their peers' expectations (optimism/pessimism are contagious), while people living in the same neighborhood may influence their peers' saving patterns (stinginess/largeness are contagious). Thus, for each type of variable there exists a different underlying social network, which we refer to as a ``layer''. Each layer connects the same set of agents by a different set of links defining a different topology. In different layers, the nature of the variables and their dynamics may be different (Ising, Heisenberg, matrix models, etc). The different variables belonging to the same agent interact (the level of optimism of an agent may influence its saving level), thus coupling the various layers. We present a simple instance of such a network, where a one-dimensional Ising chain (representing the interaction between the optimist-pessimist expectations) is coupled through a random site-to-site mapping to a one-dimensional generalized Blume-Capel chain (representing the dynamics of the agents' saving ratios). In the absence of coupling between the layers, the one-dimensional systems describing respectively the expectations and the saving ratios do not feature any ordered phase (herding). Yet, such a herding phase emerges in the coupled system, highlighting the non-trivial nature of the present framework.

cond-mat.stat-mech↗