SearcharxivSearch

arXiv subjects

H. Francis Song

Publications and source records attributed to H. Francis Song.

16 recordsLinked to original sources

From Motor Control to Team Play in Simulated Humanoid Football

Intelligent behaviour in the physical world exhibits structure at multiple spatial and temporal scales. Although movements are ultimately executed at the level of instantaneous muscle tensions or joint torques, they must be selected to serve goals defined on much longer timescales, and in terms of relations that extend far beyond the body itself, ultimately involving coordination with other agents. Recent research in artificial intelligence has shown the promise of learning-based approaches to the respective problems of complex movement, longer-term planning and multi-agent coordination. However, there is limited research aimed at their integration. We study this problem by training teams of physically simulated humanoid avatars to play football in a realistic virtual environment. We develop a method that combines imitation learning, single- and multi-agent reinforcement learning and population-based training, and makes use of transferable representations of behaviour for decision making at different levels of abstraction. In a sequence of stages, players first learn to control a fully articulated body to perform realistic, human-like movements such as running and turning; they then acquire mid-level football skills such as dribbling and shooting; finally, they develop awareness of others and play as a team, bridging the gap between low-level motor control at a timescale of milliseconds, and coordinated goal-directed behaviour as a team at the timescale of tens of seconds. We investigate the emergence of behaviours at different levels of abstraction, as well as the representations that underlie these behaviours using several analysis techniques, including statistics from real-world sports analytics. Our work constitutes a complete demonstration of integrated decision-making at multiple scales in a physically embodied multi-agent setting. See project video at https://youtu.be/KHMwq9pv7mg.

cs.AI

A Distributional View on Multi-Objective Policy Optimization

Many real-world problems require trading off multiple competing objectives. However, these objectives are often in different units and/or scales, which can make it challenging for practitioners to express numerical preferences over objectives in their native units. In this paper we propose a novel algorithm for multi-objective reinforcement learning that enables setting desired preferences for objectives in a scale-invariant way. We propose to learn an action distribution for each objective, and we use supervised learning to fit a parametric policy to a combination of these distributions. We demonstrate the effectiveness of our approach on challenging high-dimensional real and simulated robotics tasks, and show that setting different preferences in our framework allows us to trace out the space of nondominated solutions.

cs.LG

The Hanabi Challenge: A New Frontier for AI Research

From the early days of computing, games have been important testbeds for studying how well machines can do sophisticated decision making. In recent years, machine learning has made dramatic advances with artificial agents reaching superhuman performance in challenge domains like Go, Atari, and some variants of poker. As with their predecessors of chess, checkers, and backgammon, these game domains have driven research by providing sophisticated yet well-defined challenges for artificial intelligence practitioners. We continue this tradition by proposing the game of Hanabi as a new challenge domain with novel problems that arise from its combination of purely cooperative gameplay with two to five players and imperfect information. In particular, we argue that Hanabi elevates reasoning about the beliefs and intentions of other agents to the foreground. We believe developing novel techniques for such theory of mind reasoning will not only be crucial for success in Hanabi, but also in broader collaborative efforts, especially those with human partners. To facilitate future research, we introduce the open-source Hanabi Learning Environment, propose an experimental framework for the research community to evaluate algorithmic advances, and assess the performance of current state-of-the-art techniques.

cs.LG

Stabilizing Transformers for Reinforcement Learning

Owing to their ability to both effectively integrate information over long time horizons and scale to massive amounts of data, self-attention architectures have recently shown breakthrough success in natural language processing (NLP), achieving state-of-the-art results in domains such as language modeling and machine translation. Harnessing the transformer's ability to process long time horizons of information could provide a similar performance boost in partially observable reinforcement learning (RL) domains, but the large-scale transformers used in NLP have yet to be successfully applied to the RL setting. In this work we demonstrate that the standard transformer architecture is difficult to optimize, which was previously observed in the supervised learning setting but becomes especially pronounced with RL objectives. We propose architectural modifications that substantially improve the stability and learning speed of the original Transformer and XL variant. The proposed architecture, the Gated Transformer-XL (GTrXL), surpasses LSTMs on challenging memory environments and achieves state-of-the-art results on the multi-task DMLab-30 benchmark suite, exceeding the performance of an external memory architecture. We show that the GTrXL, trained using the same losses, has stability and performance that consistently matches or exceeds a competitive LSTM baseline, including on more reactive tasks where memory is less critical. GTrXL offers an easy-to-train, simple-to-implement but substantially more expressive architectural alternative to the standard multi-layer LSTM ubiquitously used for RL agents in partially observable environments.

cs.LG

V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous Control

Some of the most successful applications of deep reinforcement learning to challenging domains in discrete and continuous control have used policy gradient methods in the on-policy setting. However, policy gradients can suffer from large variance that may limit performance, and in practice require carefully tuned entropy regularization to prevent policy collapse. As an alternative to policy gradient algorithms, we introduce V-MPO, an on-policy adaptation of Maximum a Posteriori Policy Optimization (MPO) that performs policy iteration based on a learned state-value function. We show that V-MPO surpasses previously reported scores for both the Atari-57 and DMLab-30 benchmark suites in the multi-task setting, and does so reliably without importance weighting, entropy regularization, or population-based tuning of hyperparameters. On individual DMLab and Atari levels, the proposed algorithm can achieve scores that are substantially higher than has previously been reported. V-MPO is also applicable to problems with high-dimensional, continuous action spaces, which we demonstrate in the context of learning to control simulated humanoids with 22 degrees of freedom from full state observations and 56 degrees of freedom from pixel observations, as well as example OpenAI Gym tasks where V-MPO achieves substantially higher asymptotic scores than previously reported.

cs.AI

Relational Forward Models for Multi-Agent Learning

The behavioral dynamics of multi-agent systems have a rich and orderly structure, which can be leveraged to understand these systems, and to improve how artificial agents learn to operate in them. Here we introduce Relational Forward Models (RFM) for multi-agent learning, networks that can learn to make accurate predictions of agents' future behavior in multi-agent environments. Because these models operate on the discrete entities and relations present in the environment, they produce interpretable intermediate representations which offer insights into what drives agents' behavior, and what events mediate the intensity and valence of social interactions. Furthermore, we show that embedding RFM modules inside agents results in faster learning systems compared to non-augmented baselines. As more and more of the autonomous systems we develop and interact with become multi-agent in nature, developing richer analysis tools for characterizing how and why agents make decisions is increasingly necessary. Moreover, developing artificial agents that quickly and safely learn to coordinate with one another, and with humans in shared environments, is crucial.

cs.LG

Machine Theory of Mind

Theory of mind (ToM; Premack & Woodruff, 1978) broadly refers to humans' ability to represent the mental states of others, including their desires, beliefs, and intentions. We propose to train a machine to build such models too. We design a Theory of Mind neural network -- a ToMnet -- which uses meta-learning to build models of the agents it encounters, from observations of their behaviour alone. Through this process, it acquires a strong prior model for agents' behaviour, as well as the ability to bootstrap to richer predictions about agents' characteristics and mental states using only a small number of behavioural observations. We apply the ToMnet to agents behaving in simple gridworld environments, showing that it learns to model random, algorithmic, and deep reinforcement learning agents from varied populations, and that it passes classic ToM tasks such as the "Sally-Anne" test (Wimmer & Perner, 1983; Baron-Cohen et al., 1985) of recognising that others can hold false beliefs about the world. We argue that this system -- which autonomously learns how to model other agents in its world -- is an important step forward for developing multi-agent AI systems, for building intermediating technology for machine-human interaction, and for advancing the progress on interpretable AI.

cs.AI

A simple, distance-dependent formulation of the Watts-Strogatz model for directed and undirected small-world networks

Small-world networks---complex networks characterized by a combination of high clustering and short path lengths---are widely studied using the paradigmatic model of Watts and Strogatz (WS). Although the WS model is already quite minimal and intuitive, we describe an alternative formulation of the WS model in terms of a distance-dependent probability of connection that further simplifies, both practically and theoretically, the generation of directed and undirected WS-type small-world networks. In addition to highlighting an essential feature of the WS model that has previously been overlooked, this alternative formulation makes it possible to derive exact expressions for quantities such as the degree and motif distributions and global clustering coefficient for both directed and undirected networks in terms of model parameters.

cs.SI

Fluctuations and Entanglement spectrum in quantum Hall states

The measurement of quantum entanglement in many-body systems remains challenging. One experimentally relevant fact about quantum entanglement is that in systems whose degrees of freedom map to free fermions with conserved total particle number, exact relations hold relating the Full Counting Statistics associated with the bipartite charge fluctuations and the sequence of R\' enyi entropies. We draw a correspondence between the bipartite charge fluctuations and the entanglement spectrum, mediated by the R\' enyi entropies. In the case of the integer quantum Hall effect, we show that it is possible to reproduce the generic features of the entanglement spectrum from a measurement of the second charge cumulant only. Additionally, asking whether it is possible to extend the free fermion result to the $ν=1/3$ fractional quantum Hall case, we provide numerical evidence that the answer is negative in general. We further address the problem of quantum Hall edge states described by a Luttinger liquid, and derive expressions for the spectral functions of the real space entanglement spectrum at a quantum point contact realized in a quantum Hall sample.

cond-mat.mes-hall

Scaling of entanglement entropy across Lifshitz transitions

We investigate the scaling of the bipartite entanglement entropy across Lifshitz quantum phase transitions, where the topology of the Fermi surface changes without any changes in symmetry. We present both numerical and analytical results which show that Lifshitz transitions are characterized by a well-defined set of critical exponents for the entanglement entropy near the phase transition. In one dimension, we show that the entanglement entropy exhibits a length scale that diverges as the system approaches a Lifshitz critical point. In two dimensions, the leading and sub-leading coefficients of the scaling of entanglement entropy show distinct power-law singularities at critical points. The effect of weak interactions is investigated using the density matrix renormalization group algorithm.

cond-mat.str-el

Detecting Quantum Critical Points using Bipartite Fluctuations

We show that the concept of bipartite fluctuations F provides a very efficient tool to detect quantum phase transitions in strongly correlated systems. Using state of the art numerical techniques complemented with analytical arguments, we investigate paradigmatic examples for both quantum spins and bosons. As compared to the von Neumann entanglement entropy, we observe that F allows to find quantum critical points with a much better accuracy in one dimension. We further demonstrate that F can be successfully applied to the detection of quantum criticality in higher dimensions with no prior knowledge of the universality class of the transition. Promising approaches to experimentally access fluctuations are discussed for quantum antiferromagnets and cold gases.

cond-mat.str-el

d-wave Superfluid with Gapless Edges in a Cold Atom Trap

We consider a strongly repulsive fermionic gas in a two-dimensional optical lattice confined by a harmonic trapping potential. To address the strongly repulsive regime, we consider the $t-J$ Hamiltonian. The presence of the harmonic trapping potential enables the stabilization of coexisting and competing phases. In particular, at low temperatures, this allows the realization of a d-wave superfluid region surrounded by purely (gapless) normal edges. Solving the Bogoliubov-de Gennes equations and comparing with the local density approximation, we show that the proximity to the Mott insulator is revealed by a downturn of the Fermi liquid order parameter at the center of the trap where the d-wave gap has a maximum. The density profile evolves linearly with distance.

cond-mat.supr-con

Bipartite Fluctuations as a Probe of Many-Body Entanglement

We investigate in detail the behavior of the bipartite fluctuations of particle number $\hat{N}$ and spin $\hat{S}^z$ in many-body quantum systems, focusing on systems where such U(1) charges are both conserved and fluctuate within subsystems due to exchange of charges between subsystems. We propose that the bipartite fluctuations are an effective tool for studying many-body physics, particularly its entanglement properties, in the same way that noise and Full Counting Statistics have been used in mesoscopic transport and cold atomic gases. For systems that can be mapped to a problem of non-interacting fermions we show that the fluctuations and higher-order cumulants fully encode the information needed to determine the entanglement entropy as well as the full entanglement spectrum through the Rényi entropies. In this connection we derive a simple formula that explicitly relates the eigenvalues of the reduced density matrix to the Rényi entropies of integer order for any finite density matrix. In other systems, particularly in one dimension, the fluctuations are in many ways similar but not equivalent to the entanglement entropy. Fluctuations are tractable analytically, computable numerically in both density matrix renormalization group and quantum Monte Carlo calculations, and in principle accessible in condensed matter and cold atom experiments. In the context of quantum point contacts, measurement of the second charge cumulant showing a logarithmic dependence on time would constitute a strong indication of many-body entanglement.

cond-mat.mes-hall

Entanglement Entropy of the Two-Dimensional Heisenberg Antiferromagnet

We compute the von Neumann and generalized Rényi entanglement entropies in the ground-state of the spin-1/2 antiferromagnetic Heisenberg model on the square lattice using the modified spin-wave theory for finite lattices. The addition of a staggered magnetic field to regularize the Goldstone modes associated with symmetry-breaking is shown to be essential for obtaining well-behaved values for the entanglement entropy. The von Neumann and Rényi entropies obey an area law with additive logarithmic corrections, and are in good quantitative agreement with numerical results from valence bond quantum Monte Carlo and density matrix renormalization group calculations. We also compute the spin fluctuations and observe a multiplicative logarithmic correction to the area law in excellent agreement with quantum Monte Carlo calculations.

cond-mat.str-el

Entanglement from Charge Statistics: Exact Relations for Many-Body Systems

We present exact formulas for the entanglement and Rényi entropies generated at a quantum point contact (QPC) in terms of the statistics of charge fluctuations, which we illustrate with examples from both equilibrium and non-equilibrium transport. The formulas are also applicable to groundstate entanglement in systems described by non-interacting fermions in any dimension, which in one dimension includes the critical spin-1/2 XX and Ising models where conformal field theory predictions for the entanglement and Rényi entropies are reproduced from the full counting statistics. These results may play a crucial role in the experimental detection of many-body entanglement in mesoscopic structures and cold atoms in optical lattices.

cond-mat.mes-hall

General Relation between Entanglement and Fluctuations in One Dimension

In one dimension very general results from conformal field theory and exact calculations for certain quantum spin systems have established universal scaling properties of the entanglement entropy between two parts of a critical system. Using both analytical and numerical methods, we show that if particle number or spin is conserved, fluctuations in a subsystem obey identical scaling as a function of subsystem size, suggesting that fluctuations are a useful quantity for determining the scaling of entanglement, especially in higher dimensions. We investigate the effects of boundaries and subleading corrections for critical spin and bosonic chains.

cond-mat.stat-mech