SearcharxivSearch

arXiv subjects

Piotr Milos

Publications and source records attributed to Piotr Milos.

9 recordsLinked to original sources

Contrastive Representations for Temporal Reasoning

In classical AI, perception relies on learning state-based representations, while planning, which can be thought of as temporal reasoning over action sequences, is typically achieved through search. We study whether such reasoning can instead emerge from representations that capture both perceptual and temporal structure. We show that standard temporal contrastive learning, despite its popularity, often fails to capture temporal structure due to its reliance on spurious features. To address this, we introduce Combinatorial Representations for Temporal Reasoning (CRTR), a method that uses a negative sampling scheme to provably remove these spurious features and facilitate temporal reasoning. CRTR achieves strong results on domains with complex temporal structure, such as Sokoban and Rubik's Cube. In particular, for the Rubik's Cube, CRTR learns representations that generalize across all initial states and allow it to solve the puzzle using fewer search steps than BestFS, though with longer solutions. To our knowledge, this is the first method that efficiently solves arbitrary Cube states using only learned representations, without relying on an external search algorithm.

cs.LG

Model-Based Reinforcement Learning for Atari

Model-free reinforcement learning (RL) can be used to learn effective policies for complex tasks, such as Atari games, even from image observations. However, this typically requires very large amounts of interaction -- substantially more, in fact, than a human would need to learn the same games. How can people learn so quickly? Part of the answer may be that people can learn how the game works and predict which actions will lead to desirable outcomes. In this paper, we explore how video prediction models can similarly enable agents to solve Atari games with fewer interactions than model-free methods. We describe Simulated Policy Learning (SimPLe), a complete model-based deep RL algorithm based on video prediction models and present a comparison of several model architectures, including a novel architecture that yields the best results in our setting. Our experiments evaluate SimPLe on a range of Atari games in low data regime of 100k interactions between the agent and the environment, which corresponds to two hours of real-time play. In most games SimPLe outperforms state-of-the-art model-free algorithms, in some games by over an order of magnitude.

cs.LG

Existence of a phase transition of the interchange process on the Hamming graph

The interchange process on a finite graph is obtained by placing a particle on each vertex of the graph, then at rate 1, selecting an edge uniformly at random and swapping the two particles at either end of this edge. In this paper we develop new techniques to show the existence of a phase transition of the interchange process on the 2-dimensional Hamming graph. We show that in the subcritical phase, all of the cycles of the process have length $O(\log n)$, whereas in the supercritical phase a positive density of vertices lie in cycles of length at least $n^{2-\varepsilon}$ for any $\varepsilon>0$.

math.PR

Occupation times of subcritical branching immigration systems with Markov motion, clt and deviations principles

In this paper we consider two related stochastic models. The first one is a branching system consisting of particles moving according to a Markov family in R^d and undergoing subcritical branching with a constant rate of V>0. New particles immigrate to the system according to a homogeneous space time Poisson random field. The second model is the superprocess corresponding to the branching particle system. We study rescaled occupation time process and the process of its fluctuations with very mild assumptions on the Markov family. In the general setting a functional central limit theorem as well as large and moderate deviations principles are proved. The subcriticality of the branching law determines the behaviour in large time scales and in "overwhelms" the properties of the particles' motion. For this reason the results are the same for all dimensions and can be obtained for a wide class of Markov processes (both properties are unusual for systems with critical branching).

math.PR

Fluctuations of the occupation times for branching system starting from infinitely divisible point processes

In the paper the rescaled occupation time fluctuation process of a certain empirical system is investigated. The system consists of particles evolving independently according to α-stable motion in R^d, α 0. We study how the limit behaviour of the fluctuations of the occupation time depends on the \emph{initial particle configuration}. We obtain a functional central limit theorem for a vast class of infinitely divisible distributions. Our findings extend and put in a unified setting results which previously seemed to be disconnected. The limit processes form a one dimensional family of long-range dependance centred Gaussian processes.

math.PR

Occupation times of subcritical branching immigration system with Markov motions

We consider a branching system consisting of particles moving according to a Markov family in $\Rd$ and undergoing subcritical branching with a constant rate $V>0$. New particles immigrate to the system according to homogeneous space-time Poisson random field. The process of the fluctuations of the rescaled occupation time is studied with very mild assumptions on the Markov family. In this general setting a functional central limit theorem is proved. The subcriticality of the branching law is crucial for the limit behaviour and in a sense overwhelms the properties of the particles' motion. It is for this reason that the limit is the same for all dimensions and can be obtained for a wide class of Markov processes. Another consequence is the form of the limit - $\SP$-valued Wiener process with a simple temporal structure and a complicated spatial one. This behaviour contrasts sharply with the case of critical branching systems.

math.PR

Occupation time fluctuation limits of infinite variance equilibrium branching systems

We establish limit theorems for the fluctuations of the rescaled occupation time of a $(d,α,β)$-branching particle system. It consists of particles moving according to a symmetric $α$-stable motion in $\mathbb{R}^d$. The branching law is in the domain of attraction of a (1+$β$)-stable law and the initial condition is an equilibrium random measure for the system (defined below). In the paper we treat separately the cases of intermediate $α/β (1+β)α/β$ dimensions. In the most interesting case of intermediate dimensions we obtain a version of a fractional stable motion. The long-range dependence structure of this process is also studied. Contrary to this case, limit processes in critical and large dimensions have independent increments.

math.PR

Occupation time fluctuations of Poisson and equilibrium branching systems in critical and large dimensions

Limit theorems are presented for the rescaled occupation time fluctuation process of a critical finite variance branching particle system in $\mathbb{R}^{d}$ with symmetric $α$-stable motion starting off from either a standard Poisson random field or the equilibrium distribution for critical $d=2α$ and large $d>2α$ dimensions. The limit processes are generalised Wiener processes. The obtained convergence is in space-time, finite-dimensional distributions sense. With the addtional assumption on the branching law we obtain functional convergence.

math.PR

Occupation time fluctuations of Poisson and equilibrium finite variance branching systems

Functional limit theorems are presented for the rescaled occupation time fluctuations process of a critical finite variance branching particle system in $R^d$ with symmetric a-stable motion starting off from either a standard Poisson random field or from the equilibrium distribution for intermediate dimensions a<d<2a. The limit processes are determined sub-fractional and fractional Brownian motion respectively.

math.PR