SearcharxivSearch

arXiv subjects

Adrian Hutter

Publications and source records attributed to Adrian Hutter.

17 recordsLinked to original sources

InfAlign: Inference-aware language model alignment

Language model alignment is a critical step in training modern generative language models. Alignment targets to improve win rate of a sample from the aligned model against the base model. Today, we are increasingly using inference-time algorithms (e.g., Best-of-N, controlled decoding, tree search) to decode from language models rather than standard sampling. We show that this train/test mismatch makes standard RLHF framework sub-optimal in view of such inference-time methods. To this end, we propose a framework for inference-aware alignment (InfAlign), which aims to optimize inference-time win rate of the aligned policy against the base model. We prove that for any inference-time decoding procedure, the optimal aligned policy is the solution to the standard RLHF problem with a transformation of the reward. This motivates us to provide the calibrate-and-transform RL (InfAlign-CTRL) algorithm to solve this problem, which involves a reward calibration step and a KL-regularized reward maximization step with a transformation of the calibrated reward. For best-of-N sampling and best-of-N jailbreaking, we propose specific transformations offering up to 3-8% improvement on inference-time win rates. Finally, we also show that our proposed reward calibration method is a strong baseline for optimizing standard win rate.

cs.LG

Balancing Cooperativeness and Adaptiveness in the (Noisy) Iterated Prisoner's Dilemma

Ever since Axelrod's seminal work, tournaments served as the main benchmark for evaluating strategies in the Iterated Prisoner's Dilemma (IPD). In this work, we first introduce a strategy for the IPD which outperforms previous tournament champions when evaluated against the 239 strategies in the Axelrod library, at noise levels in the IPD ranging from 0% to 10%. The basic idea behind our strategy is to start playing a version of tit-for-tat which forgives unprovoked defections if their rate is not significantly above the noise level, while building a (memory-1) model of the opponent; then switch to a strategy which is optimally adapted to the model of the opponent. We then argue that the above strategy (like other prominent strategies) lacks a couple of desirable properties which are not well tested for by tournaments, but which will be relevant in other contexts: we want our strategy to be self-cooperating, i.e., cooperate with a clone with high probability, even at high noise levels; and we want it to be cooperation-inducing, i.e., optimal play against it should entail cooperating with high probability. We show that we can guarantee these properties, at a modest cost in tournament performance, by reverting from the strategy adapted to the opponent to the forgiving tit-for-tat strategy under suitable conditions

cs.GT

Learning in two-player games between transparent opponents

We consider a scenario in which two reinforcement learning agents repeatedly play a matrix game against each other and update their parameters after each round. The agents' decision-making is transparent to each other, which allows each agent to predict how their opponent will play against them. To prevent an infinite regress of both agents recursively predicting each other indefinitely, each agent is required to give an opponent-independent response with some probability at least epsilon. Transparency also allows each agent to anticipate and shape the other agent's gradient step, i.e. to move to regions of parameter space in which the opponent's gradient points in a direction favourable to them. We study the resulting dynamics experimentally, using two algorithms from previous literature (LOLA and SOS) for opponent-aware learning. We find that the combination of mutually transparent decision-making and opponent-aware learning robustly leads to mutual cooperation in a single-shot prisoner's dilemma. In a game of chicken, in which both agents try to manoeuvre their opponent towards their preferred equilibrium, converging to a mutually beneficial outcome turns out to be much harder, and opponent-aware learning can even lead to worst-case outcomes for both agents. This highlights the need to develop opponent-aware learning algorithms that achieve acceptable outcomes in social dilemmas involving an equilibrium selection problem.

cs.AI

Quantum Computing with Parafermions

$\mathbb{Z}_d$ Parafermions are exotic non-Abelian quasiparticles generalizing Majorana fermions, which correspond to the case $d=2$. In contrast to Majorana fermions, braiding of parafermions with $d>2$ allows to perform an entangling gate. This has spurred interest in parafermions and a variety of condensed matter systems have been proposed as potential hosts for them. In this work, we study the computational power of braiding parafermions more systematically. We make no assumptions on the underlying physical model but derive all our results from the algebraical relations that define parafermions. We find a familiy of $2d$ representations of the braid group that are compatible with these relations. The braiding operators derived this way reproduce those derived previously from physical grounds as special cases. We show that if a $d$-level qudit is encoded in the fusion space of four parafermions, braiding of these four parafermions allows to generate the entire single-qudit Clifford group (up to phases), for any $d$. If $d$ is odd, then we show that in fact the entire many-qudit Clifford group can be generated.

quant-ph

Continuous error correction for Ising anyons

Quantum gates in topological quantum computation are performed by braiding non-Abelian anyons. These braiding processes can presumably be performed with very low error rates. However, to make a topological quantum computation architecture truly scalable, even rare errors need to be corrected. Error correction for non-Abelian anyons is complicated by the fact that it needs to be performed on a continuous basis and further errors may occur while we are correcting existing ones. Here, we provide the first study of this problem and prove its feasibility, establishing non-Abelian anyons as a viable platform for scalable quantum computation. We thereby focus on Ising anyons as the most prominent example of non-Abelian anyons and show that for these a finite error rate can indeed be corrected continuously. There is a threshold error rate $p_c>0$ such that for all error rates $p<p_c$ the probability of a logical error per time-step can be made exponentially small in the distance of a logical qubit.

quant-ph

Active error correction for Abelian and non-Abelian anyons

We consider a class of decoding algorithms that are applicable to error correction for both Abelian and non-Abelian anyons. This class includes multiple algorithms that have recently attracted attention, including the Bravyi-Haah RG decoder. They are applied to both the problem of single shot error correction (with perfect syndrome measurements) and that of active error correction (with noisy syndrome measurements). For Abelian models we provide a threshold proof in both cases, showing that there is a finite noise threshold under which errors can be arbitrarily suppressed when any decoder in this class is used. For non-Abelian models such a proof is found for the single shot case. The means by which decoding may be performed for active error correction of non-Abelian anyons is studied in detail. Differences with the Abelian case are discussed.

quant-ph

Parafermions in a Kagome lattice of qubits for topological quantum computation

Engineering complex non-Abelian anyon models with simple physical systems is crucial for topological quantum computation. Unfortunately, the simplest systems are typically restricted to Majorana zero modes (Ising anyons). Here we go beyond this barrier, showing that the $\mathbb{Z}_4$ parafermion model of non-Abelian anyons can be realized on a qubit lattice. Our system additionally contains the Abelian $D(\mathbb{Z}_4)$ anyons as low-energetic excitations. We show that braiding of these parafermions with each other and with the $D(\mathbb{Z}_4)$ anyons allows the entire $d=4$ Clifford group to be generated. The error correction problem for our model is also studied in detail, guaranteeing fault-tolerance of the topological operations. Crucially, since the non-Abelian anyons are engineered through defect lines rather than as excitations, non-Abelian error correction is not required. Instead the error correction problem is performed on the underlying Abelian model, allowing high noise thresholds to be realized.

quant-ph

Improved HDRG decoders for qudit and non-Abelian quantum error correction

Hard-decision renormalization group (HDRG) decoders are an important class of decoding algorithms for topological quantum error correction. Due to their versatility, they have been used to decode systems with fractal logical operators, color codes, qudit topological codes, and non-Abelian systems. In this work, we develop a method of performing HDRG decoding which combines strenghts of existing decoders and further improves upon them. In particular, we increase the minimal number of errors necessary for a logical error in a system of linear size $L$ from $\Theta(L^{2/3})$ to $\Omega(L^{1-\epsilon})$ for any $\epsilon>0$. We apply our algorithm to decoding $D(\mathbb{Z}_d)$ quantum double models and a non-Abelian anyon model with Fibonacci-like fusion rules, and show that it indeed significantly outperforms previous HDRG decoders. Furthermore, we provide the first study of continuous error correction with imperfect syndrome measurements for the $D(\mathbb{Z}_d)$ quantum double models. The parallelized runtime of our algorithm is $\text{poly}(\log L)$ for the perfect measurement case. In the continuous case with imperfect syndrome measurements, the averaged runtime is $O(1)$ for Abelian systems, while continuous error correction for non-Abelian anyons stays an open problem.

quant-ph

Breakdown of Surface Code Error Correction Due to Coupling to a Bosonic Bath

We consider a surface code suffering decoherence due to coupling to a bath of bosonic modes at finite temperature and study the time available before the unavoidable breakdown of error correction occurs as a function of coupling and bath parameters. We derive an exact expression for the error rate on each individual qubit of the code, taking spatial and temporal correlations between the errors into account. We investigate numerically how different kinds of spatial correlations between errors in the surface code affect its threshold error rate. This allows us to derive the maximal duration of each quantum error correction period by studying when the single-qubit error rate reaches the corresponding threshold. At the time when error correction breaks down, the error rate in the code can be dominated by the direct coupling of each qubit to the bath, by mediated subluminal interactions, or by mediated superluminal interactions. For a 2D Ohmic bath, the time available per quantum error correction period vanishes in the thermodynamic limit of a large code size $L$ due to induced superluminal interactions, though it does so only like $1/\sqrt{\log L}$. For all other bath types considered, this time remains finite as $L\rightarrow\infty$.

quant-ph

Relative Thermalization

When studying thermalization of quantum systems, it is typical to ask whether a system interacting with an environment will evolve towards a local thermal state. Here, we show that a more general and relevant question is "when does a system thermalize relative to a particular reference?" By relative thermalization we mean that, as well as being in a local thermal state, the system is uncorrelated with the reference. We argue that this is necessary in order to apply standard statistical mechanics to the study of the interaction between a thermalized system and a reference. We then derive a condition for relative thermalization of quantum systems interacting with an arbitrary environment. This condition has two components: the first is state-independent, reflecting the structure of invariant subspaces, like energy shells, and the relative sizes of system and environment; the second depends on the initial correlations between reference, system and environment, measured in terms of conditional entropies. Intuitively, a small system interacting with a large environment is likely to thermalize relative to a reference, but only if, initially, the reference was not highly correlated with the system and environment. Our statement makes this intuition precise, and we show that in many natural settings this thermalization condition is approximately tight. Established results on thermalization, which usually ignore the reference, follow as special cases of our statements.

quant-ph

Enhanced thermal stability of the toric code through coupling to a bosonic bath

We propose and study a model of a quantum memory that features self-correcting properties and a lifetime growing arbitrarily with system size at non-zero temperature. This is achieved by locally coupling a 2D L x L toric code to a 3D bath of bosons hopping on a cubic lattice. When the stabilizer operators of the toric code are coupled to the displacement operator of the bosons, we solve the model exactly via a polaron transformation and show that the energy penalty to create anyons grows linearly with L. When the stabilizer operators of the toric code are coupled to the bosonic density operator, we use perturbation theory to show that the energy penalty for anyons scales with ln(L). For a given error model, these energy penalties lead to a lifetime of the stored quantum information growing respectively exponentially and polynomially with L. Furthermore, we show how to choose an appropriate coupling scheme in order to hinder the hopping of anyons (and not only their creation) with energy barriers that are of the same order as the anyon creation gaps. We argue that a toric code coupled to a 3D Heisenberg ferromagnet realizes our model in its low-energy sector. Finally, we discuss the delicate issue of the stability of topological order in the presence of perturbations. While we do not derive a rigorous proof of topological order, we present heuristic arguments suggesting that topological order remains intact when perturbative operators acting on the toric code spins are coupled to the bosonic environment.

quant-ph

Dynamic Generation of Topologically Protected Self-Correcting Quantum Memory

We propose a scheme to dynamically realize a quantum memory based on the toric code. The code is generated from qubit systems with typical two-body interactions (Ising, XY, Heisenberg) using periodic, NMR-like, pulse sequences. It allows one to encode the logical qubits without measurements and to protect them dynamically against the time evolution of the physical qubits. A weakly coupled cavity mode mediates a long-range attractive interaction between the stabilizer operators of the toric code, thereby suppressing the creation of thermal anyons. This significantly increases the lifetime of the memory compared to the code with noninteracting stabilizers. We investigate how the fidelity, with which the toric code is realized, depends on the period length T of the pulse sequence and the magnitude of possible pulse errors. We derive an optimal period T_opt that maximizes the fidelity.

quant-ph

An efficient Markov chain Monte Carlo algorithm for the surface code

Minimum-weight perfect matching (MWPM) has been been the primary classical algorithm for error correction in the surface code, since it is of low runtime complexity and achieves relatively low logical error rates [Phys. Rev. Lett. 108, 180501 (2012)]. A Markov chain Monte Carlo (MCMC) algorithm [Phys. Rev. Lett. 109, 160503 (2012)] is able to achieve lower logical error rates and higher thresholds than MWPM, but requires a classical runtime complexity which is super-polynomial in L, the linear size of the code. In this work we present an MCMC algorithm that achieves significantly lower logical error rates than MWPM at the cost of a polynomially increased classical runtime complexity. For error rates p close to the threshold, our algorithm needs a runtime complexity which is increased by O(L^2) relative to MWPM in order to achieve a lower logical error rate. If p is below an L-dependent critical value, no increase in the runtime complexity is necessary any longer. For p->0, the logical error rate achieved by our algorithm is exponentially smaller (in L) than that of MWPM, without requiring an increased runtime complexity. Our algorithm allows for trade-offs between runtime and achieved logical error rates as well as for parallelization, and can be also used to correct in the case of imperfect stabilizer measurements.

quant-ph

Effective quantum memory Hamiltonian from local two-body interactions

In [Phys. Rev. A 88, 062313 (2013)] we proposed and studied a model for a self-correcting quantum memory in which the energetic cost for introducing a defect in the memory grows without bounds as a function of system size. This positive behavior is due to attractive long-range interactions mediated by a bosonic field to which the memory is coupled. The crucial ingredients for the implementation of such a memory are the physical realization of the bosonic field as well as local five-body interactions between the stabilizer operators of the memory and the bosonic field. Here, we show that both of these ingredients appear in a low-energy effective theory of a Hamiltonian that involves only two-body interactions between neighboring spins. In particular, we consider the low-energy, long-wavelength excitations of an ordered Heisenberg ferromagnet (magnons) as a realization of the bosonic field. Furthermore, we present perturbative gadgets for generating the required five-spin operators. Our Hamiltonian involving only local two-body interactions is thus expected to exhibit self-correcting properties as long as the noise affecting it is in the regime where the effective low-energy description remains valid.

quant-ph

Self-correcting quantum memory with a boundary

We study the two-dimensional toric code Hamiltonian with effective long-range interactions between its anyonic excitations induced by coupling the toric code to external fields. It has been shown that such interactions allow to increase the lifetime of the stored quantum information arbitrarily by making $L$, the linear size of the memory, larger [Phys. Rev. A 82 022305 (2010)]. We show that for these systems the choice of boundary conditions (open boundaries as opposed to periodic boundary conditions) is not a mere technicality; the influence of anyons produced at the boundaries becomes in fact dominant for large enough $L$. This influence can be both beneficial or detrimental. In particular, we study an effective Hamiltonian proposed in [Phys. Rev. B 83 115415 (2011)] that describes repulsion between anyons and anyon holes. For this system, we find a lifetime of the stored quantum information that grows exponentially in $L^2$ for both periodic and open boundary conditions, though the exponent in the latter case is found to be less favourable. However, $L$ is upper-bounded through the breakdown of the perturbative treatment of the underlying Hamiltonian.

quant-ph

Dependence of a quantum mechanical system on its own initial state and the initial state of the environment it interacts with

We present a unifying framework to the understanding of when and how quantum mechanical systems become independent of their initial conditions and adapt macroscopic properties (like temperature) of the environment.By viewing this problem from an quantum information theory perspective, we are able to simplify it in a very natural and easy way. We first show that for any interaction between the system and the environment, and almost all initial states of the system, the question of how long the system retains memory of its initial conditions can be answered by studying the temporal evolution of just one special initial state. This special state thereby depends only on our knowledge of macroscopic parameters of the system. We provide a simple entropic inequality for this state that can be used to determine whether mosts states of the system have, or have not become independent of their initial conditions after time $t$. We discuss applications of our entropic criterion to thermalization times in systems with an effective light-cone and to quantum memories suffering depolarizing noise. We make a similar statement for almost all initial states of the environment, and finally provide a sufficient condition for which a system never thermalizes, but remains close to its initial state for all times.

quant-ph

Almost All Quantum States Have Low Entropy Rates for Any Coupling to the Environment

The joint state of a system that is in contact with an environment is called lazy, if the entropy rate of the system under any coupling to the environment is zero. Necessary and sufficient conditions have recently been established for a state to be lazy [ Phys. Rev. Lett. 106 050403 (2011)], and it was shown that almost all states of the system and the environment do not have this property [ Phys. Rev. A 81 052318 (2010)]. At first glance, this may lead us to believe that low entropy rates themselves form an exception, in the sense that most states are far from being lazy and have high entropy rates. Here, we show that in fact the opposite is true if the environment is sufficiently large. Almost all states of the system and the environment are pretty lazy-their entropy rates are low for any coupling to the environment.

quant-ph