SearcharxivSearch

arXiv subjects

Alfred Harwood

Publications and source records attributed to Alfred Harwood.

5 recordsLinked to original sources

Fragility of Value under Imperfect Alignment

As more responsibility is placed upon AI systems, it becomes increasingly important to guarantee that these systems are aligned with humanity. A common fear in AI safety is that human value is fragile -- that is, optimizing too heavily for an imperfect proxy to human values will lead to a catastrophic outcome. In this paper, we present a model of the alignment problem where an agent undergoes idealized alignment training that guarantees its value function satisfies a proxy condition before optimizing the world. Our primary results identify conditions on the human value function and the accuracy of several proxy conditions under which an agent with an $\eta$-catastrophic value function, one that is guaranteed to take the expectation of human value below $\eta$ in the limit of optimizing power, would be deployed. Our results highlight the danger of overoptimization and motivate AI designs that limit optimization pressure, such as quantilizers, rather than relying solely on pre-deployment training.

cs.AI

Calculating Mutual Information between a Reward Maximizer and its Environment

An important question in the field of AI is the extent to which successful behaviour requires an internal representation of the world. In this work, we quantify the amount of information an optimal policy provides about the underlying environment. We consider a Controlled Markov Process (CMP) with $n$ states and $m$ actions, assuming a uniform prior over the space of possible transition dynamics. We prove that observing a deterministic policy that is optimal for any non-constant reward function then conveys exactly $n \log m$ bits of information about the environment. Specifically, we show that the mutual information between the environment and the optimal policy is $n \log m$ bits. This bound holds across a broad class of objectives, including finite-horizon, infinite-horizon discounted, and time-averaged reward maximization. These findings provide a precise information-theoretic lower bound on the ``implicit world model'' necessary for optimality.

cs.AI

Unified Collision Model of Coherent and Measurement-based Quantum Feedback

We introduce a general framework, based on collision models and discrete CP-maps, to describe on an equal footing coherent and measurement-based feedback control of quantum mechanical systems. We apply our framework to prominent tasks in quantum control, ranging from cooling to Hamiltonian control. Unlike other proposed comparisons, where coherent feedback always proves superior, we find that either measurements or coherent manipulations of the controller can be advantageous depending on the task at hand. Measurement-based feedback is typically superior in cooling, whilst coherent feedback is better at assisting quantum operations. Furthermore, we show that both coherent and measurement-based feedback loops allow one to simulate arbitrary Hamiltonian evolutions, and discuss their respective effectiveness in this regard.

quant-ph

Cavity optomechanics assisted by optical coherent feedback

We consider a wide family of optical coherent feedback loops acting on an optomechanical system operating in the linearized regime. We assess the efficacy of such loops in improving key operations, such as cooling, steady-state squeezing and entanglement, as well as optical to mechanical state transfer. We find that mechanical sideband cooling can be enhanced through passive, interferometric coherent feedback, achieving lower steady-state occupancies and considerably speeding up the cooling process; we also quantify the detrimental effect of non-zero delay times on the cooling performance. Steady state entanglement generation in the blue sideband can also be assisted by passive interferometric feedback, which allows one to stabilise otherwise unstable systems, though active feedback (including squeezing elements) does not help to this aim. We show that active feedback loops only allow for the generation of optical, but not mechanical squeezing. Finally, we prove that passive feedback can assist state transfer at transient times for red-sideband driven systems in the strong coupling regime.

quant-ph

Ultimate squeezing through coherent quantum feedback: A fair comparison with measurement-based schemes

We develop a general framework to describe interferometric coherent feedback loops and prove that, under any such scheme, the steady-state squeezing of a bosonic mode subject to a rotating wave coupling with a white noise environment and to any quadratic Hamiltonian must abide by a noise-dependent bound that reduces to the 3dB limit at zero temperature. Such a finding is contrasted, at fixed dynamical parameters, with the performance of homodyne continuous monitoring of the output modes. The latter allows one to beat coherent feedback and the 3dB limit under certain dynamical conditions, which will be determined exactly.

quant-ph