SearcharxivSearch

arXiv · 2604.13096

Complexity scaling and optimal policy degeneracy in quantum reinforcement learning via analytically solvable unitary-control-then-measure models

Abstract

We propose and analyse a class of analytically solvable models of quantum reinforcement learning (QRL), formulated as finite-horizon Markov decision processes in finite-dimensional Hilbert spaces. The models are built around a `unitary-control-then-measure' protocol, in which a learning agent applies unitary transformations to a quantum state and interleaves each control step with a projective measurement onto a prescribed reference basis. Exact closed-form expressions for trajectory probabilities, rewards, and the expected return are derived for four concrete realisations: a closed-chain and an anti-periodic qubit implementation, a qutrit model with ladder coupling, and a four-level two-qubit system. Two structural features of these QRL protocols are then analysed. First, we identify and quantify the reduction in the computational complexity of the expected return, from the nominally exponential $O(e^N)$ scaling in the trajectory length~$N$ to an explicit power-law $O(N^{\mathcal{I}})$, driven by two rigorously established mechanisms, a trajectory equivalence and a sparsity of the transition graph, besides a third, conjectured one: a spectral concentration of the return, at the optimal policy, onto the polynomially populated trajectory classes. Second, we characterise the degeneracy of optimal policies. The low-dimensional models exhibit unique optima whose asymptotic behaviour with~$N$ is governed by the quantum Zeno effect, while the four-level system displays both plateau-type quasi-degeneracy at large horizons and genuine discrete degeneracy at critical energy parameters -- phenomena with no counterpart in the measurement-free quantum optimal control landscape.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Andrea Cintio, Alessandro Michelangeli, Dmitrii Tsutskov. 2026-04-09. Complexity scaling and optimal policy degeneracy in quantum reinforcement learning via analytically solvable unitary-control-then-measure models. https://arxiv.org/abs/2604.13096

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Average Chord Lengths in a Triangle

Let $P$ be a point inside a triangle $T$. We consider the average length of the chords of $T$ through $P$, where the direction of the chord is chosen uniformly. An elementary formula is obtained in terms of the distances from $P$ to the sides and vertices of the triangle. Several classical triangle centers give especially simple specializations. For example, if $I$ is the incenter, then \[ M_T(I)=\frac{2r}{\pi} \log\left(\cot\frac A4\cot\frac B4\cot\frac C4\right). \] Our main result is the sharp inequality \[ M_T(P)\le \frac{p}{\pi\sqrt3}\log(2+\sqrt3), \] valid simultaneously for every triangle of perimeter $p$ and every interior point $P$. Thus, among all such pairs $(T,P)$, the largest possible average chord length occurs only when $T$ is equilateral and $P$ is its center. The proof is an elementary symmetrization argument. We close with brief remarks relating the problem to the radial center of a convex body, the electrostatic potential center of a triangle, and dual quermassintegrals.

math.GM

A Proof of Liu's Conjecture on the Fundamental Triangle Inequality

Let $a,b,c$ be the side lengths of a triangle, and let $R$ and $r$ denote its circumradius and inradius, respectively. We prove a conjecture of Liu stating that \[\sum_{\mathrm{cyc}} \left(\frac{a(b+c-a)}{bc}\right)^k \geq 2+\left(\frac{2r}{R}\right)^k,~~k>1, \] with the reverse inequality for $0<k<1$. The proof reduces the problem to three positive variables with fixed sum and product. We also determine the equality cases.

math.GM