SearcharxivSearch

arXiv · 2609.16266

Towards Surrogate Based Dequantization of Quantum Reinforcement Learning

Abstract

In recent years, the utility of parameterized quantum circuits as function approximators has been widely studied. In the context of reinforcement learning, this approach has led to variational quantum algorithms such as quantum Q-learning. While these methods show promising empirical results, and can provide provable advantages for artificial problems, it remains unclear whether they can provide a provable quantum advantage over classical approaches for problems of practical relevance. A natural way to investigate this question is through the lens of dequantization: The construction of efficient classical algorithms capable of matching the performance of quantum variational methods. Building on recent kernel-based dequantization results for supervised learning, we take steps towards extending this surrogate-based dequantization program to reinforcement learning. Specifically, we study the simplified setting of reinforcement learning with a uniform generative model in which uniformly random state-action samples are available, which models the regime of sampling from a large experience replay buffer after sufficient exploration. Within this setting, we provide finite sample guarantees for classical kernelized Fitted Q-Iteration, with classical kernels designed to match the inductive bias of particular parameterized quantum circuits. Using these results, we then provide a set of sufficient conditions, on the data-encoding strategy of a parameterized quantum circuit, the corresponding classical kernel, and the problem structure, under which kernelized Fitted Q-Iteration provides a meaningful dequantization of quantum Q-learning, in this simplified setting. Apart from providing rigorous dequantization guarantees when these conditions are met, these results also motivate the use of kernelized fitted Q-iteration as a dequantization heuristic when these sufficient conditions cannot be verified.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Pablo Rodriguez-Grasa, Sofiene Jerbi, Mikel Sanz, Ryan Sweke. 2026-09-14. Towards Surrogate Based Dequantization of Quantum Reinforcement Learning. https://arxiv.org/abs/2609.16266

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Single-Ensemble Multiparameter Squeezing with Qudits

Conventional spin squeezing enhances a single sensing channel. Here, we show how internal qudit levels enable simultaneous multiparameter squeezing within one ensemble. In two-component magnetometry, a qutrit sensor provides two orthogonal and weakly compatible channels. A collective twisting interaction squeezes both responses while preserving joint attainability of the ultimate sensitivity. The sensing gain is quantified by using a matrix generalization of the Wineland sensitivity that retains both noise correlations and cross-channel response. An interaction-based echo amplifies the signal to overcome noise from a fixed local joint readout, yielding a simulated $13~\mathrm{dB}$ gain over the product-state standard quantum limit for $N=128$ qutrits. More generally, we use the single-site quantum Fisher information matrix to select reference states and channel quadratures for prescribed sensing tasks. The tangent geometry permits at most $d-1$ independent, weakly compatible channels around a common pure reference state for a $d$-level sensor. Our work provides a constructive task-to-protocol map for multiparameter squeezing in a single qudit ensemble.

quant-ph

A Design Space Study of Density Matrix Parameterizations for Diffusion-Based Quantum State Tomography

Diffusion-based quantum state tomography (QST) has shown promising results, but all existing methods implicitly adopt a single parameterization (typically Cholesky) without systematic evaluation. We present the first design space study of density matrix parameterizations for diffusion QST, introducing a geometric framework based on the Jacobian Gram matrix $\mathbf{J}^\top\mathbf{J}$. Our calibration of seven parameterizations at 2- and 3-qubit scales, validated by end-to-end training, reveals that \emph{geometric conditioning alone does not predict end-to-end performance}: at 3-qubit scale, Hermitian direct ($κ= 2.0\times$) performs worse than Cholesky ($κ= 27\times$) at all shot levels---a $13.5\times$ isotropy advantage that translates into a fidelity \emph{disadvantage} of up to $+0.51$. The 2-qubit ranking (Hermitian $>$ Bloch) reverses at 3 qubits (Bloch 0.907 vs.\ Hermitian 0.394). We provide a geometric explanation: unbounded parameterizations suffer projection-induced information loss because the PSD constraint couples diagonal and off-diagonal coordinates in ways the unconstrained model cannot respect, whereas the Bloch representation places the maximally mixed state at the center of the valid region, minimizing projection loss.

quant-ph

Entanglement free Metrology Exploiting Multimode Hong Ou Mandel Sensor Advantage

The Hong-Ou-Mandel (HOM) interference in the multimode frequency domain has been explored for precision metrology, with several experimental demonstrations exploiting its robustness against dispersion and phase noise, as well as its large dynamic range and compatibility with fragile samples. Conventional multimode HOM metrology exploits frequency-entangled states, which naturally satisfy bosonic exchange symmetry under any centered symmetric joint spectral distribution, to provide these advantages. However, these entangled states are typically generated via spontaneous parametric down-conversion (SPDC), requiring strong pump lasers that hinder practical implementation. In this paper, we employ frequency product states, which do not possess entanglement or path-mode exchange symmetry, as the probe state and post-select measurement outcomes exhibiting frequency anti-correlation. Our results demonstrate that these advantages,peak narrowing, dispersion cancellation, phase-noise immunity, a large dynamic range, and compatibility with fragile samples, arise neither from entanglement nor from bosonic exchange symmetry, but rather from spectral anti-correlation. We further show that entanglement is not the source of the measurement precision: the entanglement-free approach attains the same quantum Fisher information as the entangled-state scheme, indicating that the fundamental precision limit does not originate from entanglement.

quant-ph