arXiv · 2609.05808
Physical policy gradient theorem for in situ stochastic-adjoint training
Abstract
In situ adjoint training extracts parameter gradients directly from measurement, but has so far been limited to reciprocal or restricted systems. Here, we introduce the physical counterpart of the policy gradient theorem: a stochastic-adjoint gradient estimator that lifts these constraints by trading reciprocity for nondegenerate diffusion. As validation, we train a nonlinear resonator network, whose own dynamics supply the policy, against antagonistic temporal modulations with gradients from measured stochastic trajectories alone, without finite differences or a separate adjoint experiment.
Explore related subjects
Keep this discovery
William Tuxbury, Zin Lin. 2026-09-05. Physical policy gradient theorem for in situ stochastic-adjoint training. https://arxiv.org/abs/2609.05808
Cite the original work for its findings. Save a collection to share your selection of sources.
Discover connections
Connections use source metadata and explicit phrase matches, not verified experimental comparisons.