SearcharxivSearch

arXiv subjects

Yuhang Mei

Publications and source records attributed to Yuhang Mei.

4 recordsLinked to original sources

Time-Reversed BSDEs for Accurate Gradient Estimation in Diffusion Models

There is a growing literature adopting a stochastic optimal control (SOC) perspective to fine-tune diffusion models and related generative policies. A prominent class of methods, known as iterative diffusion optimization, solves the SOC problem by simulating the diffusion process, evaluating a loss function, and applying stochastic optimization algorithms, with adjoint matching emerging as a state-of-the-art approach. However, the adjoint process used in these methods is not adapted to the forward diffusion filtration, which can lead to unstable or high-variance gradient estimates. In this paper, we revisit gradient estimation in diffusion models through the lens of backward stochastic differential equations (BSDEs). We propose an alternative estimator based on a time-reversed BSDE formulation introduced in our prior work, which produces an adjoint process adapted to the underlying filtration. This adapted structure leads to more stable gradient estimates with potentially lower variance. We analyze the accuracy of the proposed estimator and compare it with adjoint matching. Numerical experiments on fine-tuning toy diffusion models demonstrate improved gradient stability and competitive performance.

math.OC

A Time-Reversal Control Synthesis for Steering the State of Stochastic Systems

This paper presents a novel approach for steering the state of a stochastic control-affine system to a desired target within a finite time horizon. Our method leverages the time-reversal of diffusion processes to construct the required feedback control law. Specifically, the control law is the so-called score function associated with the time-reversal of random state trajectories that are initialized at the target state and are simulated backwards in time. A neural network is trained to approximate the score function, enabling applicability to both linear and nonlinear stochastic systems. Numerical experiments demonstrate the effectiveness of the proposed method across several benchmark examples.

math.OC

Flow matching for stochastic linear control systems

This paper addresses the problem of steering an initial probability distribution to a target probability distribution through a deterministic or stochastic linear control system. Our proposed approach is inspired by the flow matching methodology, with the difference that we can only affect the flow through the given control channels. The motivation comes from applications such as robotic swarms and stochastic thermodynamics, where agents or particles can only be manipulated through control actions. The feedback control law that achieves the task is characterized as the conditional expectation of the control inputs for the stochastic bridges that respect the given control system dynamics. Explicit forms are derived for special cases, and a numerical procedure is presented to approximate the control law, illustrated with examples.

math.OC

Time-reversal solution of BSDEs in stochastic optimal control: a linear quadratic study

This paper addresses the numerical solution of backward stochastic differential equations (BSDEs) arising in stochastic optimal control. Specifically, we investigate two BSDEs: one derived from the Hamilton-Jacobi-Bellman equation and the other from the stochastic maximum principle. For both formulations, we analyze and compare two numerical methods. The first utilizes the least-squares Monte-Carlo (LSMC) approach for approximating conditional expectations, while the second leverages a time-reversal (TR) of diffusion processes. Although both methods extend to nonlinear settings, our focus is on the linear-quadratic case, where analytical solutions provide a benchmark. Numerical results demonstrate the superior accuracy and efficiency of the TR approach across both BSDE representations, highlighting its potential for broader applications in stochastic control.

math.OC