SearcharxivSearch

arXiv subjects

Xuefeng Bao

Publications and source records attributed to Xuefeng Bao.

3 recordsLinked to original sources

Analysis on the Derivation of the Schrödinger Equation with Analogy to Electromagnetic Wave Equation

The Schrödinger equation is universally accepted due to its excellent predictions aligning with observed results within its defined conditions. Nevertheless, it does not seem to possess the simplicity of fundamental laws, such as Newton's laws of motion. Various insightful attempts have been made to elucidate the rationale behind the Schrödinger equation. This paper seeks to review existing explanations and propose some prospectives on the derivation of the Schrödinger equation.

quant-ph

A Supplementary Condition for the Convergence of the Control Policy during Adaptive Dynamic Programming

Reinforcement learning based adaptive/approximate dynamic programming (ADP) is a powerful technique to determine an approximate optimal controller for a dynamical system. These methods bypass the need to analytically solve the nonlinear Hamilton-Jacobi-Bellman equation, whose solution is often to difficult to determine but is needed to determine the optimal control policy. ADP methods usually employ a policy iteration algorithm that evaluates and improves a value function at every step to find the optimal control policy. Previous works in ADP have been lacking a stronger condition that ensures the convergence of the policy iteration algorithm. This paper provides a sufficient but not necessary condition that guarantees the convergence of an ADP algorithm. This condition may provide a more solid theoretical framework for ADP-based control algorithm design for nonlinear dynamical systems.

math.OC

A Theoretical Difficulty in Approximate Dynamic Programming with Input Constraints

Equipping approximate dynamic programming (ADP) with inputconstraints has a tremendous significance. This enables ADP to be applied tothe systems with actuator limitations, which is quite common for dynamicalsystems. In a conventional constrained ADP framework, the optimal control issearched via a policy iteration algorithm, where the value under a constrainedcontrol is solved from a Hamilton-Jacobi-Bellman (HJB) equation while theconstrained control policy is improved based on the current estimated value.This concise and applicable method has been widely-used. However, the con-vergence of the existing policy iteration algorithm may possesses a theoreticaldifficulty, which might be caused by forcibly evaluating the same trajectoryeven though the control policy has already changed. This problem will beexplored in this paper.

math.OC