SearcharxivSearch

arXiv subjects

Yueyang Zheng

Publications and source records attributed to Yueyang Zheng.

11 recordsLinked to original sources

A linear-quadratic partially observed Stackelberg stochastic differential game with multiple followers and its application to multi-agent formation control

In this paper, we study a linear-quadratic partially observed Stackelberg stochastic differential game problem in which a single leader and multiple followers are involved. We consider more practical formulation for partial information that none of them can observed the complete information and the followers know more than the leader. Some completely different methods including a novel state decomposition and orthogonal decomposition are applied to overcome the difficulties caused by partially observability which improves the tools and relaxes the constraint condition imposed on admissible control in the existing literature. More precisely, the followers encounter the standard linear-quadratic partially observed optimal control problems, however, a kind of forward-backward indefinite linear-quadratic partially observed optimal control problem is considered by the leader. Instead of maximum principle of forward-backward control systems, inspired by the existing work related to definite case and classical forward control system, some distinct forward-backward linear-quadratic decoupling techniques including the method of completion of squares are applied to solve the leader's problem. More interestingly, we develop the deterministic formation control in multi-agent system with a framework of Stackelberg differential game and extend it to the stochastic case. The optimal strategies are obtained by our theoretical result suitably.

math.OC

Linear-Quadratic Partially Observed Mean Field Stackelberg Stochastic Differential Game with Applications

This paper is concerned with a linear-quadratic partially observed mean field Stackelberg stochastic differential game, which contains a leader and a large number of followers. Specifically, the followers confront a large-population Nash game subsequent to the leader's initial announcement of his strategy. In turn, the leader optimizes his own cost functional, taking into account the anticipated reactions of the followers. The state equations of both the leader and the followers are general stochastic differential equations, where the drift terms contain both the state average term and the state expectation term. However, the followers' state average terms enter into the drift term of the leader's state equation and the state expectation term of the leader enters into the state equation of the follower, reflecting the mutual influence between the leader and the followers. By utilizing the techniques of state decomposition and backward separation principle, we deduce the open-loop adapted decentralized strategies and feedback decentralized strategies of this leader-followers system, and demonstrate that the decentralized strategies are the corresponding $\varepsilon$-Stackelberg-Nash equilibrium. Finally, we apply the theoretical result to a product planning problem with sticky prices.

math.OC

The Global Maximum Principle for Optimal Control of Partially Observed Stochastic Systems Driven by Fractional Brownian Motion

In this paper we study the stochastic control problem of partially observed (multi-dimensional) stochastic system driven by both Brownian motions and fractional Brownian motions. In the absence of the powerful tool of Girsanov transformation, we introduce and study new stochastic processes which are used to transform the original problem to a "classical one". The adjoint backward stochastic differential equations and the necessary condition satisfied by the optimal control (maximum principle) are obtained.

math.OC

ICARUS: A Specialized Architecture for Neural Radiance Fields Rendering

The practical deployment of Neural Radiance Fields (NeRF) in rendering applications faces several challenges, with the most critical one being low rendering speed on even high-end graphic processing units (GPUs). In this paper, we present ICARUS, a specialized accelerator architecture tailored for NeRF rendering. Unlike GPUs using general purpose computing and memory architectures for NeRF, ICARUS executes the complete NeRF pipeline using dedicated plenoptic cores (PLCore) consisting of a positional encoding unit (PEU), a multi-layer perceptron (MLP) engine, and a volume rendering unit (VRU). A PLCore takes in positions \& directions and renders the corresponding pixel colors without any intermediate data going off-chip for temporary storage and exchange, which can be time and power consuming. To implement the most expensive component of NeRF, i.e., the MLP, we transform the fully connected operations to approximated reconfigurable multiple constant multiplications (MCMs), where common subexpressions are shared across different multiplications to improve the computation efficiency. We build a prototype ICARUS using Synopsys HAPS-80 S104, a field programmable gate array (FPGA)-based prototyping system for large-scale integrated circuits and systems design. We evaluate the power-performance-area (PPA) of a PLCore using 40nm LP CMOS technology. Working at 400 MHz, a single PLCore occupies 16.5 $mm^2$ and consumes 282.8 mW, translating to 0.105 uJ/sample. The results are compared with those of GPU and tensor processing unit (TPU) implementations.

cs.AR

The Global Maximum Principle for Progressive Optimal Control of Partially Observed Forward-Backward Stochastic Systems with Random Jumps

IIn this paper, we study a partially observed progressive optimal control problem of forward-backward stochastic differential equations with random jumps, where the control domain is not necessarily convex, and the control variable enter into all the coefficients. In our model, the observation equation is not only driven by a Brownian motion but also a Poisson random measure, which also have correlated noises with the state equation. For preparation, we first derive the existence and uniqueness of the solutions to the fully coupled forward-backward stochastic system with random jumps and its estimation in $L^β(β\geq2)$-space under some assumptions, and the non-linear filtering equation of partially observed stochastic system with random jumps. Then we derive the partially observed global maximum principle with random jumps. To show its applications, a partially observed linear quadratic progressive optimal control problem with random jumps is investigated, by the maximum principle and stochastic filtering. State estimate feedback representation of the optimal control is given in a more explicit form by introducing some ordinary differential equations.

math.OC

The Maximum Principle for Discounted Optimal Control of Partially Observed Forward-Backward Stochastic Systems with Jumps on Infinite Horizon

This paper is concerned with a discounted optimal control problem of partially observed forward-backward stochastic systems with jumps on infinite horizon. The control domain is convex and a kind of infinite horizon observation equation is introduced. The uniquely solvability of infinite horizon forward (backward) stochastic differential equation with jumps is obtained and more extended analysis, especially for the backward case, is made. Some new estimates are first given and proved for the critical variational inequality. Then an ergodic maximum principle is obtained by introducing some infinite horizon adjoint equations whose uniquely solvabilities are guaranteed necessarily. Finally, some comparison are made with two kinds of representative infinite horizon stochastic systems and their related optimal controls.

math.OC

A Linear Quadratic Partially Observed Stackelberg Stochastic Differential Game with Applications

This paper is concerned with a linear-quadratic partially observed Stackelberg stochastic differential game with correlated state and observation noises, where the diffusion coefficient does not contain the control variable and the control set is not necessarily convex. Both the leader and the follower have their own observation equations, and the information filtration available to the leader is contained in that to the follower. By spike variational, state decomposition and backward separation techniques, necessary and sufficient conditions of the Stackelberg equilibrium points are derived. In the follower's problem, the state estimation feedback of optimal control can be represented by a forward-backward stochastic differential filtering equation and some Riccati equation. In the leader's problem, via the innovation process, the state estimation feedback of optimal control is represented by a stochastic differential filtering equation, a semi-martingale process and three high-dimensional Riccati equations. At the same time, the uniqueness and existence of solutions to adjoint equations can be guaranteed by a new combined idea, and a kind of fully coupled forward-backward stochastic differential equations with filtering is studied as a by-product. Then we give explicit expressions of Stackelberg equilibrium points in a special case. As a practical application, an inspiring dynamic advertising problem with asymmetric information is studied, and the effectiveness and reasonability of the theoretical result is illustrated by numerical simulations. Moreover, the relationship between optimal control, state estimate and some practical parameters is analyzed in detail.

math.OC

Stackelberg Stochastic Differential Game with Asymmetric Noisy Observations

This paper is concerned with a Stackelberg stochastic differential game with asymmetric noisy observation, with one follower and one leader. In our model, the follower cannot observe the state process directly, but could observe a noisy observation process, while the leader can completely observe the state process. Open-loop Stackelberg equilibrium is considered. The follower first solve an stochastic optimal control problem with partial observation, the maximum principle and verification theorem are obtained. Then the leader turns to solve an optimal control problem for a conditional mean-field forward-backward stochastic differential equation, and both maximum principle and verification theorem are proved. An linear-quadratic Stackelberg stochastic differential game with asymmetric noisy observation is discussed to illustrate the theoretical results in this paper. With the aid of some Riccati equations, the open-loop Stackelberg equilibrium admits its state estimate feedback representation.

math.OC

A Stackelberg Game of Backward Stochastic Differential Equations with Partial Information

This paper is concerned with a Stackelberg game of backward stochastic differential equations (BSDEs) with partial information, where the information of the follower is a sub-$σ$-algebra of that of the leader. Necessary and sufficient conditions of the optimality for the follower and the leader are first given for the general problem, by the partial information stochastic maximum principles of BSDEs and forward-backward stochastic differential equations (FBSDEs), respectively. Then a linear-quadratic (LQ) Stackelberg game of BSDEs with partial information is investigated. The state estimate feedback representation for the optimal control of the follower is first given via two Riccati equations. Then the leader's problem is formulated as an optimal control problem of FBSDE. Four high-dimensional Riccati equations are introduced to represent the state estimate feedback for the optimal control of the leader. Theoretic results are applied to a pension fund management problem of two players in the financial market.

math.OC

An Optimal Investment Problem under Correlated Noises: Risk-Sensitive Stochastic Control Approach

This paper is concerned with an optimal investment problem under correlated noises in the financial market, and the expected utility functional is hyperbolic absolute risk aversion (HARA) with the exponent $γ\neq0$. The problem can be reformulated as a risk-sensitive stochastic control problem. A new stochastic maximum principle is obtained first, where the adjoint equations and maximum condition heavily depend on the risk-sensitive parameter and the correlation coefficient. The optimal investment strategy is obtained explicitly in a state feedback form via the solution to a certain Riccati equation, under the risk-seeking case. Numerical simulation and figures are given to illustrate the sensitivity for the optimal investment strategy, with respect to the risk-sensitive parameter and the correlation coefficient.

math.OC

A Stackelberg Game of Backward Stochastic Differential Equations with Applications

This paper is concerned with a Stackelberg game of backward stochastic differential equations (BSDEs), where the coefficients of the backward system and the cost functionals are deterministic, and the control domain is convex. Necessary and sufficient conditions of the optimality for the follower and the leader are first given for the general problem, by the stochastic maximum principles of BSDEs and forward-backward stochastic differential equations (FBSDEs), respectively. Then a linear-quadratic (LQ) Stackelberg game of BSDEs is investigated under standard assumptions. The state feedback representation for the optimal control of the follower is first given via two Riccati equations. Then the leader's problem is formulated as an optimal control problem of FBSDE with the control-independent diffusion term. Two high-dimensional Riccati equations are introduced to represent the state feedback for the optimal control of the leader. The solvability of the four Riccati equations are discussed. Theoretic results are applied to an optimal consumption rate problem of two players in the financial market.

math.OC