SearcharxivSearch

arXiv subjects

Zhanhao Zhang

Publications and source records attributed to Zhanhao Zhang.

14 recordsLinked to original sources

Deep Learning Method for Stationary Distribution of Reflected Brownian Motion

The stationary distribution of reflected Brownian motion (RBM) plays an important role in the analysis of high-dimensional stochastic systems, yet closed-form solutions are known only for a few special cases. Computing important performance metrics, such as tail probabilities, is even more intractable, despite their practical relevance. In this paper, we develop a deep learning approach that accurately and efficiently learns the Laplace transform of high-dimensional RBMs based on the basic adjoint relationship (BAR). Our framework combines a careful design of the loss function, training data sampling procedure, and neural network architecture. We evaluate the proposed method on RBM instances with known ground-truth tail probabilities and demonstrate near-perfect prediction in high-dimensional settings, highlighting its potential as a general tool for analyzing stochastic systems beyond analytically tractable regimes. Our code can be found at https://github.com/zhangz73/NN4MGF.

cs.LG

Optimal Batched Scheduling of Stochastic Processing Networks Using Atomic Action Decomposition

Stochastic processing networks (SPNs) have broad applications in healthcare, transportation, and communication networks. The control of SPN is to dynamically assign servers in batches under uncertainty to optimize long-run performance. This problem is challenging as the policy dimension grows exponentially with the number of servers, making standard reinforcement learning and policy optimization methods intractable at scale. We propose an atomic action decomposition framework that addresses this scalability challenge by breaking joint assignments into sequential single-server assignments. This yields policies with constant dimension, independent of the number of servers. We study two classes of atomic policies, the step-dependent and step-independent atomic policies, and prove that both achieve the same optimal long-run average reward as the original joint policies. These results establish that computing the optimal SPN control can be made scalable without loss of optimality using the atomic framework. Our results offer theoretical justification for the strong empirical success of the atomic framework in large-scale applications reported in previous articles.

eess.SY

Atomic Proximal Policy Optimization for Electric Robo-Taxi Dispatch and Charger Allocation

Pioneering companies such as Waymo have deployed robo-taxi services in several U.S. cities. These robo-taxis are electric vehicles, and their operations require the joint optimization of ride matching, vehicle repositioning, and charging scheduling in a stochastic environment. We model the operations of the ride-hailing system with robo-taxis as a discrete-time, average-reward Markov Decision Process with an infinite horizon. As the fleet size grows, dispatching becomes challenging, as both the system state space and the fleet dispatching action space grow exponentially with the number of vehicles. To address this, we introduce a scalable deep reinforcement learning algorithm, called Atomic Proximal Policy Optimization (Atomic-PPO), that reduces the action space using atomic action decomposition. We evaluate our algorithm using real-world NYC for-hire vehicle trip records and measure its performance by the long-run average reward achieved by the dispatching policy, relative to a fluid-based upper bound. Our experiments demonstrate the superior performance of Atomic-PPO compared to benchmark methods. Furthermore, we conduct extensive numerical experiments to analyze the efficient allocation of charging facilities and assess the impact of vehicle range and charger speed on system performance.

cs.AI

Generative Market Equilibrium Models with Stable Adversarial Learning via Reinforcement

We present a general computational framework for solving continuous-time financial market equilibria under minimal modeling assumptions while incorporating realistic financial frictions, such as trading costs, and supporting multiple interacting agents. Inspired by generative adversarial networks (GANs), our approach employs a novel generative deep reinforcement learning framework with a decoupling feedback system embedded in the adversarial training loop, which we term as the \emph{reinforcement link}. This architecture stabilizes the training dynamics by incorporating feedback from the discriminator. Our theoretically guided feedback mechanism enables the decoupling of the equilibrium system, overcoming challenges that hinder conventional numerical algorithms. Experimentally, our algorithm not only learns but also provides testable predictions on how asset returns and volatilities emerge from the endogenous trading behavior of market participants, where traditional analytical methods fall short. The design of our model is further supported by an approximation guarantee.

q-fin.MF

Dynamical Simulation Model of the Pyro-Process in Cement Clinker Production

This study presents a dynamic simulation model for the pyro-process of clinker production in cement plants. The study aims to construct a simulation model capable of replicating the real-world dynamics of the pyro-process to facilitate research into the improvements of operation, i.e., the development of alternative strategies for reducing CO2 emissions and ensuring clinker quality, production, and lowering fuel consumption. The presented model is an index-1 differential-algebraic equation (DAE) model based on first engineering principles and modular approaches. Using a systematic approach, the model is described on a detailed level that integrates geometric aspects, thermo-physical properties, transport phenomena, stoichiometry and kinetics, mass and energy balances, and algebraic volume and energy conservations. By manually calibrating the model to a steady-state reference, we provide dynamic simulation results that match the expected reference performance and the expected dynamic behavior from the industrial practices.

eess.SY

Numerical Discrete-Time Implementation of Continuous-Time Linear-Quadratic Model Predictive Control

This study presents the design, discretization and implementation of the continuous-time linear-quadratic model predictive control (CT-LMPC). The control model of the CT-LMPC is parameterized as transfer functions with time delays, and they are separated into deterministic and stochastic parts for relevant control and filtering algorithms. We formulate time-delay, finite-horizon CT linear-quadratic optimal control problems (LQ-OCPs) for the CT-LMPC. By assuming piece-wise constant inputs and constraints, we present the numerical discretization of the proposed LQ-OCPs and show how to convert the discrete-time (DT) equivalent into a standard quadratic program. The performance of the CT-LMPC is compared with the conventional DT-LMPC algorithm. Our numerical experiments show that, under fixed tunning parameters, the CT-LMPC shows better closed-loop performance as the sampling time increases than the conventional DT-LMPC.

math.OC

Capacity Allocation and Pricing of High Occupancy Toll Lane Systems with Heterogeneous Travelers

In this article, we study the optimal design of High Occupancy Toll (HOT) lanes. In our setup, the traffic authority determines the road capacity allocation between HOT lanes and ordinary lanes, as well as the toll price charged for travelers who use the HOT lanes but do not meet the high-occupancy eligibility criteria. We build a game-theoretic model to analyze the decisions made by travelers with heterogeneous values of time and carpool disutilities, who choose between paying or forming carpools to take the HOT lanes, or taking the ordinary lanes. Travelers' payoffs depend on the congestion cost of the lane that they take, the payment and the carpool disutilities. We provide a complete characterization of travelers' equilibrium strategies and resulting travel times for any capacity allocation and toll price. We also calibrate our model on the California Interstate highway 880 and compute the optimal capacity allocation and toll design.

cs.MA

Designing High-Occupancy Toll Lanes: A Game-Theoretic Analysis

In this article, we study the optimal design of High Occupancy Toll (HOT) lanes. The traffic authority determines the road capacity allocation between HOT lanes and ordinary lanes, as well as the toll price charged for travelers using HOT lanes who do not meet the high-occupancy eligibility criteria. We develop a game-theoretic model to analyze the decisions of travelers with heterogeneous preference parameters in values of time and carpool disutilities. These travelers choose between paying or forming carpools to use the HOT lanes, or taking the ordinary lanes. Travelers' welfare depends on the congestion cost of the lane they use, the toll payment, and the carpool disutilities. For highways with a single entrance and exit node, we provide a complete characterization of equilibrium strategies and a comparative statics analysis of how the equilibrium vehicle flow and travel time change with HOT capacity and toll price. We then extend the single segment model to highways with multiple entrance and exit nodes. We extend the equilibrium concept and propose various design objectives considering traffic congestion, toll revenue, and social welfare. Using the data collected from the HOT lane of the California Interstate Highway 880 (I-880), we formulate a convex program to estimate the travel demand and approximate the distribution of travelers' preference parameters. We then compute the optimal toll design of five segments for I-880 for achieve each one of the four objectives, and compare the optimal solution with the current toll pricing.

cs.GT

Numerical Discretization Methods for the Discounted Linear Quadratic Control Problem

This study focuses on the numerical discretization methods for the continuous-time discounted linear-quadratic optimal control problem (LQ-OCP) with time delays. By assuming piecewise constant inputs, we formulate the discrete system matrices of the discounted LQ-OCPs into systems of differential equations. Subsequently, we derive the discrete-time equivalent of the discounted LQ-OCP by solving these systems. This paper presents three numerical methods for solving the proposed differential equations systems: the fixed-time-step ordinary differential equation (ODE) method, the step-doubling method, and the matrix exponential method. Our numerical experiment demonstrates that all three methods accurately solve the differential equation systems. Interestingly, the step-doubling method emerges as the fastest among them while maintaining the same level of accuracy as the fixed-time-step ODE method.

math.OC

Numerical Discretization Methods for the Extended Linear Quadratic Control Problem

In this study, we introduce numerical methods for discretizing continuous-time linear-quadratic optimal control problems (LQ-OCPs). The discretization of continuous-time LQ-OCPs is formulated into differential equation systems, and we can obtain the discrete equivalent by solving these systems. We present the ordinary differential equation (ODE), matrix exponential, and a novel step-doubling method for the discretization of LQ-OCPs. Utilizing Euler-Maruyama discretization with a fine step, we reformulate the costs of continuous-time stochastic LQ-OCPs into a quadratic form, and show that the stochastic cost follows the $χ^2$ distribution. In the numerical experiment, we test and compare the proposed numerical methods. The results ensure that the discrete-time LQ-OCP derived using the proposed numerical methods is equivalent to the original problem.

eess.SY

Numerical Discretization Methods for Linear Quadratic Control Problems with Time Delays

This paper presents the numerical discretization methods of the continuous-time linear-quadratic optimal control problems (LQ-OCPs) with time delays. We describe the weight matrices of the LQ-OCPs as differential equations systems, allowing us to derive the discrete equivalent of the continuous-time LQ-OCPs. Three numerical methods are introduced for solving proposed differential equations systems: 1) the ordinary differential equation (ODE) method, 2) the matrix exponential method, and 3) the step-doubling method. We implement a continuous-time model predictive control (CT-MPC) on a simulated cement mill system, and the objective function of the CT-MPC is discretized using the proposed LQ discretization scheme. The closed-loop results indicate that the CT-MPC successfully stabilizes and controls the simulated cement mill system, ensuring the viability and effectiveness of LQ discretization.

eess.SY

Software principles and concepts applied in the implementation of cyber-physical systems for real-time advanced process control

Cyber-physical systems (CPSs) for real-time advanced process control (RT-APC) are a class of control systems using network communication to control industrial processes. In this paper, we use simple examples to describe the software principles and concepts used in the implementation of such systems. The key software principles are 1) shared data in the form of a database, files, or shared memory, 2) timers and threads for concurrent periodic execution of tasks, and 3) network communication between the control system and the process, and communication between the control system and the internet, e.g., the cloud to enable remote monitoring and commands. We show how to implement such systems for Linux operating systems applying the C programming language and we also comment on the implementation using the Python programming language. Finally, we present a complete simulation experiment using a real-time simulator.

eess.SY

Deep Learning Algorithms for Hedging with Frictions

This work studies the deep learning-based numerical algorithms for optimal hedging problems in markets with general convex transaction costs on the trading rates, focusing on their scalability of trading time horizon. Based on the comparison results of the FBSDE solver by Han, Jentzen, and E (2018) and the Deep Hedging algorithm by Buehler, Gonon, Teichmann, and Wood (2019), we propose a Stable Transfer Hedging (ST-Hedging) algorithm, to aggregate the convenience of the leading-order approximation formulas and the accuracy of the deep learning-based algorithms. Our ST-Hedging algorithm achieves the same state-of-the-art performance in short and moderately long time horizon as FBSDE solver and Deep Hedging, and generalize well to long time horizon when previous algorithms become suboptimal. With the transfer learning technique, ST-Hedging drastically reduce the training time, and shows great scalability to high-dimensional settings. This opens up new possibilities in model-based deep learning algorithms in economics, finance, and operational research, which takes advantages of the domain expert knowledge and the accuracy of the learning-based methods.

q-fin.MF

A Survey of Online Auction Mechanism Design Using Deep Learning Approaches

Online auction has been very widespread in the recent years. Platform administrators are working hard to refine their auction mechanisms that will generate high profits while maintaining a fair resource allocation. With the advancement of computing technology and the bottleneck in theoretical frameworks, researchers are shifting gears towards online auction designs using deep learning approaches. In this article, we summarized some common deep learning infrastructures adopted in auction mechanism designs and showed how these architectures are evolving. We also discussed how researchers are tackling with the constraints and concerns in the large and dynamic industrial settings. Finally, we pointed out several currently unresolved issues for future directions.

cs.GT