SearcharxivSearch

arXiv subjects

Dacheng Yao

Publications and source records attributed to Dacheng Yao.

11 recordsLinked to original sources

Data-Driven Brownian Reflection Control

We study a data-driven reflection control problem for a Brownian model with unknown drift and volatility. We first propose a learn-then-optimize (LTO) algorithm: it estimates the policy-relevant parameter during exploration, plugs the estimate into the optimality equation, and exploits the resulting policy---achieving an $O(\sqrt{T})$ finite-time expected regret bound. We further propose two algorithms, adaptive-updating (AU) and full-history adaptive-updating (AU-FH), which continuously update the estimator and reflecting level, attaining an improved $O(\log T)$ regret bound. Notably, AU-FH algorithms leverages all historical data, yielding better performance in numerical simulations. Our analysis decomposes regret into exploration, transient, and learning components. Transient regret from nonstationarity is bounded by the time-integrated deviation of the transition semigroup from stationarity evaluated on the holding cost, which can be further bounded via a Foster-Lyapunov inequality for exponential convergence of the controlled reflected Brownian motion (RBM). For learning regret, we establish local regularity properties together with consistency and mean-squared error bounds for the estimator, which control the stationary cost gap between the learned and optimal reflection policies. In addition, we leverage the monotonicity of the moving boundary Skorokhod map to derive moment bounds for the AU algorithms' switching states, via pathwise comparison with fixed-boundary RBMs.

math.OC

A Non-convex Optimization Approach of Searching Algebraic Degree Phase-type Representations for General Phase-type Distributions

For a continuous-time phase-type distribution, starting with its Laplace-Stieltjes transform, we obtain a necessary and sufficient condition for its minimal phase-type representation to have the same order as the algebraic degree of the Laplace-Stieltjes transform. To facilitate finding this minimal representation, we transform this condition equivalently into a quadratic nonconvex optimization problem, which can be effectively addressed using an alternating minimization algorithm. The algorithm convergence is also proved. Moreover, the method we develop for the continuous-time phase-type distributions can be directly used to the discrete-time phase-type distributions after establishing an equivalence between the minimal representation problems for continuous-time and discrete-times phase-type distributions.

math.OC

On Substochastic Inverse Eigenvalue Problems with the Corresponding Eigenvector Constraints

We consider the inverse eigenvalue problem of constructing a substochastic matrix from the given spectrum parameters with the corresponding eigenvector constraints. This substochastic inverse eigenvalue problem (SstIEP) with the specific eigenvector constraints is formulated into a nonconvex optimization problem (NcOP). The solvability for SstIEP with the specific eigenvector constraints is equivalent to identify the attainability of a zero optimal value for the formulated NcOP. When the optimal objective value is zero, the corresponding optimal solution to the formulated NcOP is just the substochastic matrix desired to be constructed. We develop the alternating minimization algorithm to solve the formulated NcOP, and its convergence is established by developing a novel method to obtain the boundedness of the optimal solution. Some numerical experiments are conducted to demonstrate the efficiency of the proposed method.

math.OC

Ergodic inventory control with diffusion demand and general ordering costs

In this work, we consider a continuous-time inventory system where the demand process follows an inventory-dependent diffusion process. The ordering cost of each order depends on the order quantity and is given by a general function, which is not even necessarily continuous and monotone. By applying a lower bound approach together with a comparison theorem, we show the global optimality of an $(s,S)$ policy for this ergodic inventory control problem.

math.OC

Impulse Control with Discontinuous Setup Costs: Discounted Cost Criterion

This paper studies a continuous-review backlogged inventory model considered by Helmes et al. (2015) but with discontinuous quantity-dependent setup cost for each order. In particular, the setup cost is characterized by a two-step function and a higher cost would be charged once the order quantity exceeds a threshold $Q$. Unlike the optimality of $(s,S)$-type policy obtained by Helmes et al. (2015) for continuous setup cost with the discounted cost criterion, we find that, in our model, although some $(s,S)$-type policy is indeed optimal in some cases, the $(s,S)$-type policy can not always be optimal. In particular, we show that there exist cases in which an $(s,S)$ policy is optimal for some initial levels but it is strictly worse than a generalized $(s,\{S(x):x\leq s\})$ policy for the other initial levels. Under $(s,\{S(x):x\leq s\})$ policy, it orders nothing for $x>s$ and orders up to level $S(x)$ for $x\leq s$, where $S(x)$ is a non-constant function of $x$. We further prove the optimality of such $(s,\{S(x):x\leq s\})$ policy in a large subset of admissible policies for those initial levels. Moreover, the optimality is obtained through establishing a more general lower bound theorem which will also be applicable in solving some other optimization problems by the common lower bound approach.

math.OC

Optimal Drift Rate Control and Impulse Control for a Stochastic Inventory/Production System

In this paper, we consider joint drift rate control and impulse control for a stochastic inventory system under long-run average cost criterion. Assuming the inventory level must be nonnegative, we prove that a $\{(0,q^{\star},Q^{\star},S^{\star}),\{μ^{\star}(x): x\in[0, S^{\star}]\}\}$ policy is an optimal joint control policy, where the impulse control follows the control band policy $(0,q^{\star},Q^{\star},S^{\star})$, that brings the inventory level up to $q^{\star}$ once it drops to $0$ and brings it down to $Q^{\star}$ once it rises to $S^{\star}$, and the drift rate only depends on the current inventory level and is given by function $μ^{\star}(x)$ for the inventory level $x\in[0,S^{\star}]$. The optimality of the $\{(0,q^{\star},Q^{\star},S^{\star}),\{μ^{\star}(x): x\in[0,S^{\star}]\}\}$ policy is proven by using a lower bound approach, in which a critical step is to prove the existence and uniqueness of optimal policy parameters. To prove the existence and uniqueness, we develop a novel analytical method to solve a free boundary problem consisting of an ordinary differential equation (ODE) and several free boundary conditions. Furthermore, we find that the optimal drift rate $μ^{\star}(x)$ is firstly increasing and then decreasing as $x$ increases from $0$ to $S^{\star}$ with a turnover point between $Q^{\star}$ and $S^{\star}$.

math.OC

Joint pricing and inventory control for a stochastic inventory system with Brownian motion demand

In this paper, we consider an infinite horizon, continuous-review, stochastic inventory system in which cumulative customers' demand is price-dependent and is modeled as a Brownian motion. Excess demand is backlogged. The revenue is earned by selling products and the costs are incurred by holding/shortage and ordering, the latter consists of a fixed cost and a proportional cost. Our objective is to simultaneously determine a pricing strategy and an inventory control strategy to maximize the expected long-run average profit. Specifically, the pricing strategy provides the price $p_t$ for any time $t\geq0$ and the inventory control strategy characterizes when and how much we need to order. We show that an $(s^*,S^*,p^*)$ policy is optimal and obtain the equations of optimal policy parameters, where $p^*=\{p_t^*:t\geq 0\}$. Furthermore, we find that at each time $t$, the optimal price $p_t^*$ depends on the current inventory level $z$, and it is increasing in $[s^*,z^*]$ and is decreasing in $[z^*,\infty)$, where $z^*$ is a negative level.

math.OC

Optimal Control of a Levy Inventory System: The Optimality of Control Band Policy

We consider an inventory system whose state is modeled by a Lévy process. There are two types of costs--the running costs and the inventory control costs. The running costs (also known as the holding/penalty costs) are incurred continuously at some rate as a function of the inventory state. The inventory control costs, incurred only when interventions of the inventory state are placed, have both a fixed and a variable component. The objective is to minimize the expectation of the infinite horizon discounted costs. We formulate this as a stochastic impulse control problem. In our setting, we obtain analytical results that are of significant implications. Specifically, we establish the existence of the optimal control, and we provide the solution in closed-form. More importantly, we prove the optimality of the simple control band policy. Furthermore, we investigate the transient and the steady-state behavior of the controlled process and the stochastic decomposition property.

math.OC

Optimal Ordering Policy for Inventory Systems with Quantity-Dependent Setup Costs

We consider a continuous-review inventory system in which the setup cost of each order is a general function of the order quantity and the demand process is modeled as a Brownian motion with a positive drift. Assuming the holding and shortage cost to be a convex function of the inventory level, we obtain the optimal ordering policy that minimizes the long-run average cost by a lower bound approach. To tackle some technical issues in the lower bound approach under the quantity-dependent setup cost assumption, we establish a comparison theorem that enables one to prove the global optimality of a policy by examining a tractable subset of admissible policies. Since the smooth pasting technique does not apply to our Brownian inventory model, we also propose a selection procedure for computing the optimal policy parameters when the setup cost is a step function.

math.OC

Optimal Control of Brownian Inventory Models with Convex Inventory Cost: Discounted Cost Case

We consider an inventory system in which inventory level fluctuates as a Brownian motion in the absence of control. The inventory continuously accumulates cost at a rate that is a general convex function of the inventory level, which can be negative when there is a backlog. At any time, the inventory level can be adjusted by a positive or negative amount, which incurs a fixed positive cost and a proportional cost. The challenge is to find an adjustment policy that balances the inventory cost and adjustment cost to minimize the expected total discounted cost. We provide a tutorial on using a three-step lower-bound approach to solving the optimal control problem under a discounted cost criterion. In addition, we prove that a four-parameter control band policy is optimal among all feasible policies. A key step is the constructive proof of the existence of a unique solution to the free boundary problem. The proof leads naturally to an algorithm to compute the four parameters of the optimal control band policy.

math.OC

Optimal Control of Brownian Inventory Models with Convex Holding Cost: Average Cost Case

We consider an inventory system in which inventory level fluctuates as a Brownian motion in the absence of control. The inventory continuously accumulates cost at a rate that is a general convex function of the inventory level, which can be negative when there is a backlog. At any time, the inventory level can be adjusted by a positive or negative amount, which incurs a fixed cost and a proportional cost. The challenge is to find an adjustment policy that balances the holding cost and adjustment cost to minimize the long-run average cost. When both upward and downward fixed costs are positive, our model is an impulse control problem. When both fixed costs are zero, our model is a singular or instantaneous control problem. For the impulse control problem, we prove that a four-parameter control band policy is optimal among all feasible policies. For the singular control problem, we prove that a two-parameter control band policy is optimal. We use a lower-bound approach, widely known as "the verification theorem", to prove the optimality of a control band policy for both the impulse and singular control problems. Our major contribution is to prove the existence of a "smooth" solution to the free boundary problem under some mild assumptions on the holding cost function. The existence proof leads naturally to numerical algorithms to compute the optimal control band parameters. We demonstrate that the lower-bound approach also works for Brownian inventory model in which no inventory backlog is allowed. In a companion paper, we will show how the lower-bound approach can be adapted to study a Brownian inventory model under a discounted cost criterion.

math.OC