SearcharxivSearch

arXiv subjects

Xian Yu

Publications and source records attributed to Xian Yu.

17 recordsLinked to original sources

Contextual Stochastic Optimization with Decision-Dependent Uncertainty via Nonparametric Learning

We study a general decision-dependent contextual stochastic program (DD-CSP) in which uncertainty depends on both exogenous contextual information and endogenous decisions. To learn the potentially complex dependence of uncertainty on decisions and contextual information, we employ several nonparametric regression models, including k nearest neighbors (kNN), classification and regression trees (CART), and ReLU neural networks. To account for estimation errors in predicting the uncertainty, we adopt an empirical residuals-based decision-dependent sample average approximation (ER-DD-SAA) framework, which adds empirical residuals to the point predictions from the learned regression models. For each nonparametric regression model, we develop exact mixed-integer programming (MIP) representations that can be seamlessly embedded within the ER-DD-SAA framework. For two-stage ER-DD-SAA problems with kNN, we further propose a tailored decomposition algorithm, named BD-CG, that combines Bender's decomposition with constraint generation. Under suitable assumptions, we prove that the proposed BD-CG converges to a global optimum within a finite number of iterations. From a statistical perspective, we establish the consistency and asymptotic optimality of ER-DD-SAA with all three nonparametric regression models under mild regularity conditions. Numerical experiments on a newsvendor problem with pricing and a two-stage facility location problem demonstrate that the ER-DD-SAA model with nonparametric learning consistently outperforms a parametric benchmark in out-of-sample performance and the proposed reformulations and algorithm substantially improve computational tractability.

math.OC

Learning to Cut: Reinforcement Learning for Benders Decomposition

Benders decomposition (BD) is a widely used solution approach for solving two-stage stochastic programs arising in real-world decision-making under uncertainty. However, it often suffers from slow convergence as the master problem grows with an increasing number of cuts. In this paper, we propose Reinforcement Learning for BD (RLBD), a framework that adaptively selects cuts using a neural network-based stochastic policy. The policy is trained using a policy gradient method via the REINFORCE algorithm. We evaluate the proposed approach on a two-stage stochastic electric vehicle charging station location problem and compare it with vanilla BD and LearnBD, a supervised learning approach that classifies cuts using a support vector machine. Numerical results demonstrate that RLBD achieves substantial improvements in computational efficiency and exhibits strong generalization to problems with similar structures but varying data inputs and decision variable dimensions.

math.OC

Residuals-based Offline Reinforcement Learning

Offline reinforcement learning (RL) has received increasing attention for learning policies from previously collected data without interaction with the real environment, which is particularly important in high-stakes applications. While a growing body of work has developed offline RL algorithms, these methods often rely on restrictive assumptions about data coverage and suffer from distribution shift. In this paper, we propose a residuals-based offline RL framework for general state and action spaces. Specifically, we define a residuals-based Bellman optimality operator that explicitly incorporates estimation error in learning transition dynamics into policy optimization by leveraging empirical residuals. We show that this Bellman operator is a contraction mapping and identify conditions under which its fixed point is asymptotically optimal and possesses finite-sample guarantees. We further develop a residuals-based offline deep Q-learning (DQN) algorithm. Using a stochastic CartPole environment, we demonstrate the effectiveness of our residuals-based offline DQN algorithm.

cs.LG

Distributionally Robust Optimization for Chemotherapy Scheduling under Asymmetric and Multi-Modal Uncertainty

We consider a real-world chemotherapy scheduling template design problem, where we cluster patient types into groups and find a representative time-slot duration for each group to accommodate all patient types assigned to that group, aiming to minimize the total expected idle time and overtime. From Mayo Clinic's real data, most patients' treatment durations are asymmetric (e.g., shorter/longer durations tend to have a longer right/left tail). Motivated by this observation, we consider a distributionally robust optimization (DRO) model under an asymmetric and multi-modal ambiguity set, where the distribution of the random treatment duration is modeled as a mixture of distributions from different patient types. The ambiguity set captures uncertainty in both the mode probabilities, modeled via a variation-distance-based set, and the distributions within each mode, characterized by moment information such as the empirical mean, variance, and semivariance. We reformulate the DRO model as a semi-infinite program, which cannot be solved by off-the-shelf solvers. To overcome this, we derive a closed-form expression for the worst-case expected cost and establish lower and upper bounds that are positively related to the variability of patient types assigned to each group, based on which we develop exact algorithms and highly efficient clustering-based heuristics. The lower and upper bounds on the worst-case cost imply that the optimal cost tends to decrease if we group patient types with similar treatment times. Through numerical experiments based on both synthetic datasets and Mayo Clinic's real data, we illustrate the effectiveness and efficiency of the proposed exact algorithms and heuristics and showcase the benefits of incorporating asymmetric information into the DRO formulation.

math.OC

Reward Redistribution via Gaussian Process Likelihood Estimation

In many practical reinforcement learning tasks, feedback is only provided at the end of a long horizon, leading to sparse and delayed rewards. Existing reward redistribution methods typically assume that per-step rewards are independent, thus overlooking interdependencies among state-action pairs. In this paper, we propose a Gaussian process based Likelihood Reward Redistribution (GP-LRR) framework that addresses this issue by modeling the reward function as a sample from a Gaussian process, which explicitly captures dependencies between state-action pairs through the kernel function. By maximizing the likelihood of the observed episodic return via a leave-one-out strategy that leverages the entire trajectory, our framework inherently introduces uncertainty regularization. Moreover, we show that conventional mean-squared-error (MSE) based reward redistribution arises as a special case of our GP-LRR framework when using a degenerate kernel without observation noise. When integrated with an off-policy algorithm such as Soft Actor-Critic, GP-LRR yields dense and informative reward signals, resulting in superior sample efficiency and policy performance on several MuJoCo benchmarks.

cs.LG

A Gauge Set Framework for Flexible Robustness Design

This paper proposes a unified framework for designing robustness in optimization under uncertainty using gauge sets, convex sets that generalize distance and capture how distributions may deviate from a nominal reference. Representing robustness through a gauge set reweighting formulation brings many classical robustness paradigms under a single convex-analytic perspective. The corresponding dual problem, the upper approximator regularization model, reveals a direct connection between distributional perturbations and objective regularization via polar gauge sets. This framework decouples the design of the nominal distribution, distance metric, and reformulation method, components often entangled in classical approaches, thus enabling modular and composable robustness modeling. We further provide a gauge set algebra toolkit that supports intersection, summation, convex combination, and composition, enabling complex ambiguity structures to be assembled from simpler components. For computational tractability under continuously supported uncertainty, we introduce two general finite-dimensional reformulation methods. The functional parameterization approach guarantees any prescribed gauge-based robustness through flexible selection of function bases, while the envelope representation approach yields exact reformulations under empirical nominal distributions and is asymptotically exact for arbitrary nominal choices. A detailed case study demonstrates how the framework accommodates diverse robustness requirements while admitting multiple tractable reformulations.

math.OC

On the Value of Risk-Averse Multistage Stochastic Programming in Capacity Planning

We consider a risk-averse stochastic capacity planning problem under uncertain demand in each period. Using a scenario tree representation of the uncertainty, we formulate a multistage stochastic integer program to adjust the capacity expansion plan dynamically as more information on the uncertainty is revealed. Specifically, in each stage, a decision maker optimizes capacity acquisition and resource allocation to minimize certain risk measures of maintenance and operational cost. We compare it with a two-stage approach that determines the capacity acquisition for all the periods up front. Using expected conditional risk measures (ECRMs), we derive a tight lower bound and an upper bound for the gaps between the optimal objective values of risk-averse multistage models and their two-stage counterparts. Based on these derived bounds, we present general guidelines on when to solve risk-averse two-stage or multistage models. Furthermore, we propose approximation algorithms to solve the two models more efficiently, which are asymptotically optimal under an expanding market assumption. We conduct numerical studies using randomly generated and real-world instances with diverse sizes, to demonstrate the tightness of the analytical bounds and efficacy of the approximation algorithms. We find that the gaps between risk-averse multistage and two-stage models increase as the variability of the uncertain parameters increases and decrease as the decision maker becomes more risk-averse. Moreover, stagewise-dependent scenario tree attains much higher gaps than stagewise-independent counterpart, while the latter produces tighter analytical bounds.

math.OC

Residuals-Based Contextual Distributionally Robust Optimization with Decision-Dependent Uncertainty: Theoretical Guarantees and Decomposition Algorithm

We consider a residuals-based distributionally robust optimization (DRO) model, where the underlying uncertainty depends on both covariate information and our decisions. We adopt both parametric and nonparametric regression models to learn the latent decision dependency and construct a nominal distribution (thereby ambiguity sets) around the learned model using empirical residuals from the regressions. We formulate the ambiguity set via the Wasserstein distance, where the nominal distribution is both decision- and covariate-dependent. We provide conditions under which desired statistical properties such as asymptotic optimality, rate of convergence, and finite sample guarantees are satisfied. To solve the resulting DRO model, we develop a specialized Bender's decomposition algorithm with nonlinear cuts and prove its finite convergence. Through numerical experiments, we illustrate the effectiveness of our approach and the benefits of integrating decision dependency into a residuals-based DRO framework.

math.OC

Policy Gradient Methods for Risk-Sensitive Distributional Reinforcement Learning with Provable Convergence

Risk-sensitive reinforcement learning (RL) is crucial for maintaining reliable performance in high-stakes applications. While traditional RL methods aim to learn a point estimate of the random cumulative cost, distributional RL (DRL) seeks to estimate the entire distribution of it, which leads to a unified framework for handling different risk measures. However, developing policy gradient methods for risk-sensitive DRL is inherently more complex as it involves finding the gradient of a probability measure. This paper introduces a new policy gradient method for risk-sensitive DRL with general coherent risk measures, where we provide an analytical form of the probability measure's gradient for any distribution. For practical use, we design a categorical distributional policy gradient algorithm (CDPG) that approximates any distribution by a categorical family supported on some fixed points. We further provide a finite-support optimality guarantee and a finite-iteration convergence guarantee under inexact policy evaluation and gradient estimation. Through experiments on stochastic Cliffwalk and CartPole environments, we illustrate the benefits of considering a risk-sensitive setting in DRL.

cs.LG

Distributionally Robust Optimization with Multimodal Decision-Dependent Ambiguity Sets

We consider a two-stage distributionally robust optimization (DRO) model with multimodal uncertainty, where both the mode probabilities and uncertainty distributions could be affected by the first-stage decisions. To address this setting, we propose a generic framework by introducing a $\phi$-divergence based ambiguity set to characterize the decision-dependent mode probabilities and further consider both moment-based and Wasserstein distance-based ambiguity sets to characterize the uncertainty distribution under each mode. We identify two special $\phi$-divergence examples (variation distance and $\chi^2$-distance) and provide specific forms of decision dependence relationships under which we can derive tractable reformulations. Furthermore, we investigate the benefits of considering multimodality in a DRO model compared to a single-modal counterpart through an analytical analysis. Additionally, we develop a separation-based decomposition algorithm to solve the resulting multimodal decision-dependent DRO models with finite convergence and optimality guarantee under certain settings. We provide a detailed computational study over two example problem settings, the facility location problem and shipment planning problem with pricing, to illustrate our results, which demonstrate that omission of multimodality or decision-dependent uncertainties within DRO frameworks result in inadequately performing solutions with worse in-sample and out-of-sample performances under various settings. We further demonstrate the speed-ups obtained by the solution algorithm against the off-the-shelf solver over various instances.

math.OC

Kernel-based Regularized Iterative Learning Control of Repetitive Linear Time-varying Systems

For data-driven iterative learning control (ILC) methods, both the model estimation and controller design problems are converted to parameter estimation problems for some chosen model structures. It is well-known that if the model order is not chosen carefully, models with either large variance or large bias would be resulted, which is one of the obstacles to further improve the modeling and tracking performances of data-driven ILC in practice. An emerging trend in the system identification community to deal with this issue is using regularization instead of the statistical tests, e.g., AIC, BIC, and one of the representatives is the so-called kernel-based regularization method (KRM). In this paper, we integrate KRM into data-driven ILC to handle a class of repetitive linear time-varying systems, and moreover, we show that the proposed method has ultimately bounded tracking error in the iteration domain. The numerical simulation results show that in contrast with the least squares method and some existing data-driven ILC methods, the proposed one can give faster convergence speed, better accuracy and robustness in terms of the tracking performance.

eess.SY

On the Global Convergence of Risk-Averse Natural Policy Gradient Methods with Expected Conditional Risk Measures

Risk-sensitive reinforcement learning (RL) has become a popular tool for controlling the risk of uncertain outcomes and ensuring reliable performance in highly stochastic sequential decision-making problems. While it has been shown that policy gradient methods can find globally optimal policies in the risk-neutral setting, it remains unclear if the risk-averse variants enjoy the same global convergence guarantees. In this paper, we consider a class of dynamic time-consistent risk measures, named Expected Conditional Risk Measures (ECRMs), and derive natural policy gradient (NPG) updates for ECRMs-based RL problems. We provide global optimality and iteration complexity of the proposed risk-averse NPG algorithm with softmax parameterization and entropy regularization under both exact and inexact policy evaluation. Furthermore, we test our risk-averse NPG algorithm on a stochastic Cliffwalk environment to demonstrate the efficacy of our method.

cs.LG

Risk-Averse Reinforcement Learning via Dynamic Time-Consistent Risk Measures

Traditional reinforcement learning (RL) aims to maximize the expected total reward, while the risk of uncertain outcomes needs to be controlled to ensure reliable performance in a risk-averse setting. In this paper, we consider the problem of maximizing dynamic risk of a sequence of rewards in infinite-horizon Markov Decision Processes (MDPs). We adapt the Expected Conditional Risk Measures (ECRMs) to the infinite-horizon risk-averse MDP and prove its time consistency. Using a convex combination of expectation and conditional value-at-risk (CVaR) as a special one-step conditional risk measure, we reformulate the risk-averse MDP as a risk-neutral counterpart with augmented action space and manipulation on the immediate rewards. We further prove that the related Bellman operator is a contraction mapping, which guarantees the convergence of any value-based RL algorithms. Accordingly, we develop a risk-averse deep Q-learning framework, and our numerical studies based on two simple MDPs show that the risk-averse setting can reduce the variance and enhance robustness of the results.

cs.LG

On the Value of Multistage Risk-Averse Stochastic Facility Location With or Without Prioritization

We consider a multiperiod stochastic capacitated facility location problem under uncertain demand and budget in each period. Using a scenario tree representation of the uncertainties, we formulate a multistage stochastic integer program to dynamically locate facilities in each period and compare it with a two-stage approach that determines the facility locations up front. In the multistage model, in each stage, a decision maker optimizes facility locations and recourse flows from open facilities to demand sites, to minimize certain risk measures of the cost associated with current facility location and shipment decisions. When the budget is also uncertain, a popular modeling framework is to prioritize the candidate sites. In the two-stage model, the priority list is decided in advance and fixed through all periods, while in the multistage model, the priority list can change adaptively. In each period, the decision maker follows the priority list to open facilities according to the realized budget, and optimizes recourse flows given the realized demand. Using expected conditional risk measures (ECRMs), we derive tight lower bounds for the gaps between the optimal objective values of risk-averse multistage models and their two-stage counterparts in both settings with and without prioritization. Moreover, we propose two approximation algorithms to efficiently solve risk-averse two-stage and multistage models without prioritization, which are asymptotically optimal under an expanding market assumption. We also design a set of super-valid inequalities for risk-averse two-stage and multistage stochastic programs with prioritization to reduce the computational time. We conduct numerical studies using both randomly generated and real-world instances with diverse sizes, to demonstrate the tightness of the analytical bounds and efficacy of the approximation algorithms and prioritization cuts.

math.OC

Resource Distribution Under Spatiotemporal Uncertainty of Disease Spread: Stochastic versus Robust Approaches

We consider the problem of optimizing locations of distribution centers (DCs) and plans for distributing resources such as test kits and vaccines, under spatiotemporal uncertainties of disease spread and demand for the resources. We aim to balance the operational cost (including costs of deploying facilities, shipping, and storage) and quality of service (reflected by demand coverage), while ensuring equity and fairness of resource distribution across multiple populations. We compare a sample-based stochastic programming (SP) approach with a distributionally robust optimization (DRO) approach using a moment-based ambiguity set. Numerical studies are conducted on instances of distributing COVID-19 vaccines in the United States and test kits, to compare SP and DRO models with a deterministic formulation using estimated demand and with the current resource distribution plans implemented in the US. We demonstrate the results over distinct phases of the pandemic to estimate the cost and speed of resource distribution depending on scale and coverage, and show the ``demand-driven'' properties of the SP and DRO solutions. Our results further indicate that if the worst-case unmet demand is prioritized, then the DRO approach is preferred despite of its higher overall cost. Nevertheless, the SP approach can provide an intermediate plan under budgetary restrictions without significant compromises in demand coverage.

math.OC

An Optimization-and-Simulation Framework for Redesigning University Campus Bus System with Social Distancing

The outbreak of coronavirus disease 2019 (COVID-19) has led to significant challenges for schools, workplaces and communities to return to operations during the pandemic, requiring policymakers to balance individuals' safety and operational efficiency. In this paper, we present our work using mixed-integer programming and simulation for redesigning routes and bus schedules for University of Michigan (UM)'s campus bus system during the COVID-19 pandemic. We propose a hub-and-spoke design and utilize real data of student activities to identify hub locations and bus stops to be used in the new routes. Using the same total number of buses to operate, each new bus route has 50\% or fewer seats being used and takes maximumly 15 minutes, to reduce disease transmission through expiratory aerosol. We sample a variety of scenarios that cover variations of peak demand, social-distancing requirements, and break-down buses, to demonstrate the system resiliency of the new routes and schedules via simulation. The new bus routes are implemented and used by all UM campuses during the academic year 2020-2021, to ensure social distancing and short travel time. Our approach can be generalized to redesign public transit systems with social distancing requirement during the pandemic to reduce passengers' infection risk.

eess.SY

Multistage Distributionally Robust Mixed-Integer Programming with Decision-Dependent Moment-Based Ambiguity Sets

We study multistage distributionally robust mixed-integer programs under endogenous uncertainty, where the probability distribution of stage-wise uncertainty depends on the decisions made in previous stages. We first consider two ambiguity sets defined by decision-dependent bounds on the first and second moments of uncertain parameters and by mean and covariance matrix that exactly match decision-dependent empirical ones, respectively. For both sets, we show that the subproblem in each stage can be recast as a mixed-integer linear program (MILP). Moreover, we extend the general moment-based ambiguity set in (Delage and Ye, 2010) to the multistage decision-dependent setting, and derive mixed-integer semidefinite programming (MISDP) reformulations of stage-wise subproblems. We develop methods for attaining lower and upper bounds of the optimal objective value of the multistage MISDPs, and approximate them using a series of MILPs. We deploy the Stochastic Dual Dynamic integer Programming (SDDiP) method for solving the problem under the three ambiguity sets with risk-neutral or risk-averse objective functions, and conduct numerical studies on multistage facility-location instances having diverse sizes under different parameter and uncertainty settings. Our results show that the SDDiP quickly finds optimal solutions for moderate-sized instances under the first two ambiguity sets, and also finds good approximate bounds for the multistage MISDPs derived under the third ambiguity set. We also demonstrate the efficacy of incorporating decision-dependent distributional ambiguity in multistage decision-making processes.

math.OC