SearcharxivSearch

arXiv subjects

Linwei Xin

Publications and source records attributed to Linwei Xin.

10 recordsLinked to original sources

DeepStock: Reinforcement Learning with Policy Regularizations for Inventory Management

Deep Reinforcement Learning (DRL) provides a general-purpose methodology for training inventory policies that can leverage big data and compute. However, off-the-shelf implementations of DRL have seen mixed success, often plagued by high sensitivity to the hyperparameters used during training. In this paper, we show that by imposing policy regularizations, grounded in classical inventory concepts such as "Base Stock", we can significantly accelerate hyperparameter tuning and improve the final performance of several DRL methods. We report details from a 100% deployment of DRL with policy regularizations on Alibaba's e-commerce platform, Tmall. We also include extensive synthetic experiments, which show that policy regularizations reshape the narrative on what is the best DRL method for inventory management.

cs.LG

Online Order Fulfillment with Replenishment

In modern e-commerce and service operations, firms must jointly manage inventory replenishment and real-time order fulfillment to maximize profit under demand uncertainty. While each component has been studied extensively in isolation, their interaction remains underexplored. This paper investigates a fundamental operational question: which lever plays a more decisive role in overall system performance, replenishment or fulfillment? We model the system as a one-location online order fulfillment problem with lost sales and stochastic customer arrivals, each offering heterogeneous rewards. Replenishment follows either a base-stock or constant-order policy, while real-time fulfillment decisions are made using online algorithms. Our core performance metric is the expected average profit per replenishment cycle, evaluated across all combinations of these policies and algorithms. Our main theoretical result shows that when the replenishment cycle is long, the cumulative regret of online fulfillment remains of the same order as in a corresponding single-cycle problem, even under repeated replenishment, revealing a form of regret stability. This phenomenon also extends to a multi-location setting. We further develop a regret-based framework that quantitatively compares the value of improving replenishment versus improving fulfillment, and we characterize regimes in which optimizing replenishment yields a larger revenue impact than refining the online fulfillment algorithm (and vice versa). Motivated by examples where myopic algorithms underperform, we introduce a novel look-ahead online algorithm that anticipates future replenishment and demand. Numerical experiments verify that this algorithm outperforms myopic baselines. Overall, our results provide both theoretical and managerial insights into situations where inventory replenishment policies are more influential and vice versa.

math.OC

Hidden Convexity in Queueing Models

We study the joint control of arrival and service rates in queueing systems with the objective of minimizing long-run expected cost minus revenue. Although the objective function is non-convex, first-order methods have been empirically observed to converge to globally optimal solutions. This paper provides a theoretical foundation for this empirical phenomenon by characterizing the optimization landscape and identifying a hidden convexity: the problem admits a convex reformulation after an appropriate change of variables. Leveraging this hidden convexity, we establish the Polyak-Lojasiewicz-Kurdyka (PLK) condition for the original control problem, which excludes spurious local minima and supports global convergence guarantees for first-order methods. Our analysis applies to a broad class of $GI/GI/1$ queueing models, including those with Gamma-distributed interarrival and service times, as well as $GI/M/1$ queues with log-concave interarrival times. As a key ingredient in the proof, we establish a new convexity property of the expected queue length under a square-root transformation of the traffic intensity.

math.OC

Auto-Formulating Dynamic Programming Problems with Large Language Models

Dynamic programming (DP) is a fundamental method in operations research, but formulating DP models has traditionally required expert knowledge of both the problem context and DP techniques. Large Language Models (LLMs) offer the potential to automate this process. However, DP problems pose unique challenges due to their inherently stochastic transitions and the limited availability of training data. These factors make it difficult to directly apply existing LLM-based models or frameworks developed for other optimization problems, such as linear or integer programming. We introduce DP-Bench, the first benchmark covering a wide range of textbook-level DP problems to enable systematic evaluation. We present Dynamic Programming Language Model (DPLM), a 7B-parameter specialized model that achieves performance comparable to state-of-the-art LLMs like OpenAI's o1 and DeepSeek-R1, and surpasses them on hard problems. Central to DPLM's effectiveness is DualReflect, our novel synthetic data generation pipeline, designed to scale up training data from a limited set of initial examples. DualReflect combines forward generation for diversity and backward generation for reliability. Our results reveal a key insight: backward generation is favored in low-data regimes for its strong correctness guarantees, while forward generation, though lacking such guarantees, becomes increasingly valuable at scale for introducing diverse formulations. This trade-off highlights the complementary strengths of both approaches and the importance of combining them.

cs.AI

VC Theory for Inventory Policies

There has been growing interest in applying reinforcement learning (RL) to inventory management, either by optimizing over temporal transitions or by learning directly from full historical demand trajectories. This contrasts sharply with classical data-driven approaches, which first estimate demand distributions from past data and then compute well-structured optimal policies via dynamic programming. This paper considers a hybrid approach that combines trajectory-based RL with policy regularization imposing base-stock and $(s, S) $ structures. We provide generalization guarantees for this combined approach for several well-known classes in a $T$-period dynamic inventory model, using tools from the celebrated Vapnik-Chervonenkis (VC) theory, such as the Pseudo-dimension and Fat-shattering dimension. Our results have implications for regret against the best-in-class policies, and allow for an arbitrary distribution over demand sequences, which makes no assumptions such as independence across time. Surprisingly, we prove that the class of policies defined by $T$ non-stationary base-stock levels exhibits a generalization error that does not grow with $T$, whereas the two-parameter $(s, S)$ policy class has a generalization error growing logarithmically with $T$. Overall, our analysis leverages specific inventory structures within the learning theory framework, and improves sample complexity guarantees even compared to existing results assuming independent demands.

stat.ML

Risk-Based Distributionally Robust Optimal Power Flow With Dynamic Line Rating

In this paper, we propose a risk-based data-driven approach to optimal power flow (DROPF) with dynamic line rating. The risk terms, including penalties for load shedding, wind generation curtailment and line overload, are embedded into the objective function. To hedge against the uncertainties on wind generation data and line rating data, we consider a distributionally robust approach. The ambiguity set is based on second-order moment and Wasserstein distance, which captures the correlations between wind generation outputs and line ratings, and is robust to data perturbation. We show that the proposed DROPF model can be reformulated as a conic program. Considering the relatively large number of constraints involved, an approximation of the proposed DROPF model is suggested, which significantly reduces the computational costs. A Wasserstein distance constrained DROPF and its tractable reformulation are also provided for practical large-scale test systems. Simulation results on the 5-bus, the IEEE 118-bus and the Polish 2736-bus test systems validate the effectiveness of the proposed models.

math.OC

Distributionally robust inventory control when demand is a martingale

Demand forecasting plays an important role in many inventory control problems. To mitigate the potential harms of model misspecification, various forms of distributionally robust optimization have been applied. Although many of these methodologies suffer from the problem of time-inconsistency, the work of Klabjan et al. established a general time-consistent framework for such problems by connecting to the literature on robust Markov decision processes. Motivated by the fact that many forecasting models exhibit special structure, as well as a desire to understand the impact of positing different dependency structures, in this paper we formulate and solve a time-consistent distributionally robust multi-stage newsvendor model which naturally unifies and robustifies several inventory models with forecasting. In particular, many simple models of demand forecasting have the feature that demand evolves as a martingale. We consider a robust variant of such models, in which the sequence of future demands may be any martingale with given mean and support. Under such a model, past realizations of demand are naturally incorporated into the structure of the uncertainty set going forwards. We explicitly compute the minimax optimal policy (and worst-case distribution) in closed form, by combining ideas from convexity, probability, and dynamic programming. We prove that at optimality the worst-case demand distribution corresponds to the setting in which inventory may become obsolete, a scenario of practical interest. To gain further insight, we prove weak convergence (as the time horizon grows large) to a simple and intuitive process. We also compare to the analogous setting in which demand is independent across periods (analyzed previously by Shapiro), and identify interesting differences between these models, in the spirit of the price of correlations studied by Agrawal et al.

math.PR

Asymptotic optimality of Tailored Base-Surge policies in dual-sourcing inventory systems

Dual-sourcing inventory systems, in which one supplier is faster (i.e. express) and more costly, while the other is slower (i.e. regular) and cheaper, arise naturally in many real-world supply chains. These systems are notoriously difficult to optimize due to the complex structure of the optimal solution and the curse of dimensionality, having resisted solution for over 40 years. Recently, so-called Tailored Base-Surge (TBS) policies have been proposed as a heuristic for the dual-sourcing problem. Under such a policy, a constant order is placed at the regular source in each period, while the order placed at the express source follows a simple order-up-to rule. Numerical experiments by several authors have suggested that such policies perform well as the lead time difference between the two sources grows large, which is exactly the setting in which the curse of dimensionality leads to the problem becoming intractable. However, providing a theoretical foundation for this phenomenon has remained a major open problem. In this paper, we provide such a theoretical foundation by proving that a simple TBS policy is indeed asymptotically optimal as the lead time of the regular source grows large, with the lead time of the express source held fixed. Our main proof technique combines novel convexity and lower-bounding arguments, an explicit implementation of the vanishing discount factor approach to analyzing infinite-horizon Markov decision processes, and ideas from the theory of random walks and queues, significantly extending the methodology and applicability of a novel framework for analyzing inventory models with large lead times recently introduced by Goldberg and co-authors in the context of lost-sales models with positive lead times.

math.PR

Optimality gap of constant-order policies decays exponentially in the lead time for lost sales models

Inventory models with lost sales and large lead times have traditionally been considered intractable due to the curse of dimensionality. Recently, Goldberg and co-authors laid the foundations for a new approach to solving these models, by proving that as the lead time grows large, a simple constant-order policy is asymptotically optimal. However, the bounds proven there require the lead time to be very large before the constant-order policy becomes effective, in contrast to the good numerical performance demonstrated by Zipkin even for small lead time values. In this work, we prove that for the infinite-horizon variant of the same lost sales problem, the optimality gap of the same constant-order policy actually converges \emph{exponentially fast} to zero, with the optimality gap decaying to zero at least as fast as the exponential rate of convergence of the expected waiting time in a related single-server queue to its steady-state value. We also derive simple and explicit bounds for the optimality gap, and demonstrate good numerical performance across a wide range of parameter values for the special case of exponentially distributed demand. Our main proof technique combines convexity arguments with ideas from queueing theory.

math.PR

Time (in)consistency of multistage distributionally robust inventory models with moment constraints

Recently, there has been a growing interest in developing inventory control policies which are robust to model misspecification. One approach is to posit that nature selects a worst-case distribution for any stochastic primitives from some pre-specified family. Several communities have observed that a subtle phenomena known as time inconsistency can arise in this framework. In particular, it becomes possible that a policy which is optimal at time zero may not be optimal for the associated optimization problem in which the decision-maker recomputes her policy at each point in time, which has implications for implementability. If there exists a policy which is optimal for both formulations, we say that the policy is time consistent, and the problem is weakly time consistent. If every optimal policy is time consistent, we say that the problem is strongly time consistent. We study these phenomena in the context of managing an inventory over time, when only the mean, variance, and support are known for the demand at each stage. We provide several illustrative examples showing that here the question of time consistency can be quite subtle, and complement these observations by providing simple sufficient conditions for weak and strong time consistency. Interestingly, our results show that time consistency may hold even when rectangularity does not. Although a similar phenomena was previously identified by Shapiro for the setting in which only the mean and support of the demand are known, there the problem was always weakly time consistent, with both formulations having the same optimal value. Here our model is rich enough to exhibit a variety of interesting behaviors, including lack of weak time consistency, strong time consistency even when both formulations have different optimal values, and non-existence of even a single optimal base-stock policy under the static formulation.

math.OC