SearcharxivSearch

arXiv subjects

J. G. Dai

Publications and source records attributed to J. G. Dai.

At least 19 recordsLinked to original sources

Functional Limits of Generalized Jackson Networks in Multi-scale Heavy Traffic

We investigate the functional limits of generalized Jackson networks in a multi-scale heavy traffic regime where stations approach full utilization at distinct, separated rates. Our main result shows that the appropriately scaled queue length processes converge weakly to a limit process whose coordinates are mutually independent. This finding reveals the underlying dynamic mechanism that explains the asymptotic independence previously observed only in stationary distributions. The specific form of the limit process is shown to depend on the initial conditions. In this paper, we consider the matching-rate and the lowest-rate initial conditions. Although the corresponding limit processes have different laws, they have the same product-form exponential limiting distribution on the positive orthant as $t\to\infty$. Moreover, we introduce and analyze a blockwise multi-scale heavy traffic regime. In this regime, the network's stations are partitioned into blocks, where stations in different blocks approach the heavy traffic at different rates, while stations within the same block share a common rate. We obtain the functional limits in this regime as well, showing that the limit process exhibits blockwise independence.

math.PR

Throughput-Optimal Scheduling Algorithms for LLM Inference and AI Agents

As demand for Large Language Models (LLMs) and AI agents grows rapidly, optimizing systems for efficient LLM inference becomes critical. While significant efforts have targeted system-level engineering, little has been explored from a mathematical modeling and queueing perspective. In this paper, we develop the queueing fundamentals for LLM inference. In particular, we study the throughput aspect of LLM inference systems. We prove that a large class of `work-conserving' scheduling algorithms achieve maximum throughput for both individual requests and AI-agent workloads with directed acyclic graph (DAG) and fork-join routing topologies, establishing `work-conserving' as a key design principle for practitioners. Technically, we develop a fluid-limit framework for multi-class batched processing networks under $K$-FCFS scheduling, which may be of independent interest. Evaluations of real-world systems confirm that Orca and Sarathi-Serve are throughput-optimal, reassuring practitioners, while FasterTransformer and vanilla vLLM are not maximally stable and should be used with caution. Our analysis also reveals how constraints such as batch size limits and cyclic routing topologies complicate the throughput picture, pointing to rich open questions at the intersection of queueing theory and LLM system design.

stat.ML

Asymptotic product-form steady-state for generalized Jackson networks in multi-scale heavy traffic

We prove that under a multi-scale heavy traffic condition, the stationary distribution of the scaled queue length vector process in any generalized Jackson network has a product-form limit. Each component in the product form follows an exponential distribution, corresponding to the Brownian approximation of a single station queue. The ``single station'' can be constructed precisely and its parameters have a good intuitive interpretation.

math.PR

Tight matrices and heavy traffic steady state convergence in queueing networks

We are interested to prove that the stationary distribution of a multiclass queueing network converges to the stationary distribution of a semimartingale reflecting Brownian motion (SRBM) in heavy traffic. A key condition for this convergence is that the sequence of the pre-limit stationary distributions under appropriate scaling is tight. In Braverman et al.(2025), a sufficient condition for this tightness is introduced in the term of the reflection matrix $R$ of the SRBM, which is coined for $R$ to be ``tight''. In this paper, we study how we can verify this tightness of $R$ of an SRBM. For a $2$-dimensional SRBM, we give necessary and sufficient conditions for $R$ to be tight, while, for a general dimension, we only give sufficient conditions. We then apply these results to the SRBMs arising from the diffusion approximations of multiclass queueing networks with static buffer priority service disciplines that are studied in Braverman et al.(2025). It is shown that $R$ is always tight for this network with two stations if $R$ is completely-$\sr{S}$. For the case of more than two stations, it is shown that $R$ is tight for reentrant lines with last-buffer-first-service (LBFS) discipline, but it is not always tight for reentrant line with first-buffer-first-service (FBFS) discipline.

math.PR

Asymptotic Product-form Steady-state Distribution for Semimartingale Reflecting Brownian Motion in Multi-scaling Regime

Inspired by Dai et al. [2023], we develop a novel multi-scaling asymptotic regime for semimartingale reflecting Brownian motion (SRBM). In this regime, we establish the steady-state convergence of SRBM to a product-form limit with exponentially distributed components by assuming the P-reflection matrix and a uniform moment bound condition. We further demonstrate that the uniform moment bound condition holds in several subclasses of P-matrices. Our proof approach is rooted in the basic adjoint relationship (BAR) for SRBM proposed by Harrison and Williams [1987a].

math.PR

Asymptotic Product-form Steady-state for Multiclass Queueing Networks with SBP Service Policies in Multi-scale Heavy Traffic

In this work, we study the stationary distribution of the scaled queue length vector process in multiclass queueing networks operating under static buffer priority service policies. We establish that when subjected to a multi-scale heavy traffic condition, the stationary distribution converges to a product-form limit, with each component in the product form following an exponential distribution. A major assumption in proving the desired product-form limit is the uniform moment bound for scaled queue lengths. We prove this assumption holds if the unscaled high-priority queue lengths have uniform moment bound and a certain reflection matrix is a P-matrix.

math.PR

Steady-State Convergence of the Continuous-Time Routing System with General Distributions in Heavy Traffic

This paper examines a continuous-time routing system with general interarrival and service time distributions, operating under the join-the-shortest-queue and power-of-two-choices policies. Under a weaker set of assumptions than those commonly found in the literature, we prove that the scaled steady-state queue length at each station converges weakly to an identical exponential random variable in heavy traffic. Specifically, our results hold under the assumption of the $(2 + δ_0)$th moment for the interarrival and service distributions with some $δ_0 > 0$. The proof leverages the Palm version of the basic adjoint relationship (BAR) as a key technique.

math.PR

Uniform Moment Bounds for Generalized Jackson Networks in Multi-scale Heavy Traffic

We establish uniform moment bounds for steady-state queue lengths of generalized Jackson networks (GJNs) in multi-scale heavy traffic as recently proposed by Dai et al. [2023]. Uniform moment bounds lay the foundation for further analysis of the limit stationary distribution. Our result can be used to verify the crucial moment state space collapse (SSC) assumption in Dai et al. [2023] to establish a product-form limit of GJN in the multi-scale heavy traffic regime. Our proof critically utilizes the Palm version of the basic adjoint relationship (BAR) as developed in Braverman et al. [2023].

math.PR

The BAR approach for multiclass queueing networks with SBP service policies

The basic adjoint relationship (BAR) approach is an analysis technique based on the stationary equation of a Markov process. This approach was introduced to study heavy-traffic, steady-state convergence of generalized Jackson networks in which each service station has a single job class. We extend it to multiclass queueing networks operating under static-buffer-priority (SBP) service disciplines. Our extension makes a connection with Palm distributions that allows one to attack a difficulty arising from queue-length truncation, which appears to be unavoidable in the multiclass setting. For multiclass queueing networks operating under SBP service disciplines, our BAR approach provides an alternative to the "interchange of limits" approach that has dominated the literature in the last twenty years. The BAR approach can produce sharp results and allows one to establish steady-state convergence under three additional conditions: stability, state space collapse (SSC) and a certain matrix being "tight." These three conditions do not appear to depend on the interarrival and service-time distributions beyond their means, and their verification can be studied as three separate modules. In particular, they can be studied in a simpler, continuous-time Markov chain setting when all distributions are exponential. As an example, these three conditions are shown to hold in reentrant lines operating under last-buffer-first-serve discipline. In a two-station, five-class reentrant line, under the heavy-traffic condition, the tight-matrix condition implies both the stability condition and the SSC condition. Whether such a relationship holds generally is an open problem.

math.PR

High order steady-state diffusion approximations

We derive and analyze new diffusion approximations of stationary distributions of Markov chains that are based on second- and higher-order terms in the expansion of the Markov chain generator. Our approximations achieve a higher degree of accuracy compared to diffusion approximations widely used for the past fifty years, while retaining a similar computational complexity. To support our approximations, we present a combination of theoretical and numerical results across three different models. Our approximations are derived recursively through Stein/Poisson equations, and the theoretical results are proved using Stein's method.

math.PR

Queueing Network Controls via Deep Reinforcement Learning

Novel advanced policy gradient (APG) methods, such as Trust Region policy optimization and Proximal policy optimization (PPO), have become the dominant reinforcement learning algorithms because of their ease of implementation and good practical performance. A conventional setup for notoriously difficult queueing network control problems is a Markov decision problem (MDP) that has three features: infinite state space, unbounded costs, and long-run average cost objective. We extend the theoretical framework of these APG methods for such MDP problems. The resulting PPO algorithm is tested on a parallel-server system and large-size multiclass queueing networks. The algorithm consistently generates control policies that outperform state-of-art heuristics in literature in a variety of load conditions from light to heavy traffic. These policies are demonstrated to be near-optimal when the optimal policy can be computed. A key to the successes of our PPO algorithm is the use of three variance reduction techniques in estimating the relative value function via sampling. First, we use a discounted relative value function as an approximation of the relative value function. Second, we propose regenerative simulation to estimate the discounted relative value function. Finally, we incorporate the approximating martingale-process method into the regenerative estimator.

math.OC

Refined Policy Improvement Bounds for MDPs

The policy improvement bound on the difference of the discounted returns plays a crucial role in the theoretical justification of the trust-region policy optimization (TRPO) algorithm. The existing bound leads to a degenerate bound when the discount factor approaches one, making the applicability of TRPO and related algorithms questionable when the discount factor is close to one. We refine the results in \cite{Schulman2015, Achiam2017} and propose a novel bound that is "continuous" in the discount factor. In particular, our bound is applicable for MDPs with the long-run average rewards as well.

cs.LG

A High-fidelity, Machine-learning Enhanced Queueing Network Simulation Model for Hospital Ultrasound Operations

We collaborate with a large teaching hospital in Shenzhen, China and build a high-fidelity simulation model for its ultrasound center to predict key performance metrics, including the distributions of queue length, waiting time and sojourn time, with high accuracy. The key challenge to build an accurate simulation model is to understanding the complicated patient routing at the ultrasound center. To address the issue, we propose a novel two-level routing component to the queueing network model. We apply machine learning tools to calibrate the key components of the queueing model from data with enhanced accuracy.

cs.LG

Scalable Deep Reinforcement Learning for Ride-Hailing

Ride-hailing services, such as Didi Chuxing, Lyft, and Uber, arrange thousands of cars to meet ride requests throughout the day. We consider a Markov decision process (MDP) model of a ride-hailing service system, framing it as a reinforcement learning (RL) problem. The simultaneous control of many agents (cars) presents a challenge for the MDP optimization because the action space grows exponentially with the number of cars. We propose a special decomposition for the MDP actions by sequentially assigning tasks to the drivers. The new actions structure resolves the scalability problem and enables the use of deep RL algorithms for control policy optimization. We demonstrate the benefit of our proposed decomposition with a numerical experiment based on real data from Didi Chuxing.

math.OC

Empty-car routing in ridesharing systems

This paper considers a closed queueing network model of ridesharing systems such as Didi Chuxing, Lyft, and Uber. We focus on empty-car routing, a mechanism by which we control car flow in the network to optimize system-wide utility functions, e.g. the availability of empty cars when a passenger arrives. We establish both process-level and steady-state convergence of the queueing network to a fluid limit in a large market regime where demand for rides and supply of cars tend to infinity, and use this limit to study a fluid-based optimization problem. We prove that the optimal network utility obtained from the fluid-based optimization is an upper bound on the utility in the finite car system for any routing policy, both static and dynamic, under which the closed queueing network has a stationary distribution. This upper bound is achieved asymptotically under the fluid-based optimal routing policy. Simulation results with real-world data released by Didi Chuxing demonstrate the benefit of using the fluid-based optimal routing policy compared to various other policies.

math.PR

Heavy traffic approximation for the stationary distribution of a generalized Jackson network: the BAR approach

In the seminal paper of Gamarnik and Zeevi (2006), the authors justify the steady-state diffusion approximation of a generalized Jackson network (GJN) in heavy traffic. Their approach involves the so-called limit interchange argument, which has since become a popular tool employed by many others who study diffusion approximations. In this paper we illustrate a novel approach by using it to justify the steady-state approximation of a GJN in heavy traffic. Our approach involves working directly with the basic adjoint relationship (BAR), an integral equation that characterizes the stationary distribution of a Markov process. As we will show, the BAR approach is a more natural choice than the limit interchange approach for justifying steady-state approximations, and can potentially be applied to the study of other stochastic processing networks such as multiclass queueing networks.

math.PR

Stein's method for steady-state diffusion approximations: an introduction through the Erlang-A and Erlang-C models

This paper provides an introduction to the Stein method framework in the context of steady-state diffusion approximations. The framework consists of three components: the Poisson equation and gradient bounds, generator coupling, and moment bounds. Working in the setting of the Erlang-A and Erlang-C models, we prove that both Wasserstein and Kolmogorov distances between the stationary distribution of a normalized customer count process, and that of an appropriately defined diffusion process decrease at a rate of $1/\sqrt{R}$, where $R$ is the offered load. Futhermore, these error bounds are \emph{universal}, valid in any load condition from lightly loaded to heavily loaded.

math.PR

High order steady-state diffusion approximation of the Erlang-C system

In this paper we introduce a new diffusion approximation for the steady-state customer count of the Erlang-C system. Unlike previous diffusion approximations, which use the steady-state distribution of a diffusion process with a constant diffusion coefficient, our approximation uses the steady-state distribution of a diffusion process with a \textit{state-dependent} diffusion coefficient. We show, both analytically and numerically, that our new approximation is an order of magnitude better than its counterpart. To obtain the analytical results, we use Stein's to show that a variant of the Wasserstein distance between the normalized customer count distribution and our approximation vanishes at a rate of $1/R$, where $R$ is the offered load to the system. In contrast, the previous approximation only achieved a rate of $1/R$. We hope our results motivate others to consider diffusion approximations with state-dependent diffusion coefficients.

math.PR