SearcharxivSearch

arXiv subjects

Yishun Luo

Publications and source records attributed to Yishun Luo.

2 recordsLinked to original sources

Load Balancing Policies in Heterogeneous Systems: Non-Monotone Stability and Heavy-Traffic Optimality

We consider a discrete-time queueing system with $n$ heterogeneous parallel single-server queues. Jobs arrive at a central dispatcher and must be assigned immediately to one of the queues. We develop a unified framework for a broad family of load-balancing policies, including Join the Shortest Queue (JSQ), Join the Shortest Expected Delay (JSED), and Power-of-$d$ Choices (Po$d$). In this framework, the dispatcher updates queue-length information periodically, possibly at arbitrarily long intervals, and dispatches jobs based on the sampled permutation of scaled queue lengths and the servers' service rates. Leveraging this structure, we derive a closed-form, easily verifiable sufficient condition for stability. We further show that, for general policies, stability above the induced threshold need not be monotone in the arrival rate, and we obtain an exact characterization under a persistent bottleneck dominance condition. When the stability condition holds strictly, we prove state-space collapse and heavy-traffic delay optimality. We also show that the steady-state queue-length vector converges in distribution to a deterministic vector scaled by an exponential random variable in heavy traffic. Methodologically, we extend Lyapunov-drift and transform techniques to a cycle-based analysis with multi-step updates. Our results connect the policy-induced dispatch fractions and sampled permutations to stability, delay, and distributional performance, providing guidance for designing scalable load-balancing schemes with limited queue-length information.

cs.PF

Heavy-traffic Optimality of Skip-the-Longest-Queues in Heterogeneous Service Systems

We consider a discrete-time parallel service system consisting of $n$ heterogeneous single server queues with infinite capacity. Jobs arrive to the system as an i.i.d. process with rate proportional to $n$, and must be immediately dispatched in the time slot that they arrive. The dispatcher is assumed to be able to exchange messages with the servers to obtain their queue lengths and make dispatching decisions, introducing an undesirable communication overhead. In this setting, we propose a ultra-low communication overhead load balancing policy dubbed $k$-Skip-the-$d$-Longest-Queues ($k$-SLQ-$d$), where queue lengths are only observed every $k(n-d)$ time slots and, between observations, incoming jobs are sent to a queue that is not one of the $d$ longest ones at the time that the queues were last observed. For this policy, we establish conditions on $d$ for it to be throughput optimal and we show that, under that condition, it is asymptotically delay-optimal in heavy-traffic for arbitrarily low communication overheads (i.e., for arbitrarily large $k$).

cs.PF