Searcharxiv⌕ Search

arXiv subjects

Sanne van Kempen

Publications and source records attributed to Sanne van Kempen.

4 recordsLinked to original sources

Learning Adaptive SED for heterogeneous load balancing

We study a two-server load balancing system with heterogeneous service rates that are a priori unknown to the dispatcher. The goal is to route customers according to the Shortest--Expected--Delay (SED) policy, but this requires knowledge of the service rates. Empirical policies that route based on estimates perform poorly: due to estimation error, the empirical policy disagrees with the oracle on an infinite region of the state space. We propose an online learning algorithm that converges to SED while learning the service rates. The algorithm carefully balances empirical SED routing with forced exploration phases that guarantee sufficient sampling of both servers. We prove that our algorithm achieves finite regret; this differs from classical Multi-Armed Bandit settings where regret typically grows logarithmically in time. Finally, numerical experiments demonstrate the performance of our algorithm and highlight the regimes in which forced exploration is especially beneficial.

cs.LG↗

Stability of skill-based queues with deterministic waiting time thresholds

We consider skill-based queues where service is First--Come--First--Served, but some compatibility lines are only available once a customer's waiting time exceeds a threshold. For this model, we establish necessary and sufficient stability conditions. We find that the thresholds do not affect stability; an intuitive result that has not been proved in the literature up to this point. While necessity follows from a coupling argument, the sufficiency proof is more challenging. It follows by formulating the First--in--Line waiting time process as a Piecewise Deterministic Markov Process (PDMP) and constructing a Lyapunov function that accounts for the threshold structure. We show that this Lyapunov function satisfies the boundary condition of the PDMP framework, and that its generator has negative drift when waiting times are large. Finally, we show that bounded waiting-time sets are petite, so that the Foster--Lyapunov criterion applies.

math.PR↗

Demonstration of effective UCB-based routing in skill-based queues on real-world data

This paper is about optimally controlling skill-based queueing systems such as data centers, cloud computing networks, and service systems. By means of a case study using a real-world data set, we investigate the practical implementation of a recently developed reinforcement learning algorithm for optimal customer routing. Our experiments show that the algorithm efficiently learns and adapts to changing environments and outperforms static benchmark policies, indicating its potential for live implementation. We also augment the real-world applicability of this algorithm by introducing a new heuristic routing rule to reduce delays. Moreover, we show that the algorithm can optimize for multiple objectives: next to payoff maximization, secondary objectives such as server load fairness and customer waiting time reduction can be incorporated. Tuning parameters are used for balancing inherent performance trade--offs. Lastly, we investigate the sensitivity to estimation errors and parameter tuning, providing valuable insights for implementing adaptive routing algorithms in complex real-world queueing systems.

cs.LG↗

Learning payoffs while routing in skill-based queues

Motivated by applications in service systems, we consider queueing systems where each customer must be handled by a server with the right skill set. We focus on optimizing the routing of customers to servers in order to maximize the total payoff of customer--server matches. In addition, customer--server dependent payoff parameters are assumed to be unknown a priori. We construct a machine learning algorithm that adaptively learns the payoff parameters while maximizing the total payoff and prove that it achieves polylogarithmic regret. Moreover, we show that the algorithm is asymptotically optimal up to logarithmic terms by deriving a regret lower bound. The algorithm leverages the basic feasible solutions of a static linear program as the action space. The regret analysis overcomes the complex interplay between queueing and learning by analyzing the convergence of the queue length process to its stationary behavior. We also demonstrate the performance of the algorithm numerically, and have included an experiment with time-varying parameters highlighting the potential of the algorithm in non-static environments.

cs.LG↗