SearcharxivSearch

arXiv subjects

Yanlin Qu

Publications and source records attributed to Yanlin Qu.

8 recordsLinked to original sources

A Harris recurrent continuous-time Markov process without wide-sense regenerative structure

While Harris recurrent Markov chains (in discrete time) automatically exhibit wide-sense regenerative structure, we construct a Harris recurrent Markov process (in continuous time) that is not wide-sense regenerative, thereby giving a negative answer to the open problem first raised in the 1990s and later posed in Glynn(2011). The counterexample exhibits the following rigidity property: every almost surely finite random time that is independent of the state observed at that time must be almost surely constant. A Cantor set linearly independent over the rationals plays a key role in the construction, turning calendar time into an algebraic record of the path already traversed.

math.PR

A Broader View of Thompson Sampling

Thompson Sampling is one of the most widely used and studied bandit algorithms, known for its simple structure, low regret performance, and solid theoretical guarantees. Yet, in stark contrast to most other families of bandit algorithms, the exact mechanism through which posterior sampling (as introduced by Thompson) is able to "properly" balance exploration and exploitation, remains a mystery. In this paper, we show that the core insight to address this question stems from recasting Thompson Sampling as an online optimization algorithm. To distill this, we introduce a suitable time invariant notion of regret that leads to a stationarized bandit problem, and a stationary Bellman-optimal policy. We then show that Thompson Sampling admits an online optimization form that mimics the structure of the aforementioned Bellman-optimal policy, where "greediness" is regularized by a measure of residual uncertainty. This new lens of online optimization allows both a better understanding of Thompson Sampling dynamics, as well as a principled manner for policy improvement that mimics the Bellman-optimal benchmark.

cs.LG

Deep Learning for Markov Chains: Lyapunov Functions, Poisson's Equation, and Stationary Distributions

Lyapunov functions are fundamental to establishing the stability of Markovian models, yet their construction typically demands substantial creativity and analytical effort. In this paper, we show that deep learning can automate this process by training neural networks to satisfy integral equations derived from first-transition analysis. Beyond stability analysis, our approach can be adapted to solve Poisson's equation and estimate stationary distributions. While neural networks are inherently function approximators on compact domains, it turns out that our approach remains effective when applied to Markov chains on non-compact state spaces. We demonstrate the effectiveness of this methodology through several examples from queueing theory and beyond.

cs.LG

Rubik's Cube Scrambling Requires at Least 26 Random Moves

Scrambling the standard 3x3x3 Rubik's Cube corresponds to a random walk on a group containing approximately 43 quintillion elements. Viewing the random walk as a Markov chain, its mixing time determines the number of random moves required to sufficiently scramble a solved cube. With the aid of a supercomputer, we show that the mixing time is at least 26, providing the first non-trivial bound.

math.PR

Double Distributionally Robust Bid Shading for First Price Auctions

Bid shading has become a standard practice in the digital advertising industry, in which most auctions for advertising (ad) opportunities are now of first price type. Given an ad opportunity, performing bid shading requires estimating not only the value of the opportunity but also the distribution of the highest bid from competitors (i.e. the competitive landscape). Since these two estimates tend to be very noisy in practice, first-price auction participants need a bid shading policy that is robust against relatively significant estimation errors. In this work, we provide a max-min formulation in which we maximize the surplus against an adversary that chooses a distribution both for the value and the competitive landscape, each from a Kullback-Leibler-based ambiguity set. As we demonstrate, the two ambiguity sets are essential to adjusting the shape of the bid-shading policy in a principled way so as to effectively cope with uncertainty. Our distributionally robust bid shading policy is efficient to compute and systematically outperforms its non-robust counterpart on real datasets provided by Yahoo DSP.

cs.GT

Deep Learning for Computing Convergence Rates of Markov Chains

Convergence rate analysis for general state-space Markov chains is fundamentally important in areas such as Markov chain Monte Carlo and algorithmic analysis (for computing explicit convergence bounds). This problem, however, is notoriously difficult because traditional analytical methods often do not generate practically useful convergence bounds for realistic Markov chains. We propose the Deep Contractive Drift Calculator (DCDC), the first general-purpose sample-based algorithm for bounding the convergence of Markov chains to stationarity in Wasserstein distance. The DCDC has two components. First, inspired by the new convergence analysis framework in Qu, Blanchet and Glynn (2023), we introduce the Contractive Drift Equation (CDE), the solution of which leads to an explicit convergence bound. Second, we develop an efficient neural-network-based CDE solver. Equipped with these two components, DCDC solves the CDE and converts the solution into a convergence bound. We analyze the sample complexity of the algorithm and further demonstrate the effectiveness of the DCDC by generating convergence bounds for realistic Markov chains arising from stochastic processing networks as well as constant step-size stochastic optimization.

cs.LG

Computable Bounds on Convergence of Markov Chains in Wasserstein Distance via Contractive Drift

We introduce a unified framework to estimate the convergence of Markov chains to equilibrium in Wasserstein distance. The framework can provide convergence bounds with rates ranging from polynomial to exponential, all derived from a contractive drift condition that integrates not only contraction and drift but also coupling and metric design. The resulting bounds are computable, as they contain simple constants, one-step transition expectations, but no equilibrium-related quantities. We introduce the large M technique and the boundary removal technique to enhance the applicability of the framework, which is further enhanced by deep learning in Qu, Blanchet and Glynn (2024). We apply the framework to non-contractive or even expansive Markov chains arising from queueing theory, stochastic optimization, and Markov chain Monte Carlo.

math.PR

Closed-form Solutions of Relativistic Black-Scholes Equations

Drawing insights from the triumph of relativistic over classical mechanics when velocities approach the speed of light, we explore a similar improvement to the seminal Black-Scholes (Black and Scholes (1973)) option pricing formula by considering a relativist version of it, and then finding a respective solution. We show that our solution offers a significant improvement over competing solutions (e.g., Romero and Zubieta-Martinez (2016)), and obtain a new closed-form option pricing formula, containing the speed limit of information transfer c as a new parameter. The new formula is rigorously shown to converge to the Black-Scholes formula as c goes to infinity. When c is finite, the new formula can flatten the standard volatility smile which is more consistent with empirical observations. In addition, an alternative family of distributions for stock prices arises from our new formula, which offer a better fit, are shown to converge to lognormal, and help to better explain the volatility skew.

q-fin.MF