SearcharxivSearch

arXiv subjects

Peter Glynn

Publications and source records attributed to Peter Glynn.

At least 19 recordsLinked to original sources

Extending Subsampling to Sequential Stopping

Fixed-width sequential stopping rules terminate a stochastic simulation once an estimated confidence interval reaches a prescribed width. Classical fixed-width theory typically relies on a strongly consistent estimator of the asymptotic variance. This makes the normalized stopping time asymptotically deterministic, allowing fixed-sample-size limit theory to be transferred to the estimator at termination. This mechanism can fail when simulation output has infinite variance or long-range dependence. Although self-normalization and subsampling can yield asymptotically valid confidence intervals at a fixed sample size, the scaling process and the stopping time may retain nondegenerate randomness, so fixed-sample-size quantiles need not provide valid coverage at termination. In this paper, we develop a unified framework based on a joint functional limit theorem for the estimation process and a scaling process. We characterize the asymptotic behavior of both the stopping time and the self-normalized estimator evaluated at termination, thereby obtaining asymptotically valid sequential confidence intervals in classical finite-variance and infinite-variance settings. Moreover, we introduce a sequential subsampling procedure that consistently estimates the distribution relevant at the stopping time without directly estimating nuisance parameters in the limit distribution. The framework is verified for heavy-tailed moving-average processes, stochastic approximation, and an M/G/1 queue with heavy-tailed service times.

math.ST

A Harris recurrent continuous-time Markov process without wide-sense regenerative structure

While Harris recurrent Markov chains (in discrete time) automatically exhibit wide-sense regenerative structure, we construct a Harris recurrent Markov process (in continuous time) that is not wide-sense regenerative, thereby giving a negative answer to the open problem first raised in the 1990s and later posed in Glynn(2011). The counterexample exhibits the following rigidity property: every almost surely finite random time that is independent of the state observed at that time must be almost surely constant. A Cantor set linearly independent over the rationals plays a key role in the construction, turning calendar time into an algebraic record of the path already traversed.

math.PR

Statistical Inference for Stochastic Gradient Descent: Beyond Finite Variance

Stochastic gradient descent (SGD) is foundational to large-scale statistical learning and stochastic optimization. However, in some modern statistical learning problems, stochastic gradients can exhibit infinite-variance behavior. Consequently, classical inference methods for SGD that rely on a finite-variance assumption break down. We develop a model-agnostic methodology for constructing confidence regions from SGD iterates in both the finite- and infinite-variance regimes. We first show that Polyak--Ruppert averaging has an asymptotic directional scale no larger than that of the fastest-rate final iterate, analogous to its lower asymptotic variance in the finite-variance setting. Accordingly, we focus our inference methodology on the Polyak--Ruppert averaged estimator. Specifically, we establish a joint central limit theorem for this estimator and an empirical second-moment normalizer from the same iterates. This joint limit yields a self-normalized statistic in which the leading tail-dependent scaling terms cancel. We then use subsampling to estimate the relevant quantiles, avoiding explicit estimation of nuisance parameters including tail indices, slowly varying functions, or stable-law parameters. The resulting confidence regions are straightforward to implement and asymptotically valid in both the finite- and infinite-variance regimes. Empirical studies show reliable coverage in various settings, supporting the proposed method as a practical tool for uncertainty quantification in stochastic optimization.

stat.ML

Fast Convergence of Policy Regret in Learning Stochastic Optimal Control

Policy learning in modern operations environments faces a fundamental tension between limited operational data and the large, often continuous, state and action spaces over which good decisions must be identified and deployed. We study value-based policy learning in stochastic optimal control: a greedy policy induced by an estimate of the optimal action-value function $Q^*$ is deployed, and its performance is measured by regret. The empirical success of this approach calls for statistical insight into the structures that enable fast regret convergence. We show that, in continuous action spaces, fast policy learning is induced by three geometric structures: a growth exponent $p$, which quantifies how quickly $Q^*$ separates suboptimal actions from its maximizers; a margin-mass exponent $m$, which controls how much deployment mass lies on states with weak growth; and an action-wise regularity exponent $q$, which measures the smoothness of the $Q^*$-estimation error across actions. Given a $n^{-1/2}$-accurate estimator of $Q^*$, we show that the minimax-optimal policy regret convergence rate is \[ \widetilde{\Theta}\left( n^{-\min\left\{\frac{p}{2(p-q)},\frac{m+1}{2m}\right\}} \right), \] up to a logarithmic factor at the boundary between the two regimes. The exponent $q$ is crucial: $q>0$ yields faster-than-$n^{-1/2}$ regret. This regime is natural in operations applications. In particular, we verify $q>0$ under mild regularity conditions in dynamic inventory control and service allocation examples, while the mechanism underlying this fast rate regime extends beyond these settings.

math.OC

A Single-Sample Polylogarithmic Regret Bound for Nonstationary Online Linear Programming

We study nonstationary Online Linear Programming (OLP), where $n$ orders arrive sequentially with reward-resource consumption pairs that form a sequence of independent, but not necessarily identically distributed, random vectors. At the beginning of the planning horizon, the decision-maker is provided with a resource endowment that is sufficient to fulfill a significant portion of the requests. The decision-maker seeks to maximize the expected total reward by making immediate and irrevocable acceptance or rejection decisions for each order, subject to this resource endowment. We focus on the challenging single-sample setting, where only one sample from each of the $n$ distributions is available at the start of the planning horizon. We propose a novel re-solving algorithm that integrates a dynamic programming perspective with the dual-based frameworks traditionally employed in stationary environments. In the large-resource regime, where the resource endowment scales linearly with the number of orders, we prove that our algorithm achieves $O((\log n)^2)$ regret across a broad class of nonstationary distribution sequences. Our results demonstrate that polylogarithmic regret is attainable even under significant environmental shifts and minimal data availability, bridging the gap between stationary OLP and more volatile real-world resource allocation problems.

cs.DS

Statistical Inference in Causal Partial Identification with Smooth Densities

Many causal quantities are only partially identifiable due to the inherent missingness of potential outcomes, and the associated partial identification (PI) sets can be obtained by solving an optimal transport (OT) problem. Covariates often provide additional information about the potential outcomes and thus yield tighter PI sets, which can be obtained via conditional optimal transport (COT). However, COT-based PI set estimators are susceptible to the curse of dimensionality in the covariates and outcomes, which precludes the asymptotic normality and hinders statistical inference. In this paper, we exploit smoothness in the marginal densities of covariates and potential outcomes and develop a wavelet-based primal method for COT with multivariate outcomes and covariates. Moreover, for quadratic cost functions, we establish a stability result for COT and prove asymptotic normality of the proposed estimator. This characterization of the asymptotic distribution enables valid statistical inference for the partial identification set. Empirically, we validate the estimation and inference performance of our approach through numerical experiments in comparison with existing benchmarks.

math.ST

A Two-Layer Framework for Joint Online Configuration Selection and Admission Control

We study online configuration selection with admission control problem, which arises in LLM serving, GPU scheduling, and revenue management. In a planning horizon with $T$ periods, we consider a two-layer framework for the decisions made within each time period. In the first layer, the decision maker selects one of the $K$ configurations (ex. quantization, parallelism, fare class) which induces distribution over the reward-resource pair of the incoming request. In the second layer, the decision maker observes the request and then decides whether to accept it or not. Benchmarking this framework requires care. We introduce a \textbf{switching-aware fluid oracle} that accounts for the value of mixing configurations over time, provably upper-bounding any online policy. We derive a max-min formulation for evaluating the benchmark, and we characterize saddle points of the max-min problem via primal-dual optimality conditions linking equilibrium, feasibility, and complementarity. This guides the design of \textbf{SP-UCB--OLP} algorithm, which solves an optimistic saddle point problem and achieves $\tilde{O}(\sqrt{KT})$ regret.

math.OC

On Recurrence of the Infinite Server Queue

This paper concerns the recurrence structure of the infinite server queue, as viewed through the prism of the maximum dater sequence, namely the time to drain the current work in the system as seen at arrival epochs. Despite the importance of this model in queueing theory, we are aware of no complete analysis of the stability behavior of this model, especially in settings in which either or both the inter arrival and service time distributions have infinite mean. In this paper, we fully develop the analog of the Loynes construction of the stationary version in the context of stationary ergodic inputs, extending earlier work of E.Altman (2005), and then classify the Markov chain when the inputs are independent and identically distributed. This allows us to classify the chain, according to transience, recurrence in the sense of Harris, and positive recurrence in the sense of Harris. We further go on to develop tail asymptotics for the stationary distribution of the maximum dater sequence, when the service times have tails that are asymptotically exponential or Pareto, and we contrast the stability theory for the infinite server queue relative to that for the single server queue.

math.PR

Logarithmic Accuracy in Importance Sampling via Large Deviations

Importance sampling (IS) is a widely used simulation method for estimating rare event probabilities. In IS, the relative variance of an estimator is the most common measure of estimator accuracy, and the focus of existing literature is on constructing an importance measure under which the relative variance of the estimator grows sub-exponentially as the parameter increases. In practice, constructing such an estimator is not easy. In this work, we study the behavior of IS estimators under an importance measure which is not necessarily optimal using large deviations theory. This provides new insights into asymptotic efficiency of IS estimators and the required sample size. Based on the study, we also propose new diagnostics of IS for rare event simulation.

math.ST

Deep Learning for Markov Chains: Lyapunov Functions, Poisson's Equation, and Stationary Distributions

Lyapunov functions are fundamental to establishing the stability of Markovian models, yet their construction typically demands substantial creativity and analytical effort. In this paper, we show that deep learning can automate this process by training neural networks to satisfy integral equations derived from first-transition analysis. Beyond stability analysis, our approach can be adapted to solve Poisson's equation and estimate stationary distributions. While neural networks are inherently function approximators on compact domains, it turns out that our approach remains effective when applied to Markov chains on non-compact state spaces. We demonstrate the effectiveness of this methodology through several examples from queueing theory and beyond.

cs.LG

Causal Partial Identification via Conditional Optimal Transport

We study the estimation of causal estimand involving the joint distribution of treatment and control outcomes for a single unit. In typical causal inference settings, it is impossible to observe both outcomes simultaneously, which places our estimation within the domain of partial identification (PI). Pre-treatment covariates can substantially reduce estimation uncertainty by shrinking the partially identified set. Recent work has shown that covariate-assisted PI sets can be characterized through conditional optimal transport (COT) problems. However, finite-sample estimation of COT poses significant challenges, primarily because the COT functional is discontinuous under the weak topology, rendering the direct plug-in estimator inconsistent. To address this issue, existing literature relies on relaxations or indirect methods involving the estimation of non-parametric nuisance statistics. In this work, we demonstrate the continuity of the COT functional under a stronger topology induced by the adapted Wasserstein distance. Leveraging this result, we propose a direct, consistent, non-parametric estimator for COT value that avoids nuisance parameter estimation. We derive the convergence rate for our estimator and validate its effectiveness through comprehensive simulations, demonstrating its improved performance compared to existing approaches.

stat.ME

Tightening Causal Bounds via Covariate-Aware Optimal Transport

Causal estimands can vary significantly depending on the relationship between outcomes in treatment and control groups, potentially leading to wide partial identification (PI) intervals that impede decision making. Incorporating covariates can substantially tighten these bounds, but requires determining the range of PI over probability models consistent with the joint distributions of observed covariates and outcomes in treatment and control groups. This problem is known to be equivalent to a conditional optimal transport (COT) optimization task, which is more challenging than standard optimal transport (OT) due to the additional conditioning constraints. In this work, we study a tight relaxation of COT that effectively reduces it to standard OT, leveraging its well-established computational and theoretical foundations. Our relaxation incorporates covariate information and ensures narrower PI intervals for any value of the penalty parameter, while becoming asymptotically exact as a penalty increases to infinity. This approach preserves the benefits of covariate adjustment in PI and results in a data-driven estimator for the PI set that is easy to implement using existing OT packages. We analyze the convergence rate of our estimator and demonstrate the effectiveness of our approach through extensive simulations, highlighting its practical use and superior performance compared to existing methods.

stat.ME

Wasserstein Projection Tests: Higher-Order Asymptotics and Connections to Empirical Likelihood

Tests of moment restrictions can be constructed by projecting the empirical distribution onto the set of laws satisfying the null. Wasserstein projection (WP) performs this projection by transporting mass in the sample space, rather than only reweighting observed atoms as in empirical likelihood (EL), and therefore incorporates both the ground geometry and the local variation of the moment function. Existing WP theory provides a first-order null limit but does not quantify finite-sample calibration or distinguish tests with the same limiting local power. We derive stochastic expansions of the rescaled WP statistic under the null and under Pitman-type local alternatives, identifying its order-$n^{-1/2}$ correction with a uniform remainder of order $\widetilde{O}_p(n^{-1})$. The resulting Edgeworth theory gives an order-$\widetilde{O}(n^{-1})$ size error for the plug-in WP test and an explicit order-$n^{-1/2}$ correction to local power. For scalar moments under local location shifts, WP, EL, and Hotelling's $T^2$ have the same first-order power, whereas their second-order ranking is determined by derivative-based curvature for WP and value-based skewness for EL. Extending the expansion by one further order, we construct two Bartlett-type corrections that improve coverage accuracy to $\widetilde{O}(n^{-3/2})$. Finally, we develop a certified localized dual algorithm that controls numerical error in the test decision. A theory-guided smooth fairness experiment using a common asymptotic chi-square reference quantile finds WP power gains over the corresponding EL and $T^2$ tests in the derivative-curvature regime predicted by the comparison theory.

math.ST

An Efficient High-Dimensional Gradient Estimator for Stochastic Differential Equations

Overparameterized stochastic differential equation (SDE) models have achieved remarkable success in various complex environments, such as PDE-constrained optimization, stochastic control and reinforcement learning, financial engineering, and neural SDEs. These models often feature system evolution coefficients that are parameterized by a high-dimensional vector $\theta \in \mathbb{R}^n$, aiming to optimize expectations of the SDE, such as a value function, through stochastic gradient ascent. Consequently, designing efficient gradient estimators for which the computational complexity scales well with $n$ is of significant interest. This paper introduces a novel unbiased stochastic gradient estimator--the generator gradient estimator--for which the computation time remains stable in $n$. In addition to establishing the validity of our methodology for general SDEs with jumps, we also perform numerical experiments that test our estimator in linear-quadratic control problems parameterized by high-dimensional neural networks. The results show a significant improvement in efficiency compared to the widely used pathwise differentiation method: Our estimator achieves near-constant computation times, increasingly outperforms its counterpart as $n$ increases, and does so without compromising estimation variance. These empirical findings highlight the potential of our proposed methodology for optimizing SDEs in contemporary applications.

math.OC

Deep Learning for Computing Convergence Rates of Markov Chains

Convergence rate analysis for general state-space Markov chains is fundamentally important in areas such as Markov chain Monte Carlo and algorithmic analysis (for computing explicit convergence bounds). This problem, however, is notoriously difficult because traditional analytical methods often do not generate practically useful convergence bounds for realistic Markov chains. We propose the Deep Contractive Drift Calculator (DCDC), the first general-purpose sample-based algorithm for bounding the convergence of Markov chains to stationarity in Wasserstein distance. The DCDC has two components. First, inspired by the new convergence analysis framework in Qu, Blanchet and Glynn (2023), we introduce the Contractive Drift Equation (CDE), the solution of which leads to an explicit convergence bound. Second, we develop an efficient neural-network-based CDE solver. Equipped with these two components, DCDC solves the CDE and converts the solution into a convergence bound. We analyze the sample complexity of the algorithm and further demonstrate the effectiveness of the DCDC by generating convergence bounds for realistic Markov chains arising from stochastic processing networks as well as constant step-size stochastic optimization.

cs.LG

Optimal Sample Complexity for Average Reward Markov Decision Processes

We resolve the open question regarding the sample complexity of policy learning for maximizing the long-run average reward associated with a uniformly ergodic Markov decision process (MDP), assuming a generative model. In this context, the existing literature provides a sample complexity upper bound of $\widetilde O(|S||A|t_{\text{mix}}^2 ε^{-2})$ and a lower bound of $Ω(|S||A|t_{\text{mix}} ε^{-2})$. In these expressions, $|S|$ and $|A|$ denote the cardinalities of the state and action spaces respectively, $t_{\text{mix}}$ serves as a uniform upper limit for the total variation mixing times, and $ε$ signifies the error tolerance. Therefore, a notable gap of $t_{\text{mix}}$ still remains to be bridged. Our primary contribution is the development of an estimator for the optimal policy of average reward MDPs with a sample complexity of $\widetilde O(|S||A|t_{\text{mix}}ε^{-2})$. This marks the first algorithm and analysis to reach the literature's lower bound. Our new algorithm draws inspiration from ideas in Li et al. (2020), Jin and Sidford (2021), and Wang et al. (2023). Additionally, we conduct numerical experiments to validate our theoretical findings.

cs.LG

Optimal $δ$-Correct Best-Arm Selection for Heavy-Tailed Distributions

Given a finite set of unknown distributions or arms that can be sampled, we consider the problem of identifying the one with the maximum mean using a $δ$-correct algorithm (an adaptive, sequential algorithm that restricts the probability of error to a specified $δ$) that has minimum sample complexity. Lower bounds for $δ$-correct algorithms are well known. $δ$-correct algorithms that match the lower bound asymptotically as $δ$ reduces to zero have been previously developed when arm distributions are restricted to a single parameter exponential family. In this paper, we first observe a negative result that some restrictions are essential, as otherwise, under a $δ$-correct algorithm, distributions with unbounded support would require an infinite number of samples in expectation. We then propose a $δ$-correct algorithm that matches the lower bound as $δ$ reduces to zero under the mild restriction that a known bound on the expectation of $(1+ε)^{th}$ moment of the underlying random variables exists, for $ε> 0$. We also propose batch processing and identify near-optimal batch sizes to speed up the proposed algorithm substantially. The best-arm problem has many learning applications, including recommendation systems and product selection. It is also a well-studied classic problem in the simulation community.

cs.LG

Optimal Sample Complexity of Reinforcement Learning for Mixing Discounted Markov Decision Processes

We consider the optimal sample complexity theory of tabular reinforcement learning (RL) for maximizing the infinite horizon discounted reward in a Markov decision process (MDP). Optimal worst-case complexity results have been developed for tabular RL problems in this setting, leading to a sample complexity dependence on $γ$ and $ε$ of the form $\tilde Θ((1-γ)^{-3}ε^{-2})$, where $γ$ denotes the discount factor and $ε$ is the solution error tolerance. However, in many applications of interest, the optimal policy (or all policies) induces mixing. We establish that in such settings, the optimal sample complexity dependence is $\tilde Θ(t_{\text{mix}}(1-γ)^{-2}ε^{-2})$, where $t_{\text{mix}}$ is the total variation mixing time. Our analysis is grounded in regeneration-type ideas, which we believe are of independent interest, as they can be used to study RL problems for general state space MDPs.

cs.LG