SearcharxivSearch

arXiv subjects

Itai Gurvich

Publications and source records attributed to Itai Gurvich.

9 recordsLinked to original sources

Goggin's corrected Kalman Filter: Guarantees and Filtering Regimes

In this paper we revisit a non-linear filter for {\em non-Gaussian} noises that was introduced in [1]. Goggin proved that transforming the observations by the score function and then applying the Kalman Filter (KF) to the transformed observations results in an asymptotically optimal filter. In the current paper, we study the convergence rate of Goggin's filter in a pre-limit setting that allows us to study a range of signal-to-noise regimes which includes, as a special case, Goggin's setting. Our guarantees are explicit in the level of observation noise, and unlike most other works in filtering, we do not assume Gaussianity of the noises. Our proofs build on combining simple tools from two separate literature streams. One is a general posterior Cramér-Rao lower bound for filtering. The other is convergence-rate bounds in the Fisher information central limit theorem. Along the way, we also study filtering regimes for linear state-space models, characterizing clearly degenerate regimes -- where trivial filters are nearly optimal -- and a {\em balanced} regime, which is where Goggin's filter has the most value. \footnote{This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible.

cs.IT

A Hierarchical Approach to Robust Stability of Multiclass Queueing Networks

We re-visit the global - relative to control policies - stability of multiclass queueing networks. In these, as is known, it is generally insufficient that the nominal utilization at each server is below 100%. Certain policies, although work conserving, may destabilize a network that satisfies the nominal-load conditions; additional conditions on the primitives are needed for global stability (stability under any work-conserving policy). The global-stability region was fully characterized for two-station networks in [13], but a general framework for networks with more than two stations remains elusive. In this paper, we offer progress on this front by considering a subset of non-idling control policies, namely queue-ratio (QR) policies. These include as special cases all static-priority policies. With this restriction, we are able to introduce a complete framework that applies to networks of any size. Our framework breaks the analysis of robust QR stability (stability under any QR policy) into (i) robust state-space collapse and (ii) robust stability of the Skorohod problem (SP) representing the fluid workload. Sufficient conditions for both are specified in terms of simple optimization problems. We use these optimization problems to prove that the family of QR policies satisfies a weak form of convexity relative to policies. A direct implication of this convexity is that: if the SP is stable for all static-priority policies (the "extreme" QR policies), then it is also stable under any QR policy. While robust QR stability is weaker than global stability, our framework recovers necessary and sufficient conditions for global stability in specific networks.

math.OC

A Low-rank Approximation for MDPs via Moment Coupling

We introduce a framework to approximate a Markov Decision Process that stands on two pillars: state aggregation -- as the algorithmic infrastructure; and central-limit-theorem-type approximations -- as the mathematical underpinning of optimality guarantees. The theory is grounded in recent work Braverman et al (2020} that relates the solution of the Bellman equation to that of a PDE where, in the spirit of the central limit theorem, the transition matrix is reduced to its local first and second moments. Solving the PDE is $\textit{not}$ required by our method. Instead, we construct a "sister" (controlled) Markov chain whose two local transition moments are approximately identical with those of the focal chain. Because of this $\textit{moment matching}$, the original chain and its "sister" are coupled through the PDE, a coupling that facilitates optimality guarantees. Embedded into standard soft aggregation algorithms, moment matching provided a disciplined mechanism to tune the aggregation and disaggregation probabilities. The computational gains arise from the reduction of the effective state space from $N$ to $N^{\frac{1}{2}+ε}$ is as one might intuitively expect from approximations grounded in the central limit theorem.

math.OC

Online Allocation and Pricing: Constant Regret via Bellman Inequalities

We develop a framework for designing simple and efficient policies for a family of online allocation and pricing problems, that includes online packing, budget-constrained probing, dynamic pricing, and online contextual bandits with knapsacks. In each case, we evaluate the performance of our policies in terms of their regret (i.e., additive gap) relative to an offline controller that is endowed with more information than the online controller. Our framework is based on Bellman Inequalities, which decompose the loss of an algorithm into two distinct sources of error: (1) arising from computational tractability issues, and (2) arising from estimation/prediction of random trajectories. Balancing these errors guides the choice of benchmarks, and leads to policies that are both tractable and have strong performance guarantees. In particular, in all our examples, we demonstrate constant-regret policies that only require re-solving an LP in each period, followed by a simple greedy action-selection rule; thus, our policies are practical as well as provably near optimal.

math.OC

Uniformly bounded regret in the multi-secretary problem

In the secretary problem of Cayley (1875) and Moser (1956), $n$ non-negative, independent, random variables with common distribution are sequentially presented to a decision maker who decides when to stop and collect the most recent realization. The goal is to maximize the expected value of the collected element. In the $k$-choice variant, the decision maker is allowed to make $k \leq n$ selections to maximize the expected total value of the selected elements. Assuming that the values are drawn from a known distribution with finite support, we prove that the best regret---the expected gap between the optimal online policy and its offline counterpart in which all $n$ values are made visible at time $0$---is uniformly bounded in the the number of candidates $n$ and the budget $k$. Our proof is constructive: we develop an adaptive Budget-Ratio policy that achieves this performance. The policy selects or skips values depending on where the ratio of the residual budget to the remaining time stands relative to multiple thresholds that correspond to middle points of the distribution. We also prove that being adaptive is crucial: in general, the minimal regret among non-adaptive policies grows like the square root of $n$. The difference is the value of adaptiveness.

math.PR

On the Taylor Expansion of Value Functions

We introduce a framework for approximate dynamic programming that we apply to discrete time chains on $\mathbb{Z}_+^d$ with countable action sets. Our approach is grounded in the approximation of the (controlled) chain's generator by that of another Markov process. In simple terms, our approach stipulates applying a second-order Taylor expansion to the value function to replace the Bellman equation with one in continuous space and time where the transition matrix is reduced to its first and second moments. In some cases, the resulting equation (which we label {\bf TCP}) can be interpreted as corresponding to a Brownian control problem. When tractable, the TCP serves as a useful modeling tool. More generally, the TCP is a starting point for approximation algorithms. We develop bounds on the optimality gap---the sub-optimality introduced by using the control produced by the "Taylored" equation. These bounds can be viewed as a conceptual underpinning, analytical rather than relying on weak convergence arguments, for the good performance of controls derived from Brownian control problems. We prove that, under suitable conditions and for suitably "large" initial states, (i) the optimality gap is smaller than a $1-α$ fraction of the optimal value, where $α\in (0,1)$ is the discount factor, and (ii) the gap can be further expressed as the infinite horizon discounted value with a "lower-order" per period reward. Computationally, our framework leads to an "aggregation" approach with performance guarantees. While the guarantees are grounded in PDE theory, the practical use of this approach requires no knowledge of that theory.

math.OC

Diffusion models and steady-state approximations for exponentially ergodic Markovian queues

Motivated by queues with many servers, we study Brownian steady-state approximations for continuous time Markov chains (CTMCs). Our approximations are based on diffusion models (rather than a diffusion limit) whose steady-state, we prove, approximates that of the Markov chain with notable precision. Strong approximations provide such "limitless" approximations for process dynamics. Our focus here is on steady-state distributions, and the diffusion model that we propose is tractable relative to strong approximations. Within an asymptotic framework, in which a scale parameter $n$ is taken large, a uniform (in the scale parameter) Lyapunov condition imposed on the sequence of diffusion models guarantees that the gap between the steady-state moments of the diffusion and those of the properly centered and scaled CTMCs shrinks at a rate of $\sqrt{n}$. Our proofs build on gradient estimates for solutions of the Poisson equations associated with the (sequence of) diffusion models and on elementary martingale arguments. As a by-product of our analysis, we explore connections between Lyapunov functions for the fluid model, the diffusion model and the CTMC.

math.PR

Scheduling parallel servers in the nondegenerate slowdown diffusion regime: Asymptotic optimality results

We consider the problem of minimizing queue-length costs in a system with heterogenous parallel servers, operating in a many-server heavy-traffic regime with nondegenerate slowdown. This regime is distinct from the well-studied heavy traffic diffusion regimes, namely the (single server) conventional regime and the (many-server) Halfin-Whitt regime. It has the distinguishing property that waiting times and service times are of comparable magnitudes. We establish an asymptotic lower bound on the cost and devise a sequence of policies that asymptotically attain this bound. As in the conventional regime, the asymptotics can be described by means of a Brownian control problem, the solution of which exhibits a state space collapse.

math.PR

On optimality gaps in the Halfin--Whitt regime

We consider optimal control of a multi-class queue in the Halfin--Whitt regime, and revisit the notion of asymptotic optimality and the associated optimality gaps. The existing results in the literature for such systems provide asymptotically optimal controls with optimality gaps of $o(\sqrt{n})$ where $n$ is the system size, for example, the number of servers. We construct a sequence of asymptotically optimal controls where the optimality gap grows logarithmically with the system size. Our analysis relies on a sequence of Brownian control problems, whose refined structure helps us achieve the improved optimality gaps.

math.PR