SearcharxivSearch

arXiv subjects

Dmitry B. Rokhlin

Publications and source records attributed to Dmitry B. Rokhlin.

At least 19 recordsLinked to original sources

Dynamic Regret for Online Regression in RKHS via Discounted VAW and Subspace Approximation

We study online regression with the square loss in a reproducing kernel Hilbert space under a dynamic regret criterion. The learner is compared with a time-varying comparator sequence, and the bounds depend on its path length in the RKHS norm. The proposed method transfers the finite-dimensional discounted Vovk--Azoury--Warmuth approach of Jacobsen \& Cutkosky (2024) to the RKHS setting by means of finite-dimensional subspace approximations. For a fixed subspace, we run a VAW-based ensemble of discounted VAW forecasters over a geometric grid of discount factors. The additional approximation error is controlled by the uniform projection error of kernel sections. We then introduce a general orthogonal truncation method: starting from a feature expansion of the kernel, we construct the associated RKHS by introducing an inner product that makes the feature functions orthonormal, and then use the spans of the first basis functions as finite-dimensional approximation spaces. The resulting subspace reduction is applied to several approximation schemes. Explicit feature expansions yield fast-regime bounds for Gaussian and analytic dot-product kernels. Mercer truncations provide a spectral approximation method and lead to dynamic regret bounds in fast and slow regimes, depending on the eigenvalue decay. Finally, we study subspaces spanned by kernel sections and apply this construction to Matérn kernels.

cs.LG

A hierarchical Vovk-Azoury-Warmuth forecaster with discounting for online regression in RKHS

We study the problem of online regression with the unconstrained quadratic loss against a time-varying sequence of functions from a Reproducing Kernel Hilbert Space (RKHS). Recently, Jacobsen and Cutkosky (2024) introduced a discounted Vovk-Azoury-Warmuth (DVAW) forecaster that achieves optimal dynamic regret in the finite-dimensional case. In this work, we lift their approach to the non-parametric domain by synthesizing the DVAW framework with a random feature approximation. We propose a fully adaptive, hierarchical algorithm, which we call H-VAW-D (Hierarchical Vovk-Azoury-Warmuth with Discounting), that learns both the discount factor and the number of random features. We prove that this algorithm, which has a per-iteration computational complexity of $O(T\ln T)$, achieves an expected dynamic regret of $O(T^{2/3}P_T^{1/3} + \sqrt{T}\ln T)$, where $P_T$ is the functional path length of a comparator sequence.

cs.LG

Random feature-based double Vovk-Azoury-Warmuth algorithm for online multi-kernel learning

We introduce a novel multi-kernel learning algorithm, VAW$^2$, for online least squares regression in reproducing kernel Hilbert spaces (RKHS). VAW$^2$ leverages random Fourier feature-based functional approximation and the Vovk-Azoury-Warmuth (VAW) method in a two-level procedure: VAW is used to construct expert strategies from random features generated for each kernel at the first level, and then again to combine their predictions at the second level. A theoretical analysis yields a regret bound of $O(T^{1/2}\ln T)$ in expectation with respect to artificial randomness, when the number of random features scales as $T^{1/2}$. Empirical results on some benchmark datasets demonstrate that VAW$^2$ achieves superior performance compared to the existing online multi-kernel learning algorithms: Raker and OMKL-GF, and to other theoretically grounded method methods involving convex combination of expert predictions at the second level.

cs.LG

SOLO FTRL algorithm for production management with transfer prices

We consider a firm producing and selling $d$ commodities, and consisting from $n$ production and $m$ sales divisions. The firm manager tries to stimulate the best division performance by sequentially selecting internal commodity prices (transfer prices). In the static problem under general strong convexity and compactness assumptions we show that the SOLO FTRL algorithm of Orabona and Pal (2018), applied to the dual problem, gives the estimates of order $T^{-1/4}$ in the number $T$ of iterations for the optimality gap and feasibility residuals. This algorithm uses only the information on division reactions to current prices. It does not depend on any parameters and requires no information on the production and cost functions. The results of similar nature are obtained for the dynamic problem, where these functions depend on an i.i.d. sequence of random variables. We present two computer experiments with one and two commodities. In the static case the transfer prices and the difference between the supply and demand demonstrate fast stabilization. In the dynamic case the same quantities fluctuate around equilibrium values after a short transition phase.

math.OC

Relative utility bounds for empirically optimal portfolios

We consider a single-period portfolio selection problem for an investor, maximizing the expected ratio of the portfolio utility and the utility of a best asset taken in hindsight. The decision rules are based on the history of stock returns with unknown distribution. Assuming that the utility function is Lipschitz or Hölder continuous (the concavity is not required), we obtain high probability utility bounds under the sole assumption that the returns are independent and identically distributed. These bounds depend only on the utility function, the number of assets and the number of observations. For concave utilities similar bounds are obtained for the portfolios produced by the exponentiated gradient method. Also we use statistical experiments to study risk and generalization properties of empirically optimal portfolios. Herein we consider a model with one risky asset and a dataset, containing the stock prices from NYSE.

q-fin.PM

Resource allocation in communication networks with large number of users: the stochastic gradient descent method

We consider a communication network with fixed number of links, shared by large number of users. The resource allocation is performed on the basis of an aggregate utility maximization in accordance with the popular approach, proposed by Kelly and coauthors (1998). The problem is to construct a pricing mechanism for transmission rates to stimulate an optimal allocation of the available resources. In contrast to the usual approach, the proposed algorithm does not use the information on the aggregate traffic over each link. Its inputs are the total number $N$ of users, the link capacities and optimal myopic reactions of randomly selected users to the current prices. The dynamic pricing scheme is based on the dual projected stochastic gradient descent method. For a special class of utility functions $u_i$ we obtain upper bounds for the amount of constraint violation and the deviation of the objective function from the optimal value. These estimates are uniform in $N$ and are of order $O(T^{-1/4})$ in the number $T$ of reaction measurements. We present some computer experiments for quadratic utility functions $u_i$.

math.OC

Robbins-Monro conditions for persistent exploration learning strategies

We formulate simple assumptions, implying the Robbins-Monro conditions for the $Q$-learning algorithm with the local learning rate, depending on the number of visits of a particular state-action pair (local clock) and the number of iteration (global clock). It is assumed that the Markov decision process is communicating and the learning policy ensures the persistent exploration. The restrictions are imposed on the functional dependence of the learning rate on the local and global clocks. The result partially confirms the conjecture of Bradkte (1994).

stat.ML

PDE approach to the problem of online prediction with expert advice: a construction of potential-based strategies

We consider a sequence of repeated prediction games and formally pass to the limit. The supersolutions of the resulting non-linear parabolic partial differential equation are closely related to the potential functions in the sense of N.\,Cesa-Bianci, G.\,Lugosi (2003). Any such supersolution gives an upper bound for forecaster's regret and suggests a potential-based prediction strategy, satisfying the Blackwell condition. A conventional upper bound for the worst-case regret is justified by a simple verification argument.

cs.LG

Asymptotic efficiency of the proportional compensation scheme for a large number of producers

We consider a manager, who allocates some fixed total payment amount between $N$ rational agents in order to maximize the aggregate production. The profit of $i$-th agent is the difference between the compensation (reward) obtained from the manager and the production cost. We compare (i) the \emph{normative} compensation scheme, where the manager enforces the agents to follow an optimal cooperative strategy; (ii) the \emph{linear piece rates} compensation scheme, where the manager announces an optimal reward per unit good; (iii) the \emph{proportional} compensation scheme, where agent's reward is proportional to his contribution to the total output. Denoting the correspondent total production levels by $s^*$, $\hat s$ and $\overline s$ respectively, where the last one is related to the unique Nash equilibrium, we examine the limits of the prices of anarchy $\mathscr A_N=s^*/\overline s$, $\mathscr A_N'=\hat s/\overline s$ as $N\to\infty$. These limits are calculated for the cases of identical convex costs with power asymptotics at the origin, and for power costs, corresponding to the Coob-Douglas and generalized CES production functions with decreasing returns to scale. Our results show that asymptotically no performance is lost in terms of $\mathscr A'_N$, and in terms of $\mathscr A_N$ the loss does not exceed $31\%$.

econ.GN

Minimax perfect stopping rules for selling an asset near its ultimate maximum

We study the problem of selling an asset near its ultimate maximum in the minimax setting. The regret-based notion of a perfect stopping time is introduced. A perfect stopping time is uniquely characterized by its optimality properties and has the following form: one should sell the asset if its price deviates from the running maximum by a certain time-dependent quantity. The related selling rule improves any earlier one and cannot be improved by further delay. The results, which are applicable to a quite general price model, are illustrated by several examples.

q-fin.PM

Asymptotic sequential Rademacher complexity of a finite function class

For a finite function class we describe the large sample limit of the sequential Rademacher complexity in terms of the viscosity solution of a $G$-heat equation. In the language of Peng's sublinear expectation theory, the same quantity equals to the expected value of the largest order statistics of a multidimensional $G$-normal random variable. We illustrate this result by deriving upper and lower bounds for the asymptotic sequential Rademacher complexity.

cs.LG

Rational taxation in an open access fishery model

We consider a model of fishery management, where $n$ agents exploit a single population with strictly concave continuously differentiable growth function of Verhulst type. If the agent actions are coordinated and directed towards the maximization of the discounted cooperative revenue, then the biomass stabilizes at the level, defined by the well known "golden rule". We show that for independent myopic harvesting agents such optimal (or $\varepsilon$-optimal) cooperative behavior can be stimulated by the proportional tax, depending on the resource stock, and equal to the marginal value function of the cooperative problem. To implement this taxation scheme we prove that the mentioned value function is strictly concave and continuously differentiable, although the instantaneous individual revenues may be neither concave nor differentiable.

math.OC

Optimal production and pricing strategies in a dynamic model of monopolistic firm

We consider a deterministic continuous time model of monopolistic firm, which chooses production and pricing strategies of a single good. Firm's goal is to maximize the discounted profit over infinite time horizon. The no-backlogging assumption induces the state constraint on the inventory level. The revenue and production cost functions are assumed to be continuous but, in general, we do not impose the concavity/convexity property. Using the results form the theory of viscosity solutions and Young-Fenchel duality, we derive a representation for the value function, study its regularity properties, and give a complete description of optimal strategies for this non-convex optimal control problem. In agreement with the results of Chazal et al. (2003), it is optimal to liquidate initial inventory in finite time and then use an optimal static strategy. We give a condition, allowing to distinguish if this static strategy can be represented by an ordinary or relaxed control. The latter is related to production cycles. General theory is illustrate by the example of a non-convex production cost, proposed by Arvan and Moses (1981).

math.OC

Central limit theorem under variance uncertainty

We prove the central limit theorem (CLT) for a sequence of independent zero-mean random variables $ξ_j$, perturbed by predictable multiplicative factors $λ_j$ with values in intervals $[\underlineλ_j,\overlineλ_j]$. It is assumed that the sequences $\underlineλ_j$, $\overlineλ_j$ are bounded and satisfy some stabilization condition. Under the classical Lindeberg condition we show that the CLT limit, corresponding to a "worst" sequence $λ_j$, is described by the solution $v$ of one-dimensional $G$-heat equation. The main part of the proof follows Peng's approach to the CLT under sublinear expectations, and utilizes Hölder regularity properties of $v$. Under the lack of such properties, we use the technique of half-relaxed limits from the theory of viscosity solutions.

math.PR

Central limit theorem under uncertain linear transformations

We prove a variant of the central limit theorem (CLT) for a sequence of i.i.d. random variables $ξ_j$, perturbed by a stochastic sequence of linear transformations $A_j$, representing the model uncertainty. The limit, corresponding to a "worst" sequence $A_j$, is expressed in terms of the viscosity solution of the $G$-heat equation. In the context of the CLT under sublinear expectations this nonlinear parabolic equation appeared previously in the papers of S.Peng. Our proof is based on the technique of half-relaxed limits from the theory of approximation schemes for fully nonlinear partial differential equations.

math.PR

Regular finite fuel stochastic control problems with exit time

We consider a class of exit time stochastic control problems for diffusion processes with discounted criterion, where the controller can utilize a given amount of resource, called "fuel". In contrast to the vast majority of existing literature, concerning the "finite fuel" problems, it is assumed that the intensity of fuel consumption is bounded. We characterize the value function of the optimization problem as the unique continuous viscosity solution of the Dirichlet boundary value problem for the correspondent Hamilton-Jacobi-Bellman (HJB) equation. Our assumptions concern the HJB equations, related to the problems with infinite fuel and without fuel. Also, we present computer experiments, for the problems of optimal regulation and optimal tracking of a simple stochastic system with the stable or unstable equilibrium point.

math.OC

Stochastic Perron's method for optimal control problems with state constraints

We apply the stochastic Perron method of Bayraktar and Sîrbu to a general infinite horizon optimal control problem, where the state $X$ is a controlled diffusion process, and the state constraint is described by a closed set. We prove that the value function $v$ is bounded from below (resp., from above) by a viscosity supersolution (resp., subsolution) of the related state constrained problem for the Hamilton-Jacobi-Bellman equation. In the case of a smooth domain, under some additional assumptions, these estimates allow to identify $v$ with a unique continuous constrained viscosity solution of this equation.

math.OC

Verification by stochastic Perron's method in stochastic exit time control problems

We apply the Stochastic Perron method, created by Bayraktar and Sîrbu, to a stochastic exit time control problem. Our main assumption is the validity of the Strong Comparison Result for the related Hamilton-Jacobi-Bellman (HJB) equation. Without relying on Bellman's optimality principle we prove that inside the domain the value function is continuous and coincides with a viscosity solution of the Dirichlet boundary value problem for the HJB equation.

math.OC