SearcharxivSearch

arXiv subjects

Vivek Shripad Borkar

Publications and source records attributed to Vivek Shripad Borkar.

5 recordsLinked to original sources

Stochastic approximation in non-markovian environments revisited

Based on some recent work of the author on stochastic approximation in non-markovian environments, the situation when the driving random process is non-ergodic in addition to being non-markovian is considered. Using this, we propose an analytic framework for understanding transformer based learning, specifically, the `attention' mechanism, and continual learning, both of which depend on the entire past in principle.

stat.ML

A theoretical basis for model collapse in recursive training

It is known that recursive training from generative models can lead to the so called `collapse' of the simulated probability distribution. This note shows that one in fact gets two different asymptotic behaviours depending on whether an external source, howsoever minor, is also contributing samples.

math.PR

A dynamic view of some anomalous phenomena in SGD

It has been observed by Belkin et al.\ that over-parametrized neural networks exhibit a `double descent' phenomenon. That is, as the model complexity (as reflected in the number of features) increases, the test error initially decreases, then increases, and then decreases again. A counterpart of this phenomenon in the time domain has been noted in the context of epoch-wise training, viz., the test error decreases with the number of iterates, then increases, then decreases again. Another anomalous phenomenon is that of \textit{grokking} wherein two regimes of descent are interrupted by a third regime wherein the mean loss remains almost constant. This note presents a plausible explanation for these and related phenomena by using the theory of two time scale stochastic approximation, applied to the continuous time limit of the gradient dynamics. This gives a novel perspective for an already well studied theme.

math.OC

Risk-sensitive control, single controller games and linear programming

This article recalls the recent work on a linear programming formulation of infinite horizon risk-sensitive control via its equivalence with a single controller game, using a classic work of Vrieze. This is then applied to a constrained risk-sensitive control problem with a risk-sensitive cost and risk-sensitive constraint. This facilitates a Lagrange multiplier based resolution thereof. In the process, this leads to an unconstrained linear program and its dual, parametrized by a parameter that is a surrogate for Lagrange multiplier. This also opens up the possibility of a primal - dual type numerical scheme wherein the linear program is a subroutine within the subgradient ascent based update rule for the Lagrange multiplier. This equivalent unconstrained risk-sensitive control formulation does not seem obvious without the linear programming equivalents as intermediaries. We also discuss briefly other related algorithmic possibilities for future research.

math.OC

A variational formula for risk-sensitive reward

We derive a variational formula for the optimal growth rate of reward in the infinite horizon risk-sensitive control problem for discrete time Markov decision processes with compact metric state and action spaces, extending a formula of Donsker and Varadhan for the Perron-Frobenius eigenvalue of a positive operator. This leads to a concave maximization formulation of the problem of determining this optimal growth rate.

math.OC