SearcharxivSearch

arXiv subjects

Anton Plaksin

Publications and source records attributed to Anton Plaksin.

7 recordsLinked to original sources

SlimSpec: Low-Rank Draft LM-Head for Accelerated Speculative Decoding

Speculative decoding speeds up autoregressive generation in Large Language Models (LLMs) through a two-step procedure, where a lightweight draft model proposes tokens which the target model then verifies in a single forward pass. Although the drafter network is small in modern architectures, its LM-head still performs projection to a large vocabulary, becoming one of the major computational bottlenecks. In prior work this issue has been predominantly addressed via static or dynamic vocabulary truncation. Yet mitigating the bottleneck, these methods bring in extra complexity, such as special vocabulary curation, sophisticated inference-time logic or modifications of the training setup. In this paper, we propose SlimSpec, a low-rank parameterization of the drafter's LM-head that compresses the inner representation rather than the output, preserving full vocabulary support. We evaluate our method with EAGLE-3 drafter across three target models and diverse benchmarks in both latency- and throughput-bound inference regimes. SlimSpec achieves $4\text{-}5\times$ acceleration over the standard LM-head architecture while maintaining a competitive acceptance length, surpassing existing methods by up to $8\text{-}9\%$ of the end-to-end speedup. Our method requires minimal adjustments of training and inference pipelines. Combined with the aforementioned speedup improvements, it makes SlimSpec a strong alternative across wide variety of draft LM-head architectures.

cs.LG

Domain Adaptation of Drag Reduction Policy to Partial Measurements

Feedback control of fluid-based systems poses significant challenges due to their high-dimensional, nonlinear, and multiscale dynamics, which demand real-time, three-dimensional, multi-component measurements for sensing. While such measurements are feasible in digital simulations, they are often only partially accessible in the real world. In this paper, we propose a method to adapt feedback control policies obtained from full-state measurements to setups with only partial measurements. Our approach is demonstrated in a simulated environment by minimising the aerodynamic drag of a simplified road vehicle. Reinforcement learning algorithms can optimally solve this control task when trained on full-state measurements by placing sensors in the wake. However, in real-world applications, sensors are limited and typically only on the vehicle, providing only partial measurements. To address this, we propose to train a Domain Specific Feature Transfer (DSFT) map reconstructing the full measurements from the history of the partial measurements. By applying this map, we derive optimal policies based solely on partial data. Additionally, our method enables determination of the optimal history length and offers insights into the architecture of optimal control policies, facilitating their implementation in real-world environments with limited sensor information.

cs.LG

Zero-Sum Positional Differential Games as a Framework for Robust Reinforcement Learning: Deep Q-Learning Approach

Robust Reinforcement Learning (RRL) is a promising Reinforcement Learning (RL) paradigm aimed at training robust to uncertainty or disturbances models, making them more efficient for real-world applications. Following this paradigm, uncertainty or disturbances are interpreted as actions of a second adversarial agent, and thus, the problem is reduced to seeking the agents' policies robust to any opponent's actions. This paper is the first to propose considering the RRL problems within the positional differential game theory, which helps us to obtain theoretically justified intuition to develop a centralized Q-learning approach. Namely, we prove that under Isaacs's condition (sufficiently general for real-world dynamical systems), the same Q-function can be utilized as an approximate solution of both minimax and maximin Bellman equations. Based on these results, we present the Isaacs Deep Q-Network algorithms and demonstrate their superiority compared to other baseline RRL and Multi-Agent RL algorithms in various environments.

cs.LG

Viscosity solutions of Hamilton-Jacobi equations for neutral-type systems

The paper deals with path-dependent Hamilton-Jacobi equations with a coinvariant derivative which arise in investigations of optimal control problems and differential games for neutral-type systems in Hale's form. A viscosity (generalized) solution of a Cauchy problem for such equations is considered. The existence, uniqueness, and consistency of the viscosity solution are proved. Equivalent definitions of the viscosity solution, including the definitions of minimax and Dini solutions, are obtained. Application of the results to an optimal control problem for neutral-type systems in Hale's form are given.

math.OC

Equivalence of minimax and viscosity solutions of path-dependent Hamilton-Jacobi equations

In the paper, we consider a path-dependent Hamilton-Jacobi equation with coinvariant derivatives over the space of continuous functions. Such equations arise from optimal control problems and differential games for time-delay systems. We study generalized solutions of the considered Hamilton-Jacobi equation both in the minimax and in the viscosity sense. A minimax solution is defined as a functional which epigraph and subgraph satisfy certain conditions of weak invariance, while a viscosity solution is defined in terms of a pair of inequalities for coinvariant sub- and super-gradients. We prove that these two notions are equivalent, which is the main result of the paper. As a corollary, we obtain comparison and uniqueness results for viscosity solutions of a Cauchy problem for the considered Hamilton-Jacobi equation and a right-end boundary condition. The proof is based on a certain property of the coinvariant subdifferential. To establish this property, we develop a technique going back to the proofs of multidirectional mean-value inequalities. In particular, the absence of the local compactness property of the underlying continuous function space is overcome by using Borwein-Preiss variational principle with an appropriate guage-type functional.

math.OC

Viscosity solutions of Hamilton-Jacobi-Bellman-Isaacs equations for time-delay systems

The paper deals with a zero-sum differential game for a dynamical system which motion is described by a nonlinear delay differential equation under an initial condition defined by a piecewise continuous function. The corresponding Cauchy problem for Hamilton-Jacobi-Bellman-Isaacs equation with coinvariant derivatives is derived and the definition of a viscosity solution of this problem is considered. It is proved that the differential game has the value that is the unique viscosity solution. Moreover, based on notions of sub- and superdifferentials corresponding to coinvariant derivatives, the infinitesimal description of the viscosity solution is obtained. The example of applying these results is given.

math.OC

Minimax and Viscosity Solutions of Hamilton-Jacobi-Bellman Equations for Time-Delay Systems

The paper deals with a Bolza optimal control problem for a dynamical system which motion is described by a delay differential equation under an initial condition defined by a piecewise continuous function. For the value functional in this problem, the Cauchy problem for the Hamilton-Jacobi-Bellman equation with coinvariant derivatives is considered. Minimax and viscosity solutions of this problem are studied. It is proved that both of these solutions exist, are unique and coincide with the value functional.

math.OC