Searcharxiv⌕ Search

arXiv subjects

Nikita Yudin

Publications and source records attributed to Nikita Yudin.

5 recordsLinked to original sources

Mathematical methods of reinforcement learning

Reinforcement learning (RL) is increasingly grounded in tools from probability, optimization, and operator theory. This survey organizes the mathematical structures that underpin the design and analysis of modern algorithms in RL. We begin from Markov decision processes (MDPs) and the Bellman operators, emphasizing contraction mappings, monotonicity, and fixed-point theory that yield convergence guarantees and rates for value and policy iteration, and temporal-difference schemes. We then develop the optimization perspective: stochastic approximation and martingale methods, convex duality and the role of regularization linking mirror/proximal methods. Function approximation is treated through linear and non-linear settings, covering stabilization, error decomposition, and sample-complexity via concentration inequalities for dependent data and mixing processes. We further cover off-policy evaluation/learning, constrained RL and constrained MDPs (CMDPs). Throughout we unify algorithmic templates under common operator and variational lenses, highlighting both finite-sample bounds and asymptotic results. Our presentation is intended to provide a unified mathematical entry point for researchers in probability, optimization, and statistics interested in reinforcement learning.

math.OC↗

Non-linear in-band interference cancellation on base of conjugate gradients method

This paper investigates one possible solution to the problem of self-interference cancellation (SIC) arising in the design of in-band full-duplex (IBFD) communication systems. Self-interference cancellation is performed in the digital domain using multilayer nonlinear models adapted via gradient-based optimization. The presence of local minima and saddle points during the adaptation of multilayer models limits the direct use of second-order methods due to the indefiniteness of the hessian matrix. The mixed Newton method can address the saddle-point issue; however, it requires significant computational resources. In this work, a conjugate gradient (CG) method constructed on the base of the mixed Newton method (MNM) is proposed. The method exploits information from mixed second-order derivatives of the loss function without explicit computation the full hessian matrix. As a result, the proposed approach achieves a higher convergence rate than first-order methods while requiring significantly lower computational resources than conventional second-order methods when adapting multilayer nonlinear self-interference cancellers in full-duplex communication systems.

math.OC↗

Optimization in complex spaces with the Mixed Newton Method

We propose a second-order method for unconditional minimization of functions $f(z)$ of complex arguments. We call it the Mixed Newton Method due to the use of the mixed Wirtinger derivative $\frac{\partial^2f}{\partial\bar z\partial z}$ for computation of the search direction, as opposed to the full Hessian $\frac{\partial^2f}{\partial(z,\bar z)^2}$ in the classical Newton method. The method has been developed for specific applications in wireless network communications, but its global convergence properties are shown to be superior on a more general class of functions $f$, namely sums of squares of absolute values of holomorphic functions. In particular, for such objective functions minima are surrounded by attraction basins, while the iterates are repelled from other types of critical points. We provide formulas for the asymptotic convergence rate and show that in the scalar case the method reduces to the well-known complex Newton method for the search of zeros of holomorphic functions. In this case, it exhibits generically fractal global convergence patterns.

math.OC↗

Mixed Newton Method for Optimization in Complex Spaces

In this paper, we modify and apply the recently introduced Mixed Newton Method, which is originally designed for minimizing real-valued functions of complex variables, to the minimization of real-valued functions of real variables by extending the functions to complex space. We show that arbitrary regularizations preserve the favorable local convergence properties of the method, and construct a special type of regularization used to prevent convergence to complex minima. We compare several variants of the method applied to training neural networks with real and complex parameters.

math.OC↗

Flexible Modification of Gauss-Newton Method and Its Stochastic Extension

This work presents a novel version of recently developed Gauss-Newton method for solving systems of nonlinear equations, based on upper bound of solution residual and quadratic regularization ideas. We obtained for such method global convergence bounds and under natural non-degeneracy assumptions we present local quadratic convergence results. We developed stochastic optimization algorithms for presented Gauss-Newton method and justified sub-linear and linear convergence rates for these algorithms using weak growth condition (WGC) and Polyak-Lojasiewicz (PL) inequality. We show that Gauss-Newton method in stochastic setting can effectively find solution under WGC and PL condition matching convergence rate of the deterministic optimization method. The suggested method unifies most practically used Gauss-Newton method modifications and can easily interpolate between them providing flexible and convenient method easily implementable using standard techniques of convex optimization.

math.OC↗