Searcharxiv⌕ Search

arXiv subjects

Nadezda Sukhorukova

Publications and source records attributed to Nadezda Sukhorukova.

At least 19 recordsLinked to original sources

KAN vs LSTM Performance in Time Series Forecasting

This study presents a controlled comparison of baseline Kolmogorov-Arnold Networks (KAN), implemented via PyKAN, and Long Short-Term Memory (LSTM) networks for the forecasting of stochastic, non-stationary financial time series. The two architectures are assessed in terms of predictive accuracy, computational efficiency, and interpretability, with accuracy measured by the Root Mean Square Error (RMSE) in normalised feature space. Under a direct multi-output forecasting protocol, LSTM attains clearly superior accuracy across all tested prediction horizons, consistent with its well-established effectiveness for sequential data modelling. Baseline KAN, although offering theoretical interpretability through the Kolmogorov-Arnold representation theorem, exhibits substantially higher error rates and limited practical applicability for time series forecasting in its standard form. Several specialised temporal variants -- including Temporal KAN and Time-Frequency KAN -- have since been proposed to address these sequential modelling limitations, but they lie outside the scope of the present study. KAN is observed to converge faster during training under the configurations tested, although direct runtime comparisons are constrained by methodological factors. These findings support the adoption of LSTM for accuracy-critical financial forecasting and establish an empirical baseline for standard KAN on stochastic sequential data, motivating further investigation of temporally-aware KAN architectures. The study benchmarks baseline KAN against baseline LSTM only; the results do not extend to specialised KAN variants designed for sequential data, nor to the broader family of temporal models.

cs.LG↗

Difference of Convex (DC) approach for neural network approximation with uniform loss function

Neural networks (NNs) can be viewed as approximation tools. Traditionally, NNs are relying on gradient and stochastic gradient (SG) methods. There are a number of available computational packages for constructing least squares approximations, while uniform (minimax) approximations are hard due to their nonsmooth nature. It was recently demonstrated that a difference convex (DC) programming approach is an efficient alternative optimiser for NNs. In this paper, we demonstrate that a DC programming approach is also efficient for minimax approximation. In our numerical experiments, we compare a DC-programming approach and ADAMAX, a commonly used method for minimax NN approximations.

math.OC↗

KKT-based optimality conditions for neural network approximation

In this paper, we obtain necessary optimality conditions for neural network approximation. We consider neural networks in Manhattan ($l_1$ norm) and Chebyshev ($\max$ norm). The optimality conditions are based on neural networks with at most one hidden layer. We reformulate nonsmooth unconstrained optimisation problems as larger dimension constrained problems with smooth objective functions and constraints. Then we use KKT conditions to develop the necessary conditions and present the optimality conditions in terms of convex analysis and convex sets.

math.OC↗

Nonsmooth Optimisation and neural networks

In this paper, we study neural networks from the point of view of nonsmooth optimisation, namely, quasidifferential calculus. We restrict ourselves to the case of uniform approximation by a neural network without hidden layers, the activation functions are restricted to continuous strictly increasing functions. We develop an algorithm for computing the approximation with one hidden layer through a step-by-step procedure. The nonsmooth analysis techniques demonstrated their efficiency. In particular, they partially explain why the developed step-by-step procedure may run without any objective function improvement after just one step of the procedure.

math.OC↗

Bivariate rational approximations of the general temperature integral

The non-isothermal analysis of materials with the application of the Arrhenius equation involves temperature integration. If the frequency factor in the Arrhenius equation depends on temperature with a power-law relationship, the integral is known as the general temperature integral. This integral which has no analytical solution is estimated by the approximation functions with different accuracies. In this article, the rational approximations of the integral were obtained based on the minimization of the maximal deviation of bivariate functions. Mathematically, these problems belong to the class of quasiconvex optimization and can be solved using the bisection method. The approximations obtained in this study are more accurate than all approximates available in the literature.

math.OC↗

Variational properties of the abstract subdifferential operator

Abstract convexity generalises classical convexity by considering the suprema of functions taken from an arbitrarily defined set of functions. These are called the abstract linear (abstract affine) functions. The purpose of this paper is to study the abstract subdifferential. We obtain a number of results on the calculus of this subdifferential: summation and composition rules, and prove that under some reasonable conditions the subdifferential is a maximal abstract monotone operator. Another contribution of this paper is a counterexample that demonstrates that the separation theorem between two abstract convex sets is generally not true. The lack of the extension of separation results to the case of abstract convexity is one of the obstacles in the development of abstract convexity based numerical methods.

math.OC↗

Best free knot linear spline approximation and its application to neural networks

The problem of fixed knot approximation is convex and there are several efficient approaches to solve this problem, yet, when the knots joining the affine parts are also variable, finding conditions for a best Chebyshev approximation remains an open problem. It was noticed before that piecewise linear approximation with free knots is equivalent to neural network approximation with piecewise linear activation functions (for example ReLU). In this paper, we demonstrate that in the case of one internal free knot, the problem of linear spline approximation can be reformulated as a mixed-integer linear programming problem and solved efficiently using, for example, a branch and bound type method. We also present a new sufficient optimality condition for a one free knot piecewise linear approximation. The results of numerical experiments are provided.

math.OC↗

A comparison of rational and neural network based approximations

Rational and neural network based approximations are efficient tools in modern approximation. These approaches are able to produce accurate approximations to nonsmooth and non-Lipschitz functions, including multivariate domain functions. In this paper we compare the efficiency of function approximation using rational approximation, neural network and their combinations. It was found that rational approximation is superior to neural network based approaches with the same number of decision variables. Our numerical experiments demonstrate the efficiency of rational approximation, even when the number of approximation parameters (that is, the dimension of the corresponding optimisation problems) is small. Another important contribution of this paper lies in the improvement of rational approximation algorithms. Namely, the optimisation based algorithms for rational approximation can be adjusted to in such a way that the conditioning number of the constraint matrices are controlled. This simple adjustment enables us to work with high dimension optimisation problems and improve the design of the neural network. The main strength of neural networks is in their ability to handle models with a large number of variables: complex models are decomposed in several simple optimisation problems. Therefore the the large number of decision variables is in the nature of neural networks.

math.OC↗

Application and issues in abstract convexity

The theory of abstract convexity, also known as convexity without linearity, is an extension of the classical convex analysis. There are a number of remarkable results, mostly concerning duality, and some numerical methods, however, this area has not found many practical applications yet. In this paper we study the application of abstract convexity to function approximation. Another important research direction addressed in this paper is the connection with the so-called axiomatic convexity.

math.OC↗

Deep Learning with Nonsmooth Objectives

We explore the potential for using a nonsmooth loss function based on the max-norm in the training of an artificial neural network. We hypothesise that this may lead to superior classification results in some special cases where the training data is either very small or unbalanced. Our numerical experiments performed on a simple artificial neural network with no hidden layers (a setting immediately amenable to standard nonsmooth optimisation techniques) appear to confirm our hypothesis that uniform approximation based approaches may be more suitable for the datasets with reliable training data that either is limited size or biased in terms of relative cluster sizes.

cs.LG↗

The extension of linear inequality method for generalised rational Chebyshev approximation to approximation by general quasilinear functions

In this paper we demonstrate that a well known linear inequality method developed for rational Chebyshev approximation is equivalent to the application of the bisection method used in quasiconvex optimisation. Although this correspondence is not surprising, it naturally connects rational and generalised rational Chebyshev approximation problems with modern developments in the area of quasiconvex functions and therefore offers more theoretical and computational tools for solving this problem. %In particular, this observation leads a straightforward extension of the linear inequality method to generalised rational approximation in the sense of Loeb (ratio of two linear forms). The second important contribution of this paper is the extension of the linear inequality method to a broader class of Chebyshev approximation problems, where the corresponding objective functions remain quasiconvex. In this broader class of functions, the inequalities are no longer required to be linear: it is enough for each inequality to define a convex set and the computational challenge is in solving the corresponding convex feasibility problems. Therefore, we propose a more systematic and general approach for treating Chebyshev approximation problems. In particular, we are looking at the problems where the approximations are quasilinear functions with respect to their parameters, that are also the decision variables in the corresponding optimisation problems.

math.OC↗

An algorithm for best generalised rational approximation of continuous functions

The motivation of this paper is the development of an optimisation method for solving optimisation problems appearing in Chebyshev rational and generalised rational approximation problems, where the approximations are constructed as ratios of linear forms (linear combinations of basis functions). The coefficients of the linear forms are subject to optimisation and the basis functions are continuous function. It is known that the objective functions in generalised rational approximation problems are quasi-convex. In this paper we also prove a stronger result, the objective functions are pseudo-convex in the sense of Penot and Quang. Then we develop numerical methods, that are efficient for a wide range of pseudo-convex functions and test them on generalised rational approximation problems.

math.OC↗

Solving multi-resource allocation and location problems in disaster management through linear programming

In this paper we propose a new efficient linear programming based approach for multi-resource allocation and location problems in disaster management. Such problems require an integer solution and therefore, in most cases, the computations rely on integer and mixed-integer linear programming solvers. In general, these solvers can not handle large scaled problem. In this paper we demonstrate that there exists a large class of disaster management problems whose exact solutions can be obtained by applying the simplex method (linear programming). The results of numerical experiments are provided. Another important contribution of this paper is related to general cluster analysis and allocation. Namely, we demonstrate that the classical $k$-medoid clustering method can be implemented using linear programming techniques (simplex method) without relying on integer solvers.

math.OC↗

Two curve Chebyshev approximation and its application to signal clustering

In this paper we extend a number of important results of the classical Chebyshev approximation theory to the case of simultaneous approximation of two or more functions. The need for this extension is application driven, since such kind of problems appears in the area of curve (signal) clustering. In this paper we propose a new efficient algorithm for signal clustering and develop a procedure that allows one to reuse the results obtained at the previous iteration without recomputing the cluster centres from scratch. This approach is based on the extension of the classical de la Vallee-Poussin's procedure originally developed for polynomial approximation. In this paper, we also develop necessary and sufficient optimality conditions for two curve Chebyshev approximation, that is our core tool for curve clustering. These results are based on application of nonsmooth convex analysis.

math.OC↗

Schur functions for approximation problems

In this paper we propose a new approach to least squares approximation problems. This approach is based on partitioning and Schur function. The nature of this approach is combinatorial, while most existing approaches are based on algebra and algebraic geometry. This problem has several practical applications. One of them is curve clustering. We use this application to illustrate the results.

math.NA↗

Alternance Theorems and Chebyshev Splines Approximation

One of the purposes in this paper is to provide a better understanding of the alternance property which occurs in Chebyshev polynomial approximation and piecewise polynomial approximation problems. In the first part of this paper, we propose an original approach to obtain new proofs of the well known necessary and sufficient optimality conditions. There are two main advantages of this approach. First of all, the proofs are much simpler and easier to understand than the existing proofs. Second, these proofs are constructive and therefore they lead to alternative-based algorithms that can be considered as Remez-type approximation algorithms. In the second part of this paper, we develop new local optimality conditions for free knot polynomial spline approximation. The proofs for free knot approximation are relying on the techniques developed in the first part of this paper.

math.NA↗

Chebyshev multivariate polynomial approximation and point reduction procedure

The theory of Chebyshev (uniform) approximation for univariate polynomial and piecewise polynomial functions has been studied for decades. The optimality conditions are based on the notion of alternating sequence. However, the extension the notion of alternating sequence to the case of multivariate functions is not trivial. The contribution of this paper is two-fold. First of all, we give a geometrical interpretation of the necessary and sufficient optimality condition for multivariate approximation. These optimality conditions are not limited to the case polynomial approximation, where the basis functions are monomials. Second, we develop an algorithm for fast necessary optimality conditions verifications (polynomial case only). Although, this procedure only verifies the necessity, it is much faster than the necessary and sufficient conditions verification. This procedure is based on a point reduction procedure and resembles the univariate alternating sequence based optimality conditions. In the case of univariate approximation, however, these conditions are both necessary and sufficient. Third, we propose a procedure for necessary and sufficient optimality conditions verification that is based on a generalisation of the notion of alternating sequence to the case of multivariate polynomials.

math.NA↗

A generalisation of de la Vallée-Poussin procedure to multivariate approximations

The theory of Chebyshev approximation has been extensively studied. In most cases, the optimality conditions are based on the notion of alternance or alternating sequence (that is, maximal deviation points with alternating deviation signs). There are a number of approximation methods for polynomial and polynomial spline approximation. Some of them are based on the classical de la Vallée-Poussin procedure. In this paper we demonstrate that under certain assumptions the classical de la Vallée-Poussin procedure, developed for univariate polynomial approximation, can be extended to the case of multivariate approximation. The corresponding basis functions are not restricted to be monomials.

math.FA↗