SearcharxivSearch

arXiv subjects

Clemens Sirotenko

Publications and source records attributed to Clemens Sirotenko.

8 recordsLinked to original sources

A neural network approach to learning solutions of a class of elliptic variational inequalities

We develop a weak adversarial approach to solving obstacle problems using neural networks. By employing (generalised) regularised gap functions and their properties we rewrite the obstacle problem (which is an elliptic variational inequality) as a minmax problem, providing a natural formulation amenable to learning. Our approach, in contrast to much of the literature, does not require the elliptic operator to be symmetric. We provide an error analysis for suitable discretisations of the continuous problem, estimating in particular the approximation and statistical errors. Parametrising the solution and test function as neural networks, we apply a modified gradient descent ascent algorithm to treat the problem and conclude the paper with various examples and experiments. Our solution algorithm is in particular able to easily handle obstacle problems that feature biactivity (or lack of strict complementarity), a situation that poses difficulty for traditional numerical methods.

math.OC

On the Universality of Simple Trust-Region Algorithms

We establish universal complexity guarantees for quadratic trust-region methods and identify a common mechanism underlying their universal behavior under convexity, based on a function-gap model-decrease estimate that appears to be new in the trust-region literature. First, we prove that the basic trust-region method with inexact subproblem solves is universal under convexity. Under a $ν$-Hölder-continuous Hessian, it attains the global complexity bound $\mathcal{O}(\varepsilon^{-1/(1+ν)})$ for computing an $\varepsilon$-approximate minimizer, without knowledge of $ν\in[0,1]$ or the corresponding Hölder constant. In the nonconvex regime, the method retains the classical $\mathcal{O}(\varepsilon^{-2})$ first-order complexity bound under the usual additional bounded-Hessian assumption. With suitably vanishing inexactness, it also recovers Q-superlinear local convergence for $ν=0$ and convergence of order $1+ν$ for $ν\in(0,1]$. Second, we show that the same convex-universal mechanism applies to a trust-region variant with exact subproblem solves and a simple modification of the acceptance ratio. This variant is universal simultaneously in the nonconvex, convex, and local regimes: it attains the optimal nonconvex first-order complexity $\mathcal{O}(\varepsilon^{-(2+ν)/(1+ν)})$, while preserving the universal convex complexity and the local Newton rates. These guarantees require no knowledge of $ν$ or its Hölder constant. Both methods use the usual quadratic trust-region model and the classical radius-update mechanism, without gradient-dependent radii or model modifications such as cubic, gradient, or tensor regularization. The results show that the trust-region mechanism is inherently adaptive across nonconvex, convex, and locally strongly convex regimes, providing further theoretical support for the practical success of trust-region methods.

math.OC

Skip the Hessian, Keep the Rates: Globalized Semismooth Newton with Lazy Hessian Updates

Second-order methods are provably faster than first-order methods, and their efficient implementations for large-scale optimization problems have attracted significant attention. Yet, optimization problems in ML often have nonsmooth derivatives, which makes the existing convergence rate theory of second-order methods inapplicable. In this paper, we propose a new semismooth Newton method (SSN) that enjoys both global convergence rates and asymptotic superlinear convergence without requiring second-order differentiability. Crucially, our method does not require (generalized) Hessians to be evaluated at each iteration but only periodically, and it reuses stale Hessians otherwise (i.e., it performs lazy Hessian updates), saving compute cost and often leading to significant speedups in time, whilst still maintaining strong global and local convergence rate guarantees. We develop our theory in an infinite-dimensional setting and illustrate it with numerical experiments on matrix factorization and neural networks with Lipschitz constraints.

math.OC

LeAP-SSN: A Semismooth Newton Method with Global Convergence Rates

We propose LeAP-SSN (Levenberg--Marquardt Adaptive Proximal Semismooth Newton method), a semismooth Newton-type method with a simple, parameter-free globalisation strategy that guarantees convergence from arbitrary starting points in nonconvex settings to stationary points, and under a Polyak--Lojasiewicz condition, to a global minimum, in Hilbert spaces. The method employs an adaptive Levenberg--Marquardt regularisation for the Newton steps, combined with backtracking, and does not require knowledge of problem-specific constants. We establish global nonasymptotic rates: $\mathcal{O}(1/k)$ for convex problems in terms of objective values, $\mathcal{O}(1/\sqrt{k})$ under nonconvexity in terms of subgradients, and linear convergence under a Polyak--Lojasiewicz condition. The algorithm achieves superlinear convergence under mild semismoothness and Dennis--Moré or partial smoothness conditions, even for non-isolated minimisers. By combining strong global guarantees with superlinear local rates in a fully parameter-agnostic framework, LeAP-SSN bridges the gap between globally convergent algorithms and the fast asymptotics of Newton's method. The practical efficiency of the method is illustrated on representative problems from imaging, contact mechanics, and machine learning.

math.OC

Dictionary Learning Based Regularization in Quantitative MRI: A Nested Alternating Optimization Framework

In this article, we propose a novel regularization method for a class of nonlinear inverse problems that is inspired by an application in quantitative magnetic resonance imaging (qMRI). The latter is a special instance of a general dynamical image reconstruction technique, wherein a radio-frequency pulse sequence gives rise to a time-discrete physics-based mathematical model which acts as a side constraint in our inverse problem. To enhance reconstruction quality, we employ dictionary learning as a data-adaptive regularizer, capturing complex tissue structures beyond handcrafted priors. For computing a solution of the resulting non-convex and non-smooth optimization problem, we alternate between updating the physical parameters of interest via a Levenberg-Marquardt approach and performing several iterations of a dictionary learning algorithm. This process falls under the category of nested alternating optimization schemes. We develop a general overall algorithmic framework whose convergence theory is not directly available in the literature. Global sub-linear and local strong linear convergence in infinite dimensions under certain regularity conditions for the sub-differentials are investigated based on the Kurdyka-Lojasiewicz inequality. Eventually, numerical experiments demonstrate the practical potential and unresolved challenges of the method.

math.OC

Data-driven methods for quantitative imaging

In the field of quantitative imaging, the image information at a pixel or voxel in an underlying domain entails crucial information about the imaged matter. This is particularly important in medical imaging applications, such as quantitative Magnetic Resonance Imaging (qMRI), where quantitative maps of biophysical parameters can characterize the imaged tissue and thus lead to more accurate diagnoses. Such quantitative values can also be useful in subsequent, automatized classification tasks in order to discriminate normal from abnormal tissue, for instance. The accurate reconstruction of these quantitative maps is typically achieved by solving two coupled inverse problems which involve a (forward) measurement operator, typically ill-posed, and a physical process that links the wanted quantitative parameters to the reconstructed qualitative image, given some underlying measurement data. In this review, by considering qMRI as a prototypical application, we provide a mathematically-oriented overview on how data-driven approaches can be employed in these inverse problems eventually improving the reconstruction of the associated quantitative maps.

math.OC

Unrolled three-operator splitting for parameter-map learning in Low Dose X-ray CT reconstruction

We propose a method for fast and automatic estimation of spatially dependent regularization maps for total variation-based (TV) tomography reconstruction. The estimation is based on two distinct sub-networks, with the first sub-network estimating the regularization parameter-map from the input data while the second one unrolling T iterations of the Primal-Dual Three-Operator Splitting (PD3O) algorithm. The latter approximately solves the corresponding TV-minimization problem incorporating the previously estimated regularization parameter-map. The overall network is then trained end-to-end in a supervised learning fashion using pairs of clean-corrupted data but crucially without the need of having access to labels for the optimal regularization parameter-maps.

math.OC

Learning Regularization Parameter-Maps for Variational Image Reconstruction using Deep Neural Networks and Algorithm Unrolling

We introduce a method for fast estimation of data-adapted, spatio-temporally dependent regularization parameter-maps for variational image reconstruction, focusing on total variation (TV)-minimization. Our approach is inspired by recent developments in algorithm unrolling using deep neural networks (NNs), and relies on two distinct sub-networks. The first sub-network estimates the regularization parameter-map from the input data. The second sub-network unrolls $T$ iterations of an iterative algorithm which approximately solves the corresponding TV-minimization problem incorporating the previously estimated regularization parameter-map. The overall network is trained end-to-end in a supervised learning fashion using pairs of clean-corrupted data but crucially without the need of having access to labels for the optimal regularization parameter-maps. We prove consistency of the unrolled scheme by showing that the unrolled energy functional used for the supervised learning $Γ$-converges as $T$ tends to infinity, to the corresponding functional that incorporates the exact solution map of the TV-minimization problem. We apply and evaluate our method on a variety of large scale and dynamic imaging problems in which the automatic computation of such parameters has been so far challenging: 2D dynamic cardiac MRI reconstruction, quantitative brain MRI reconstruction, low-dose CT and dynamic image denoising. The proposed method consistently improves the TV-reconstructions using scalar parameters and the obtained parameter-maps adapt well to each imaging problem and data by leading to the preservation of detailed features. Although the choice of the regularization parameter-maps is data-driven and based on NNs, the proposed algorithm is entirely interpretable since it inherits the properties of the respective iterative reconstruction method from which the network is implicitly defined.

math.OC