SearcharxivSearch

arXiv subjects

Romain Petit

Publications and source records attributed to Romain Petit.

6 recordsLinked to original sources

On the global convergence of gradient flow for wide shallow models beyond homogeneous nonlinearities

A surprising phenomenon in the training of neural networks is the ability of gradient descent to find global minimizers of the training loss despite its non-convexity. Following earlier work, we investigate this behavior for wide shallow models. Existing global convergence results primarily concern models with positively one-homogeneous nonlinearities, such as ReLU activations, and models with scalar output weights and bounded nonlinearities, such as sigmoid activations. We study a broader class of models, including multi-head attention layers and two-layer networks with bounded or asymptotically positively one-homogeneous activations and vector output weights. Building upon [Chizat and Bach, 2018], we prove that, in the limit of many hidden neurons or attention heads, non-global minimizers of the training loss are unstable under mean-field gradient flow dynamics by constructing "escape regions" in the parameter space. Our global convergence statements are conditional in the following sense: if the mean-field gradient flow converges in W2, then its limit must be a global minimizer. We revisit the bounded nonlinearity, scalar-output setting of [CB18], giving an escape region construction adapted to unbounded nonlinear parameter domains. We also propose new constructions for nonlinearities with at most linear growth under a non-degeneracy assumption and for asymptotically positively one-homogeneous nonlinearities. Finally, we show the well-posedness and stability estimates for the mean-field training dynamics under sub-Gaussian initializations.

math.OC

On the non-convexity issue in the radial Calder\'on problem

A classical approach to the Calder\'on problem is to estimate the unknown conductivity by solving a nonlinear least-squares problem. It leads to a nonconvex optimization problem which is generally believed to be riddled with bad local minimums. We revisit this issue in the case of piecewise constant radial conductivities and prove that, contrary to previous claims, there are no spurious critical points in the case of two scalar unknowns with no measurement noise. We also provide a partial proof of this result in the general setting which holds under a numerically verifiable assumption. Finally, we investigate whether a recently proposed approach based on convexification yields better reconstructions. For the first time, we propose a way to implement it in practice and show that it is consistently outperformed by some least squares solvers, which are also faster and require less measurements.

math.NA

A convex lifting approach for the Calder\'on problem

The Calder\'on problem consists in recovering an unknown coefficient of a partial differential equation from boundary measurements of its solution. These measurements give rise to a highly nonlinear forward operator. As a consequence, the development of reconstruction methods for this inverse problem is challenging, as they usually suffer from the problem of local convergence. To circumvent this issue, we propose an alternative approach based on lifting and convex relaxation techniques, that have been successfully developed for solving finite-dimensional quadratic inverse problems. This leads to a convex optimization problem whose solution coincides with the sought-after coefficient, provided that a non-degenerate source condition holds. We demonstrate the validity of our approach on a toy model where the solution of the partial differential equation is known everywhere in the domain. In this simplified setting, we verify that the non-degenerate source condition holds under certain assumptions on the unknown coefficient. We leave the investigation of its validity in the Calder\'on setting for future works.

math.AP

Localization of point scatterers via sparse optimization on measures

We consider the inverse scattering problem for time-harmonic acoustic waves in a medium with pointwise inhomogeneities. In the Foldy-Lax model, the estimation of the scatterers' locations and intensities from far field measurements can be recast as the recovery of a discrete measure from nonlinear observations. We propose a "linearize and locally optimize" approach to perform this reconstruction. We first solve a convex program in the space of measures (known as the Beurling LASSO), which involves a linearization of the forward operator (the far field pattern in the Born approximation). Then, we locally minimize a second functional involving the nonlinear forward map, using the output of the first step as initialization. We provide guarantees that the output of the first step is close to the sought-after measure when the scatterers have small intensities and are sufficiently separated. We also provide numerical evidence that the second step still allows for accurate recovery in settings that are more involved.

math.NA

Exact recovery of the support of piecewise constant images via total variation regularization

This work is concerned with the recovery of piecewise constant images from noisy linear measurements. We study the noise robustness of a variational reconstruction method, which is based on total (gradient) variation regularization. We show that, if the unknown image is the superposition of a few simple shapes, and if a non-degenerate source condition holds, then, in the low noise regime, the reconstructed images have the same structure: they are the superposition of the same number of shapes, each a smooth deformation of one of the unknown shapes. Moreover, the reconstructed shapes and the associated intensities converge to the unknown ones as the noise goes to zero.

math.NA

Towards Off-the-grid Algorithms for Total Variation Regularized Inverse Problems

We introduce an algorithm to solve linear inverse problems regularized with the total (gradient) variation in a gridless manner. Contrary to most existing methods, that produce an approximate solution which is piecewise constant on a fixed mesh, our approach exploits the structure of the solutions and consists in iteratively constructing a linear combination of indicator functions of simple polygons.

eess.SP