SearcharxivSearch

arXiv subjects

Marcello Carioni

Publications and source records attributed to Marcello Carioni.

At least 19 recordsLinked to original sources

From plans to maps: Nonlocal regularization of optimal transport

We introduce a nonlocal regularization of optimal transport that bridges Kantorovich and Monge formulations. The regularization penalizes oscillations through an interaction kernel singular near the diagonal. We show that, for sufficiently strong singularities, every finite-energy transport plan is induced by a map, yielding equivalence between the regularized Kantorovich and Monge problems. For fractional kernels, the regularizer reduces to a fractional Sobolev seminorm that can allow jump discontinuities, leading to existence results for broad classes of source and target measures. We derive first-order optimality conditions and analyze the vanishing-regularization limit. The regularized functionals $Γ$-converge to the classical Kantorovich problem, while the first-order $Γ$-limit selects, whenever they exist, finite-energy Monge minimizers with minimal nonlocal energy. When no such minimizer exists, we identify a different asymptotic regime and establish a sharp scaling law in a prototypical mass-splitting example. Finally, after lifting the problem to the product of transport plans, we derive a dual formulation using the cone of copositive functions and propose a fixed-point numerical scheme illustrating the regularization.

math.OC

An invertible generative model for forward and inverse problems

We formulate inverse problems in a Bayesian framework and aim to train an invertible generative model that is capable of simulation (i.e., sampling from the likelihood) and inference (i.e., sampling from the posterior). We call such a generative model a Reversible Simulator. We first construct an explicit instance of a reversible simulator by combining lower and upper triangular normalizing flows associated with conditional sampling into a single invertible map. We then establish basic structural properties of this construction, investigate its non-uniqueness, and derive a variational formulation that enables the generative model to be trained directly from paired samples. Finally, we illustrate the proposed framework on analytically tractable and stylized numerical examples, demonstrating its potential as a unified approach to conditional generative modeling for forward and inverse problems.

stat.ML

Approximation Rates for Metaplectic Neural Networks

In this paper we develop quantitative approximation results for shallow neural networks constructed using a dictionary based on metaplectic operators. First, we extend the concept of Barron spaces by considering a symplectically motivated extension of the Fourier transform, known as the metaplectic transform. Then, after establishing embedding between metaplectic Barron spaces and Sobolev spaces we consider a neural metaplectic dictionary and we prove Monte-Carlo approximation bounds for metaplectic Barron functions using finite linear combinations of atoms of the dictionary. Finally, we validate the introduction of the neural metaplectic dictionary by devising a deep neural network architecture that uses as building blocks the atoms of the dictionary. We test it to approximate solutions of time-dependent Schrödinger equations, demonstrating better performance compared to classical phyisics informed neural networks architectures.

cs.LG

Dynamic Optimal Transport with Optimal Preferential Paths

We study a dynamic optimal transport type problem on a domain that consists of two parts: a compact set $Ω\subset \mathbb{R}^d$ (bulk) and a non-intersecting and sufficiently regular curve $Γ\subset Ω$. On each of them, a Benamou-Brenier type dynamic optimal transport problem is considered, yet with an additional mechanism that allows the exchange (at a cost) of mass between bulk and curve. In the respective actions, we also allow for non-linear mobilities. We first ensure the existence of minimizers by relying on the direct method of calculus of variations and we study the asymptotic properties of the minimizers under changes in the parameters regulating the dynamics in $Ω$ and $Γ$. Then, we study the case when the curve $Γ$ is also allowed to change, being the main interest in this paper. To this end, the Tangent-Point energy is added to the action functional in order to preserve the regularity properties of the curve and prevent self-intersections. Also in this case, by relying on suitable compactness estimates both for the time-dependent measures and the curve $Γ$, the existence of optimizers is shown. We extend these analytical findings by numerical simulations based on a primal-dual approach that illustrate the behaviour of geodesics, for fixed and varying curves.

math.AP

A Distributionally Robust Framework for Learned Reconstructions in Inverse Problems

Learned reconstruction operators for inverse problems are typically trained under a fixed noise model, and generalize poorly when the distribution during testing differs from the one assumed during training. Distributionally robust optimization (DRO) addresses this by optimizing against the worst-case distribution within a prescribed ambiguity set, but standard Wasserstein DRO perturbs the full joint distribution uniformly, which can be overly conservative and ignores the physics of the measurement process. We develop a structured DRO framework in which the ambiguity set is restricted to structured perturbations aligned with the data-acquisition process. This allows us to learn data-driven reconstruction operators that remain robust to distributional shifts. By constraining perturbations to subsets such as $P(Y|X)$, our framework models uncertainty in the forward operator and noise model more faithfully, accommodating any noise model expressible as a stochastic forward operator. We establish strong duality for this general formulation and derive explicit finite-dimensional dual representations for perturbations in the joint, marginal, and conditional distributions. A central result is an explicit worst-case risk bound that induces Tikhonov regularization on the Lipschitz constant of the reconstruction operator, and is less conservative relative to standard DRO for well-posed problems. Numerical experiments on deblurring and sinogram-to-CT reconstruction demonstrate improved robustness, stability, and interpretability over standard DRO and MSE baselines. In the linear setting, the learned operator becomes effectively low-rank, truncating at the intrinsic dimension of the data and recovering a data-driven analogue of truncated-SVD regularization.

math.OC

Multi-Headed Transformer Architectures as Time-dependent Wasserstein Gradient Flows

In recent years, transformer architectures have revolutionized the field of language processing, opening the door to previously unforeseen possibilities. However, from a theoretical point of view, the mathematical models proposed in the literature often lack direct contact with the actual architectures and depend on strong simplifying assumptions. In this paper, we reduce this gap by modelling the data flow in multi-headed transformer architectures as time-dependent gradient flows for a suitable interaction energy capturing the design of the attention mechanism. The explicit dependence on time allows us to consider different weights for each head and for each layer, without imposing constraints on the initialization method. Moreover, we prove that, under a suitable integrability assumption on the evolution of the weights, each element of the $ω$-limit set of the gradient flows is a stationary point of the interaction energy at a limiting weight distribution. Finally, we analyse the stability of the gradient flows considering perturbations of both the initial data and the weights. Specifically, on the one hand, we study the robustness of the proposed models with respect to noisy inputs, establishing a continuous dependence of the gradient flows on the initial data and uniqueness of the flows. On the other hand, we prove the $Γ$-convergence of the perturbed interaction energy to the unperturbed one, leading to the convergence of the corresponding gradient flows. We complement these theoretical results with numerical experiments that confirm the predicted energy-dissipation identity and clarify the asymptotic behavior of the dynamics in both the autonomous-like (Ornstein--Uhlenbeck) and the genuinely non-autonomous (oscillating-weights) regimes.

cs.LG

Atomic Gradient Flows: Gradient Flows on Sparse Representations

One of the most popular approaches for solving total variation-regularized optimization problems in the space of measures are Particle Gradient Flows (PGFs). These restrict the problem to linear combinations of Dirac deltas and then perform a Euclidean gradient flow in the weights and positions, significantly reducing the computational cost while still decreasing the energy. In this work, we generalize PGFs to convex optimization problems in arbitrary Banach spaces, which we call Atomic Gradient Flows (AGFs). To this end, the crucial ingredient turns out to be the right notion of particles, chosen here as the extremal points of the unit ball of the regularizer. This choice is motivated by the Krein-Milman theorem, which ensures that minimizers can be approximated by linear combinations of extremal points. We investigate metric gradient flows of the optimization problem when restricted to such sparse representations, for which we define a suitable discretized functional that we show to be to be consistent with the original problem via the means of $Γ$-convergence. We prove that the resulting evolution of the latter is well-defined using a minimizing movement scheme, and we establish conditions ensuring $λ$-convexity and uniqueness of the flow. Then, using Choquet's theorem, we lift the problem into the Wasserstein space on weights and extremal points, and consider Wasserstein gradient flows in this lifted setting. Our main result is that the lifting of the AGF evolution is again a metric gradient flow in the Wasserstein space, verifying the consistency of the approach with respect to a Wasserstein-type dynamic. Finally, we illustrate the applicability of AGFs to several relevant infinite-dimensional problems, including optimization of functions of bounded variation and curves of measures regularized by Optimal Transport-type penalties.

math.OC

A Dual Certificate Approach to Sparsity in Infinite-Width Shallow Neural Networks

In this paper, we study total variation (TV)-regularized training of infinite-width shallow ReLU neural networks, formulated as a convex optimization problem over measures on the unit sphere. Our approach leverages the duality theory of TV-regularized optimization problems to establish rigorous guarantees on the sparsity of the solutions to the training problem. Our analysis further characterizes how and when this sparsity persists in a low noise regime and for small regularization parameter. The key observation that motivates our analysis is that, for ReLU activations, the associated dual certificate is piecewise linear in the weight space. Its linearity regions, which we name dual regions, are determined by the activation patterns of the data via the induced hyperplane arrangement. Taking advantage of this structure, we prove that, on each dual region, the dual certificate admits at most one extreme value. As a consequence, the support of any minimizer is finite, and its cardinality can be bounded from above by a constant depending only on the geometry of the data-induced hyperplane arrangement. Then, we further investigate sufficient conditions ensuring uniqueness of such sparse solution. Finally, under a suitable non-degeneracy condition on the dual certificate along the boundaries of the dual regions, we prove that in the presence of low label noise and for small regularization parameter, solutions to the training problem remain sparse with the same number of Dirac deltas. Additionally, their location and the amplitudes converge, and, in case the locations lie in the interior of a dual region, the convergence happens with a rate that depends linearly on the noise and the regularization parameter.

math.OC

A Lipschitz spaces view of infinitely wide shallow neural networks

We revisit the mean field parametrization of shallow neural networks, using signed measures on unbounded parameter spaces and duality pairings that take into account the regularity and growth of activation functions. This setting directly leads to the use of unbalanced Kantorovich-Rubinstein norms defined by duality with Lipschitz functions, and of spaces of measures dual to those of continuous functions with controlled growth. These allow to make transparent the need for total variation and moment bounds or penalization to obtain existence of minimizers of variational formulations, under which we prove a compactness result in strong Kantorovich-Rubinstein norm, and in the absence of which we show several examples demonstrating undesirable behavior. Further, the Kantorovich-Rubinstein setting enables us to combine the advantages of a completely linear parametrization and ensuing reproducing kernel Banach space framework with optimal transport insights. We showcase this synergy with representer theorems and uniform large data limits for empirical risk minimization, and in proposed formulations for distillation and fusion applications.

math.FA

Nonlocal perimeters and variations: Extremality and decomposability for finite and infinite horizons

We analyze the extremality and decomposability properties with respect to two types of nonlocal perimeters available in the literature, the Gagliardo perimeter based on the eponymous seminorms and the nonlocal distributional Caccioppoli perimeter, both with finite and infinite interaction ranges. A nonlocal notion of indecomposability associated to these perimeters is introduced, and we prove that in both cases it can be characterized solely in terms of the interaction range or horizon $\varepsilon$. Utilizing this, we show that it is possible to uniquely decompose a set into its $\varepsilon$-connected components, establishing a nonlocal analogue of the decomposition theorem of Ambrosio, Caselles, Masnou and Morel. Moreover, the extreme points of the balls induced by the Gagliardo and nonlocal total variation seminorm are identified, which naturally correspond to the two nonlocal perimeters. Surprisingly, while the extreme points in the former case are normalized indicator functions of $\varepsilon$-simple sets, akin to the classical TV-ball, in the latter case they are instead obtained from a nonlocal transformation applied to the extreme points of the TV-ball. Finally, we explore the nonlocal-to-local transition via a $Γ$-limit as $\varepsilon \rightarrow 0$ for both perimeters, recovering the classical Caccioppoli perimeter.

math.FA

Dynamic Optimal Transport with optimal star shaped graphs

We study an optimal transport problem in a compact convex set $Ω\subset\mathbb{R}^d$ where bulk transport is coupled to dynamic optimal transport on a metric graph $ \mathsf{G} = (\mathsf{V},\mathsf{E})$ which is embedded in $Ω$. We prove existence of solutions for fixed graphs. Next, we consider varying graphs, yet only for the case of star-shaped ones. Here, the action functional is augmented by an additional penalty that prevents the edges of the graph to overlap. This allows to preserve the graph topology and thus to rely on standard techniques in Calculus of Variations in order to show existence of minimizers.

math.AP

Sparsity for dynamic inverse problems on Wasserstein curves with bounded variation

We investigate a dynamic inverse problem using a regularization which implements the so-called Wasserstein-$1$ distance. It naturally extends well-known static problems such as lasso or total variation regularized problems to a (temporally) dynamic setting. Further, the decision variables, realized as BV curves, are allowed to exhibit discontinuities, in contrast to the design variables in classical optimal transport based regularization techniques. We prove the existence and a characterization of a sparse solution. Further, we use an adaption of the fully-corrective generalized conditional gradient method to experimentally justify that the determination of BV curves in the Wasserstein-$1$ space is numerically implementable.

math.OC

Perturbation-Aware Distributionally Robust Optimization for Inverse Problems

This paper builds on classical distributionally robust optimization techniques to construct a comprehensive framework that can be used for solving inverse problems. Given an estimated distribution of inputs in $X$ and outputs in $Y$, an ambiguity set is constructed by collecting all the perturbations that belong to a prescribed set $K$ and are inside an entropy-regularized Wasserstein ball. By finding the worst-case reconstruction within $K$ one can produce reconstructions that are robust with respect to various types of perturbations: $X$-robustness, $Y|X$-robustness and, more general, targeted robustness depending on noise type, imperfect forward operators and noise anisotropies. After defining the general robust optimization problem, we derive its (weak) dual formulation and we use it to design an efficient algorithm. Finally, we demonstrate the effectiveness of our general framework to solve matrix inversion and deconvolution problems defining $K$ as the set of multivariate Gaussian perturbations in $Y|X$.

math.OC

A sparse optimization approach to infinite infimal convolution regularization

In this paper we introduce the class of infinite infimal convolution functionals and apply these functionals to the regularization of ill-posed inverse problems. The proposed regularization involves an infimal convolution of a continuously parametrized family of convex, positively one-homogeneous functionals defined on a common Banach space $X$. We show that, under mild assumptions, this functional admits an equivalent convex lifting in the space of measures with values in $X$. This reformulation allows us to prove well-posedness of a Tikhonov regularized inverse problem and opens the door to a sparse analysis of the solutions. In the case of finite-dimensional measurements we prove a representer theorem, showing that there exists a solution of the inverse problem that is sparse, in the sense that it can be represented as a linear combination of the extremal points of the ball of the lifted infinite infimal convolution functional. Then, we design a generalized conditional gradient method for computing solutions of the inverse problem without relying on an a priori discretization of the parameter space and of the Banach space $X$. The iterates are constructed as linear combinations of the extremal points of the lifted infinite infimal convolution functional. We prove a sublinear rate of convergence for our algorithm and apply it to denoising of signals and images using, as regularizer, infinite infimal convolutions of fractional-Laplacian-type operators with adaptive orders of smoothness and anisotropies.

math.OC

Exact Sparse Representation Recovery in Signal Demixing and Group BLASSO

In this short article we present the theory of sparse representations recovery in convex regularized optimization problems introduced in (Carioni and Del Grande, arXiv:2311.08072, 2023). We focus on the scenario where the unknowns belong to Banach spaces and measurements are taken in Hilbert spaces, exploring the properties of minimizers of optimization problems in such settings. Specifically, we analyze a Tikhonov-regularized convex optimization problem, where $y_0$ are the measured data, $w$ denotes the noise, and $λ$ is the regularization parameter. By introducing a Metric Non-Degenerate Source Condition (MNDSC) and considering sufficiently small $λ$ and $w$, we establish Exact Sparse Representation Recovery (ESRR) for our problems, meaning that the minimizer is unique and precisely recovers the sparse representation of the original data. We then emphasize the practical implications of this theoretical result through two novel applications: signal demixing and super-resolution with Group BLASSO. These applications underscore the broad applicability and significance of our result, showcasing its potential across different domains.

math.OC

Unsupervised approaches based on optimal transport and convex analysis for inverse problems in imaging

Unsupervised deep learning approaches have recently become one of the crucial research areas in imaging owing to their ability to learn expressive and powerful reconstruction operators even when paired high-quality training data is scarcely available. In this chapter, we review theoretically principled unsupervised learning schemes for solving imaging inverse problems, with a particular focus on methods rooted in optimal transport and convex analysis. We begin by reviewing the optimal transport-based unsupervised approaches such as the cycle-consistency-based models and learned adversarial regularization methods, which have clear probabilistic interpretations. Subsequently, we give an overview of a recent line of works on provably convergent learned optimization algorithms applied to accelerate the solution of imaging inverse problems, alongside their dedicated unsupervised training schemes. We also survey a number of provably convergent plug-and-play algorithms (based on gradient-step deep denoisers), which are among the most important and widely applied unsupervised approaches for imaging problems. At the end of this survey, we provide an overview of a few related unsupervised learning frameworks that complement our focused schemes. Together with a detailed survey, we provide an overview of the key mathematical results that underlie the methods reviewed in the chapter to keep our discussion self-contained.

cs.CV

A General Theory for Exact Sparse Representation Recovery in Convex Optimization

In this paper, we investigate the recovery of the sparse representation of data in general infinite-dimensional optimization problems regularized by convex functionals. We show that it is possible to define a suitable non-degeneracy condition on the minimal-norm dual certificate, extending the well-established non-degeneracy source condition (NDSC) associated with total variation regularized problems in the space of measures, as introduced in (Duval and Peyré, FoCM, 15:1315-1355, 2015). In our general setting, we need to study how the dual certificate is acting, through the duality product, on the set of extreme points of the ball of the regularizer, seen as a metric space. This justifies the name Metric Non-Degenerate Source Condition (MNDSC). More precisely, we impose a second-order condition on the dual certificate, evaluated on curves with values in small neighbourhoods of a given collection of n extreme points. By assuming the validity of the MNDSC, together with the linear independence of the measurements on these extreme points, we establish that, for a suitable choice of regularization parameters and noise levels, the minimizer of the minimization problem is unique and is uniquely represented as a linear combination of n extreme points. The paper concludes by obtaining explicit formulations of the MNDSC for three problems of interest. First, we examine total variation regularized deconvolution problems, showing that the classical NDSC implies our MNDSC, and recovering a result similar to (Duval and Peyré, FoCM, 15:1315-1355, 2015). Then, we consider 1-dimensional BV functions regularized with their BV-seminorm and pairs of measures regularized with their mutual 1-Wasserstein distance. In each case, we provide explicit versions of the MNDSC and formulate specific sparse representation recovery results.

math.OC

Implicit neural representations for unsupervised super-resolution and denoising of 4D flow MRI

4D flow MRI is a non-invasive imaging method that can measure blood flow velocities over time. However, the velocity fields detected by this technique have limitations due to low resolution and measurement noise. Coordinate-based neural networks have been researched to improve accuracy, with SIRENs being suitable for super-resolution tasks. Our study investigates SIRENs for time-varying 3-directional velocity fields measured in the aorta by 4D flow MRI, achieving denoising and super-resolution. We trained our method on voxel coordinates and benchmarked our approach using synthetic measurements and a real 4D flow MRI scan. Our optimized SIREN architecture outperformed state-of-the-art techniques, producing denoised and super-resolved velocity fields from clinical data. Our approach is quick to execute and straightforward to implement for novel cases, achieving 4D super-resolution.

eess.IV