SearcharxivSearch

arXiv subjects

Simon Foucart

Publications and source records attributed to Simon Foucart.

At least 19 recordsLinked to original sources

Optimal Algorithms for Nonlinear Estimation with Convex Models

A linear functional of an object from a convex symmetric set can be optimally estimated, in a worst-case sense, by a linear functional of observations made on the object. This well-known fact is extended here to a nonlinear setting: other simple functionals of the object can be optimally estimated by functionals of the observations that share a similar simple structure. This is established for the maximum of several linear functionals and even for the $\ell$th largest among them. Proving the latter requires an unusual refinement of the analytical Hahn--Banach theorem. The existence results are accompanied by practical recipes relying on convex optimization to construct the desired functionals, thereby justifying the term of estimation algorithms.

math.FA

When is a subspace of $\ell_\infty^N$ isometrically isomorphic to $\ell_\infty^n$?

It is shown in this note that one can decide whether an $n$-dimensional subspace of $\ell_\infty^N$ is isometrically isomorphic to $\ell_\infty^n$ by testing a finite number of determinental inequalities. As a byproduct, an elementary proof is provided for the fact that an $n$-dimensional subspace of $\ell_\infty^N$ with projection constant equal to one must be isometrically isomorphic to $\ell_\infty^n$.

math.FA

Learning the Maximum of a H\"older Function from Inexact Data

Within the theoretical framework of Optimal Recovery, one determines in this article the {\em best} procedures to learn a quantity of interest depending on a H\"older function acquired via inexact point evaluations at fixed datasites. {\em Best} here refers to procedures minimizing worst-case errors. The elementary arguments hint at the possibility of tackling nonlinear quantities of interest, with a particular focus on the function maximum. In a local setting, i.e., for a fixed data vector, the optimal procedure (outputting the so-called Chebyshev center) is precisely described relatively to a general model of inexact evaluations. Relatively to a slightly more restricted model and in a global setting, i.e., uniformly over all data vectors, another optimal procedure is put forward, showing how to correct the natural underestimate that simply returns the data vector maximum. Jitterred data are also briefly discussed as a side product of evaluating the minimal worst-case error optimized over the datasites.

math.GM

Optimal Prediction of Multivalued Functions from Point Samples

Predicting the value of a function $f$ at a new point given its values at old points is an ubiquitous scientific endeavor, somewhat less developed when $f$ produces multiple values that depend on one another, e.g. when it outputs likelihoods or concentrations. Considering the points as fixed (not random) entities and focusing on the worst-case, this article uncovers a prediction procedure that is optimal relatively to some model-set information about $f$. When the model sets are convex, this procedure turns out to be an affine map constructed by solving a convex optimization program. The theoretical result is specified in the two practical frameworks of (reproducing kernel) Hilbert spaces and of spaces of continuous functions.

math.OC

Worst-Case Learning under a Multi-fidelity Model

Inspired by multi-fidelity methods in computer simulations, this article introduces procedures to design surrogates for the input/output relationship of a high-fidelity code. These surrogates should be learned from runs of both the high-fidelity and low-fidelity codes and be accompanied by error guarantees that are deterministic rather than stochastic. For this purpose, the article advocates a framework tied to a theory focusing on worst-case guarantees, namely Optimal Recovery. The multi-fidelity considerations triggered new theoretical results in three scenarios: the globally optimal estimation of linear functionals, the globally optimal approximation of arbitrary quantities of interest in Hilbert spaces, and their locally optimal approximation, still within Hilbert spaces. The latter scenario boils down to the determination of the Chebyshev center for the intersection of two hyperellipsoids. It is worth noting that the mathematical framework presented here, together with its possible extension, seems to be relevant in several other contexts briefly discussed.

math.NA

Least multivariate Chebyshev polynomials on diagonally determined sets

We consider a new multivariate generalization of the classical monic (univariate) Chebyshev polynomial that minimizes the uniform norm on the interval $[-1,1]$. Let $\Pi^*_n$ be the subset of polynomials of degree at most $n$ in $d$ variables, whose homogeneous part of degree $n$ has coefficients summing up to $1$. The problem is determining a polynomial in $\Pi^*_n$ with the smallest uniform norm on a domain $\Omega$, which we call a least Chebyshev polynomial (associated with $\Omega$). Our main result solves the problem for $\Omega$ belonging to a non-trivial class of sets that we call diagonally-determined, and establishes the remarkable result that a least Chebyshev polynomial can be given via the classical, univariate, Chebyshev polynomial. In particular, the solution can be independent of the dimension. Diagonally-determined domains include centered balls in $\mathbb{R}^d$ in any norm, but can be non-convex and even non-simply connected. We also introduce a computational procedure, based on semidefinite programming hierarchies, to detect if a given semi-algebraic set is diagonally-determined.

math.OC

Optimization-Aided Construction of Multivariate Chebyshev Polynomials

This article is concerned with an extension of univariate Chebyshev polynomials of the first kind to the multivariate setting, where one chases best approximants to specific monomials by polynomials of lower degree relative to the uniform norm. Exploiting the Moment-SOS hierarchy, we devise a versatile semidefinite-programming-based procedure to compute such best approximants, as well as associated signatures. Applying this procedure in three variables leads to the values of best approximation errors for all monomials up to degree six on the euclidean ball, the simplex, and the cross-polytope. Furthermore, inspired by numerical experiments, we obtain explicit expressions for Chebyshev polynomials in two cases unresolved before, namely for the monomial $x_1^2 x_2^2 x_3$ on the euclidean ball and for the monomial $x_1^2 x_2 x_3$ on the simplex.

math.OC

Radius of Information for Two Intersected Centered Hyperellipsoids and Implications in Optimal Recovery from Inaccurate Data

For objects belonging to a known model set and observed through a prescribed linear process, we aim at determining methods to recover linear quantities of these objects that are optimal from a worst-case perspective. Working in a Hilbert setting, we show that, if the model set is the intersection of two hyperellipsoids centered at the origin, then there is an optimal recovery method which is linear. It is specifically given by a constrained regularization procedure whose parameters, short of being explicit, can be precomputed by solving a semidefinite program. This general framework can be swiftly applied to several scenarios: the two-space problem, the problem of recovery from $\ell_2$-inaccurate data, and the problem of recovery from a mixture of accurate and $\ell_2$-inaccurate data. With more effort, it can also be applied to the problem of recovery from $\ell_1$-inaccurate data. For the latter, we reach the conclusion of existence of an optimal recovery method which is linear, again given by constrained regularization, under a computationally verifiable sufficient condition. Experimentally, this condition seems to hold whenever the level of $\ell_1$-inaccuracy is small enough. We also point out that, independently of the inaccuracy level, the minimal worst-case error of a linear recovery method can be found by semidefinite programming.

math.OC

Linearly Embedding Sparse Vectors from $\ell_2$ to $\ell_1$ via Deterministic Dimension-Reducing Maps

This note is concerned with deterministic constructions of $m \times N$ matrices satisfying a restricted isometry property from $\ell_2$ to $\ell_1$ on $s$-sparse vectors. Similarly to the standard ($\ell_2$ to $\ell_2$) restricted isometry property, such constructions can be found in the regime $m \asymp s^2$, at least in theory. With effectiveness of implementation in mind, two simple constructions are presented in the less pleasing but still relevant regime $m \asymp s^4$. The first one, executing a Las Vegas strategy, is quasideterministic and applies in the real setting. The second one, exploiting Golomb rulers, is explicit and applies to the complex setting. As a stepping stone, an explicit isometric embedding from $\ell_2^n(\mathbb{C})$ to $\ell_4^{cn^2}(\mathbb{C})$ is presented. Finally, the extension of the problem from sparse vectors to low-rank matrices is raised as an open question.

math.FA

S-Procedure Relaxation: a Case of Exactness Involving Chebyshev Centers

Optimal recovery is a mathematical framework for learning functions from observational data by adopting a worst-case perspective tied to model assumptions on the functions to be learned. Working in a finite-dimensional Hilbert space, we consider model assumptions based on approximability and observation inaccuracies modeled as additive errors bounded in $\ell_2$. We focus on the local recovery problem, which amounts to the determination of Chebyshev centers. Earlier work by Beck and Eldar presented a semidefinite recipe for the determination of Chebyshev centers. The result was valid in the complex setting only, but not necessarily in the real setting, since it relied on the S-procedure with two quadratic constraints, which offers a tight relaxation only in the complex setting. Our contribution consists in proving that this semidefinite recipe is exact in the real setting, too, at least in the particular instance where the quadratic constraints involve orthogonal projectors. Our argument exploits a previous work of ours, where exact Chebyshev centers were obtained in a different way. We conclude by stating some open questions and by commenting on other recent results in optimal recovery.

math.OC

Full Recovery from Point Values: an Optimal Algorithm for Chebyshev Approximability Prior

Given pointwise samples of an unknown function belonging to a certain model set, one seeks in Optimal Recovery to recover this function in a way that minimizes the worst-case error of the recovery procedure. While it is often known that such an optimal recovery procedure can be chosen to be linear, e.g. when the model set is based on approximability by a subspace of continuous functions, a construction of the procedure is rarely available. This note uncovers a practical algorithm to construct a linear optimal recovery map when the approximation space is a Chevyshev space of dimension at least three and containing the constant functions.

math.NA

Near-Optimal Estimation of Linear Functionals with Log-Concave Observation Errors

This note addresses the question of optimally estimating a linear functional of an object acquired through linear observations corrupted by random noise, where optimality pertains to a worst-case setting tied to a symmetric, convex, and closed model set containing the object. It complements the article "Statistical Estimation and Optimal Recovery" published in the Annals of Statistics in 1994. There, Donoho showed (among other things) that, for Gaussian noise, linear maps provide near-optimal estimation schemes relatively to a performance measure relevant in Statistical Estimation. Here, we advocate for a different performance measure arguably more relevant in Optimal Recovery. We show that, relatively to this new measure, linear maps still provide near-optimal estimation schemes even if the noise is merely log-concave. Our arguments, which make a connection to the deterministic noise situation and bypass properties specific to the Gaussian case, offer an alternative to parts of Donoho's proof.

math.ST

On the Optimal Recovery of Graph Signals

Learning a smooth graph signal from partially observed data is a well-studied task in graph-based machine learning. We consider this task from the perspective of optimal recovery, a mathematical framework for learning a function from observational data that adopts a worst-case perspective tied to model assumptions on the function to be learned. Earlier work in the optimal recovery literature has shown that minimizing a regularized objective produces optimal solutions for a general class of problems, but did not fully identify the regularization parameter. Our main contribution provides a way to compute regularization parameters that are optimal or near-optimal (depending on the setting), specifically for graph signal processing problems. Our results offer a new interpretation for classical optimization techniques in graph-based learning and also come with new insights for hyperparameter selection. We illustrate the potential of our methods in numerical experiments on several semi-synthetic graph signal processing datasets.

cs.LG

On the value of the fifth maximal projection constant

Let $λ(m)$ denote the maximal absolute projection constant over real $m$-dimensional subspaces. This quantity is extremely hard to determine exactly, as testified by the fact that the only known value of $λ(m)$ for $m>1$ is $λ(2)=4/3$. There is also numerical evidence indicating that $λ(3)=(1+\sqrt{5})/2$. In this paper, relying on a new construction of certain mutually unbiased equiangular tight frames, we show that $λ(5)\geq 5(11+6\sqrt{5})/59 \approx 2.06919$. This value coincides with the numerical estimation of $λ(5)$ obtained by B. L. Chalmers, thus reinforcing the belief that this is the exact value of $λ(5)$.

math.FA

The Sparsity of LASSO-type Minimizers

This note extends an attribute of the LASSO procedure to a whole class of related procedures, including square-root LASSO, square LASSO, LAD-LASSO, and an instance of generalized LASSO. Namely, under the assumption that the input matrix satisfies an $\ell_p$-restricted isometry property (which in some sense is weaker than the standard $\ell_2$-restricted isometry property assumption), it is shown that if the input vector comes from the exact measurement of a sparse vector, then the minimizer of any such LASSO-type procedure has sparsity comparable to the sparsity of the measured vector. The result remains valid in the presence of moderate measurement error when the regularization parameter is not too small.

cs.IT

On the sparsity of LASSO minimizers in sparse data recovery

We present a detailed analysis of the unconstrained $\ell_1$-weighted LASSO method for recovery of sparse data from its observation by randomly generated matrices, satisfying the Restricted Isometry Property (RIP) with constant $δ<1$, and subject to negligible measurement and compressibility errors. We prove that if the data is $k$-sparse, then the size of support of the LASSO minimizer, $s$, maintains a comparable sparsity, $s\leq C_δk$. For example, if $δ=0.7$ then $s< 11k$ and a slightly smaller $δ=0.4$ yields $s< 4k$. We also derive new $\ell_2/\ell_1$ error bounds which highlight precise dependence on $k$ and on the LASSO parameter $λ$, before the error is driven below the scale of negligible measurement/ and compressiblity errors.

cs.IT

Optimal Recovery from Inaccurate Data in Hilbert Spaces: Regularize, but what of the Parameter?

In Optimal Recovery, the task of learning a function from observational data is tackled deterministically by adopting a worst-case perspective tied to an explicit model assumption made on the functions to be learned. Working in the framework of Hilbert spaces, this article considers a model assumption based on approximability. It also incorporates observational inaccuracies modeled via additive errors bounded in $\ell_2$. Earlier works have demonstrated that regularization provide algorithms that are optimal in this situation, but did not fully identify the desired hyperparameter. This article fills the gap in both a local scenario and a global scenario. In the local scenario, which amounts to the determination of Chebyshev centers, the semidefinite recipe of Beck and Eldar (legitimately valid in the complex setting only) is complemented by a more direct approach, with the proviso that the observational functionals have orthonormal representers. In the said approach, the desired parameter is the solution to an equation that can be resolved via standard methods. In the global scenario, where linear algorithms rule, the parameter elusive in the works of Micchelli et al. is found as the byproduct of a semidefinite program. Additionally and quite surprisingly, in case of observational functionals with orthonormal representers, it is established that any regularization parameter is optimal.

math.OC

Learning from Non-Random Data in Hilbert Spaces: An Optimal Recovery Perspective

The notion of generalization in classical Statistical Learning is often attached to the postulate that data points are independent and identically distributed (IID) random variables. While relevant in many applications, this postulate may not hold in general, encouraging the development of learning frameworks that are robust to non-IID data. In this work, we consider the regression problem from an Optimal Recovery perspective. Relying on a model assumption comparable to choosing a hypothesis class, a learner aims at minimizing the worst-case error, without recourse to any probabilistic assumption on the data. We first develop a semidefinite program for calculating the worst-case error of any recovery map in finite-dimensional Hilbert spaces. Then, for any Hilbert space, we show that Optimal Recovery provides a formula which is user-friendly from an algorithmic point-of-view, as long as the hypothesis class is linear. Interestingly, this formula coincides with kernel ridgeless regression in some cases, proving that minimizing the average error and worst-case error can yield the same solution. We provide numerical experiments in support of our theoretical findings.

cs.LG