Searcharxiv⌕ Search

arXiv subjects

Ramon van Handel

Publications and source records attributed to Ramon van Handel.

At least 37 records · Page 2Linked to original sources

The Extremals of Minkowski's Quadratic Inequality

In a seminal paper "Volumen und Oberfläche" (1903), Minkowski introduced the basic notion of mixed volumes and the corresponding inequalities that lie at the heart of convex geometry. The fundamental importance of characterizing the extremals of these inequalities was already emphasized by Minkowski himself, but has to date only been resolved in special cases. In this paper, we completely settle the extremals of Minkowski's quadratic inequality, confirming a conjecture of R. Schneider. Our proof is based on the representation of mixed volumes of arbitrary convex bodies as Dirichlet forms associated to certain highly degenerate elliptic operators. A key ingredient of the proof is a quantitative rigidity property associated to these operators.

math.MG↗

A Theory of Universal Learning

How quickly can a given class of concepts be learned from examples? It is common to measure the performance of a supervised machine learning algorithm by plotting its "learning curve", that is, the decay of the error rate as a function of the number of training examples. However, the classical theoretical framework for understanding learnability, the PAC model of Vapnik-Chervonenkis and Valiant, does not explain the behavior of learning curves: the distribution-free PAC model of learning can only bound the upper envelope of the learning curves over all possible data distributions. This does not match the practice of machine learning, where the data source is typically fixed in any given scenario, while the learner may choose the number of training examples on the basis of factors such as computational resources and desired accuracy. In this paper, we study an alternative learning model that better captures such practical aspects of machine learning, but still gives rise to a complete theory of the learnable in the spirit of the PAC model. More precisely, we consider the problem of universal learning, which aims to understand the performance of learning algorithms on every data distribution, but without requiring uniformity over the distribution. The main result of this paper is a remarkable trichotomy: there are only three possible rates of universal learning. More precisely, we show that the learning curves of any given concept class decay either at an exponential, linear, or arbitrarily slow rates. Moreover, each of these cases is completely characterized by appropriate combinatorial parameters, and we exhibit optimal learning algorithms that achieve the best possible rate in each case. For concreteness, we consider in this paper only the realizable case, though analogous results are expected to extend to more general learning scenarios.

cs.LG↗

Rademacher type and Enflo type coincide

A nonlinear analogue of the Rademacher type of a Banach space was introduced in classical work of Enflo. The key feature of Enflo type is that its definition uses only the metric structure of the Banach space, while the definition of Rademacher type relies on its linear structure. We prove that Rademacher type and Enflo type coincide, settling a long-standing open problem in Banach space theory. The proof is based on a novel dimension-free analogue of Pisier's inequality on the discrete cube.

math.FA↗

Second-Order Converses via Reverse Hypercontractivity

A strong converse shows that no procedure can beat the asymptotic (as blocklength $n\to\infty$) fundamental limit of a given information-theoretic problem for any fixed error probability. A second-order converse strengthens this conclusion by showing that the asymptotic fundamental limit cannot be exceeded by more than $O(\tfrac{1}{\sqrt{n}})$. While strong converses are achieved in a broad range of information-theoretic problems by virtue of the "blowing-up method"---a powerful methodology due to Ahlswede, Gács and Körner (1976) based on concentration of measure---this method is fundamentally unable to attain second-order converses and is restricted to finite-alphabet settings. Capitalizing on reverse hypercontractivity of Markov semigroups and functional inequalities, this paper develops the "smoothing-out" method, an alternative to the blowing-up approach that does not rely on finite alphabets and that leads to second-order converses in a variety of information-theoretic problems that were out of reach of previous methods.

cs.IT↗

Improving constant in end-point Poincaré inequality on Hamming cube

We improve the constant $\fracπ{2}$ in $L^1$-Poincaré inequality on Hamming cube. For Gaussian space the sharp constant in $L^1$ inequality is known, and it is $\sqrt{\fracπ{2}}$. For Hamming cube the sharp constant is not known, and $\sqrt{\fracπ{2}}$ gives an estimate from below for this sharp constant. On the other hand, L. Ben Efraim and F. Lust-Piquard have shown an estimate from above: $C_1\le \fracπ{2}$. There are at least two other independent proofs of the same estimate from above (we write down them in this note). Since those proofs are very different from the proof of Ben Efraim and Lust-Piquard but gave the same constant, that might have indicated that constant is sharp. But here we give a better estimate from above, showing that $C_1$ is strictly smaller than $\fracπ{2}$. It is still not clear whether $C_1> \sqrt{\fracπ{2}}$. We discuss this circle of questions and the computer experiments.

math.PR↗

Mixed volumes and the Bochner method

At the heart of convex geometry lies the observation that the volume of convex bodies behaves as a polynomial. Many geometric inequalities may be expressed in terms of the coefficients of this polynomial, called mixed volumes. Among the deepest results of this theory is the Alexandrov-Fenchel inequality, which subsumes many known inequalities as special cases. The aim of this note is to give new proofs of the Alexandrov-Fenchel inequality and of its matrix counterpart, Alexandrov's inequality for mixed discriminants, that appear conceptually and technically simpler than earlier proofs and clarify the underlying structure. Our main observation is that these inequalities can be reduced by the spectral theorem to certain trivial `Bochner formulas'.

math.MG↗

The dimension-free structure of nonhomogeneous random matrices

Let $X$ be a symmetric random matrix with independent but non-identically distributed centered Gaussian entries. We show that $$ \mathbf{E}\|X\|_{S_p} \asymp \mathbf{E}\Bigg[ \Bigg(\sum_i\Bigg(\sum_j X_{ij}^2\Bigg)^{p/2}\Bigg)^{1/p} \Bigg] $$ for any $2\le p\le\infty$, where $S_p$ denotes the $p$-Schatten class and the constants are universal. The right-hand side admits an explicit expression in terms of the variances of the matrix entries. This settles, in the case $p=\infty$, a conjecture of the first author, and provides a complete characterization of the class of infinite matrices with independent Gaussian entries that define bounded operators on $\ell_2$. Along the way, we obtain optimal dimension-free bounds on the moments $(\mathbf{E}\|X\|_{S_p}^p)^{1/p}$ that are of independent interest. We develop further extensions to non-symmetric matrices and to nonasymptotic moment and norm estimates for matrices with non-Gaussian entries that arise, for example, in the study of random graphs and in applied mathematics.

math.PR↗

The equality cases of the Ehrhard-Borell inequality

The Ehrhard-Borell inequality is a far-reaching refinement of the classical Brunn-Minkowski inequality that captures the sharp convexity and isoperimetric properties of Gaussian measures. Unlike in the classical Brunn-Minkowski theory, the equality cases in this inequality are far from evident from the known proofs. The equality cases are settled systematically in this paper. An essential ingredient of the proofs are the geometric and probabilistic properties of certain degenerate parabolic equations. The method developed here serves as a model for the investigation of equality cases in a broader class of geometric inequalities that are obtained by means of a maximum principle.

math.PR↗

Structured Random Matrices

Random matrix theory is a well-developed area of probability theory that has numerous connections with other areas of mathematics and its applications. Much of the literature in this area is concerned with matrices that possess many exact or approximate symmetries, such as matrices with i.i.d. entries, for which precise analytic results and limit theorems are available. Much less well understood are matrices that are endowed with an arbitrary structure, such as sparse Wigner matrices or matrices whose entries possess a given variance pattern. The challenge in investigating such structured random matrices is to understand how the given structure of the matrix is reflected in its spectral properties. This chapter reviews a number of recent results, methods, and open problems in this direction, with a particular emphasis on sharp spectral norm inequalities for Gaussian random matrices.

math.PR↗

Chaining, Interpolation, and Convexity II: The contraction principle

The generic chaining method provides a sharp description of the suprema of many random processes in terms of the geometry of their index sets. The chaining functionals that arise in this theory are however notoriously difficult to control in any given situation. In the first paper in this series, we introduced a particularly simple method for producing the requisite multi scale geometry by means of real interpolation. This method is easy to use, but does not always yield sharp bounds on chaining functionals. In the present paper, we show that a refinement of the interpolation method provides a canonical mechanism for controlling chaining functionals. The key innovation is a simple but powerful contraction principle that makes it possible to efficiently exploit interpolation. We illustrate the utility of this approach by developing new dimension-free bounds on the norms of random matrices and on chaining functionals in Banach lattices. As another application, we give a remarkably short interpolation proof of the majorizing measure theorem that entirely avoids the greedy construction that lies at the heart of earlier proofs.

math.PR↗

Sharp nonasymptotic bounds on the norm of random matrices with independent entries

We obtain nonasymptotic bounds on the spectral norm of random matrices with independent entries that improve significantly on earlier results. If $X$ is the $n\times n$ symmetric matrix with $X_{ij}\sim N(0,b_{ij}^2)$, we show that \[\mathbf{E}\Vert X\Vert \lesssim\max_i\sqrt{\sum_jb_{ij}^2}+\max _{ij}\vert b_{ij}\vert \sqrt{\log n}.\] This bound is optimal in the sense that a matching lower bound holds under mild assumptions, and the constants are sufficiently sharp that we can often capture the precise edge of the spectrum. Analogous results are obtained for rectangular matrices and for more general sub-Gaussian or heavy-tailed distributions of the entries, and we derive tail bounds in addition to bounds on the expected norm. The proofs are based on a combination of the moment method and geometric functional analysis techniques. As an application, we show that our bounds immediately yield the correct phase transition behavior of the spectral edge of random band matrices and of sparse Wigner matrices. We also recover a result of Seginer on the norm of Rademacher matrices.

math.PR↗

The Borell-Ehrhard Game

A precise description of the convexity of Gaussian measures is provided by sharp Brunn-Minkowski type inequalities due to Ehrhard and Borell. We show that these are manifestations of a game-theoretic mechanism: a minimax variational principle for Brownian motion. As an application, we obtain a Gaussian improvement of Barthe's reverse Brascamp-Lieb inequality.

math.PR↗

Chaining, Interpolation, and Convexity

We show that classical chaining bounds on the suprema of random processes in terms of entropy numbers can be systematically improved when the underlying set is convex: the entropy numbers need not be computed for the entire set, but only for certain "thin" subsets. This phenomenon arises from the observation that real interpolation can be used as a natural chaining mechanism. Unlike the general form of Talagrand's generic chaining method, which is sharp but often difficult to use, the resulting bounds involve only entropy numbers but are nonetheless sharp in many situations in which classical entropy bounds are suboptimal. Such bounds are readily amenable to explicit computations in specific examples, and we discover some old and new geometric principles for the control of chaining functionals as special cases.

math.PR↗

On the spectral norm of Gaussian random matrices

Let $X$ be a $d\times d$ symmetric random matrix with independent but non-identically distributed Gaussian entries. It has been conjectured by Latał{a} that the spectral norm of $X$ is always of the same order as the largest Euclidean norm of its rows. A positive resolution of this conjecture would provide a sharp understanding of the probabilistic mechanisms that control the spectral norm of inhomogeneous Gaussian random matrices. This paper establishes the conjecture up to a dimensional factor of order $\sqrt{\log\log d}$. Moreover, dimension-free bounds are developed that are optimal to leading order and that establish the conjecture in special cases. The proofs of these results shed significant light on the geometry of the underlying Gaussian processes.

math.PR↗

Can local particle filters beat the curse of dimensionality?

The discovery of particle filtering methods has enabled the use of nonlinear filtering in a wide array of applications. Unfortunately, the approximation error of particle filters typically grows exponentially in the dimension of the underlying model. This phenomenon has rendered particle filters of limited use in complex data assimilation problems. In this paper, we argue that it is often possible, at least in principle, to develop local particle filtering algorithms whose approximation error is dimension-free. The key to such developments is the decay of correlations property, which is a spatial counterpart of the much better understood stability property of nonlinear filters. For the simplest possible algorithm of this type, our results provide under suitable assumptions an approximation error bound that is uniform both in time and in the model dimension. More broadly, our results provide a framework for the investigation of filtering problems and algorithms in high dimension.

math.ST↗

Phase Transitions in Nonlinear Filtering

It has been established under very general conditions that the ergodic properties of Markov processes are inherited by their conditional distributions given partial information. While the existing theory provides a rather complete picture of classical filtering models, many infinite-dimensional problems are outside its scope. Far from being a technical issue, the infinite-dimensional setting gives rise to surprising phenomena and new questions in filtering theory. The aim of this paper is to discuss some elementary examples, conjectures, and general theory that arise in this setting, and to highlight connections with problems in statistical mechanics and ergodic theory. In particular, we exhibit a simple example of a uniformly ergodic model in which ergodicity of the filter undergoes a phase transition, and we develop some qualitative understanding as to when such phenomena can and cannot occur. We also discuss closely related problems in the setting of conditional Markov random fields.

math.PR↗

Conditional ergodicity in infinite dimension

The goal of this paper is to develop a general method to establish conditional ergodicity of infinite-dimensional Markov chains. Given a Markov chain in a product space, we aim to understand the ergodic properties of its conditional distributions given one of the components. Such questions play a fundamental role in the ergodic theory of nonlinear filters. In the setting of Harris chains, conditional ergodicity has been established under general nondegeneracy assumptions. Unfortunately, Markov chains in infinite-dimensional state spaces are rarely amenable to the classical theory of Harris chains due to the singularity of their transition probabilities, while topological and functional methods that have been developed in the ergodic theory of infinite-dimensional Markov chains are not well suited to the investigation of conditional distributions. We must therefore develop new measure-theoretic tools in the ergodic theory of Markov chains that enable the investigation of conditional ergodicity for infinite dimensional or weak-* ergodic processes. To this end, we first develop local counterparts of zero-two laws that arise in the theory of Harris chains. These results give rise to ergodic theorems for Markov chains that admit asymptotic couplings or that are locally mixing in the sense of H. Föllmer, and to a non-Markovian ergodic theorem for stationary absolutely regular sequences. We proceed to show that local ergodicity is inherited by conditioning on a nondegenerate observation process. This is used to prove stability and unique ergodicity of the nonlinear filter. Finally, we show that our abstract results can be applied to infinite-dimensional Markov processes that arise in several settings, including dissipative stochastic partial differential equations, stochastic spin systems and stochastic differential delay equations.

math.PR↗

Comparison Theorems for Gibbs Measures

The Dobrushin comparison theorem is a powerful tool to bound the difference between the marginals of high-dimensional probability distributions in terms of their local specifications. Originally introduced to prove uniqueness and decay of correlations of Gibbs measures, it has been widely used in statistical mechanics as well as in the analysis of algorithms on random fields and interacting Markov chains. However, the classical comparison theorem requires validity of the Dobrushin uniqueness criterion, essentially restricting its applicability in most models to a small subset of the natural parameter space. In this paper we develop generalized Dobrushin comparison theorems in terms of influences between blocks of sites, in the spirit of Dobrushin-Shlosman and Weitz, that substantially extend the range of applicability of the classical comparison theorem. Our proofs are based on the analysis of an associated family of Markov chains. We develop in detail an application of our main results to the analysis of sequential Monte Carlo algorithms for filtering in high dimension.

math.PR↗