SearcharxivSearch

arXiv subjects

Esteban G. Tabak

Publications and source records attributed to Esteban G. Tabak.

At least 19 recordsLinked to original sources

Dynamical Clustering and its Limiting Differential Equations

A novel, non-parametric clustering algorithm is developed, based on flows in the space of soft assignments $P_k^i$, representing the probability that point $x^i$ belongs to class $k$. These flows have two drivers: a nonlinear reaction component that relaxes $P$ to an evolving Bayesian-like posterior, and a diffusive component that relaxes $P$ to its local average, with a diffusivity $ν$ that adapts dynamically so as to balance the two components. The methodology extends to semi-supervised classification, where a subset of the labels are known. In the limit of infinitely many observations, it yields a set of non-standard reaction-diffusion equations, which produces sharp boundaries between species that separate spontaneously into well-balanced domains. Including more than one diffusive network opens the way to a broader class of applications, including the detection of regime changes in time series and a clustering procedure based on numerous features that defeats the curse of dimensionality. In the continuous limit, this extension results in a novel, non-local class of diffusive operators.

math.ST

The Monge optimal transport barycenter problem

A novel methodology is developed for the solution of the data-driven Monge optimal transport barycenter problem, where the pushforward condition is formulated in terms of the statistical independence between two sets of random variables: the factors $z$ and a transformed outcome $y$. Relaxing independence to the uncorrelation between all functions of $z$ and $y$ within suitable finite-dimensional spaces leads to an adversarial formulation, for which the adversarial strategy can be found in closed form through the first principal components of a small-dimensional matrix. The resulting pure minimization problem can be solved very efficiently through gradient descent driven flows in phase space. The methodology extends beyond scenarios where only discrete factors affect the outcome, to multivariate sets of both discrete and continuous factors, for which the corresponding barycenter problems have infinitely many marginals. Corollaries include a new framework for the solution of the Monge optimal transport problem, a procedure for the data-based simulation and estimation of conditional probability densities, and a nonparametric methodology for Bayesian inference.

math.OC

Constrained Density Estimation via Optimal Transport

A novel framework for density estimation under expectation constraints is proposed. The framework minimizes the Wasserstein distance between the estimated density and a prior, subject to the constraints that the expected value of a set of functions adopts or exceeds given values. The framework is generalized to include regularization inequalities to mitigate the artifacts in the target measure. An annealing-like algorithm is developed to address non-smooth constraints, with its effectiveness demonstrated through both synthetic and proof-of-concept real world examples in finance.

stat.ML

Optimal transport with a density-dependent cost function

A new pairwise cost function is proposed for the optimal transport barycenter problem, adopting the form of the minimal action between two points, with a Lagrangian that takes into account an underlying probability distribution. Under this notion of distance, two points can only be close if there exist paths joining them that do not traverse areas of small probability. A framework is proposed and developed for the numerical solution of the corresponding data-driven optimal transport problem. The procedure parameterizes the paths of minimal action through path dependent Chebyshev polynomials and enforces the agreement between the paths' endpoints and the given source and target distributions through an adversarial penalization. The methodology and its application to clustering and matching problems is illustrated through synthetic examples.

stat.CO

The hierarchical barycenter: conditional probability simulation with structured and unobserved covariates

This paper presents a new method for conditional probability density simulation. The method is design to work with unstructured data set when data are not characterized by the same covariates yet share common information. Specific examples considered in the text are relative to two main classes: homogeneous data characterized by samples with missing value for the covariates and data set divided in two or more groups characterized by covariates that are only partially overlapping. The methodology is based on the mathematical theory of optimal transport extending the barycenter problem to the newly defined hierarchical barycenter problem. A newly, data driven, numerical procedure for the solution of the hierarchical barycenter problem is proposed and its advantages, over the use of classical barycenter, are illustrated on synthetic and real world data sets.

stat.ME

A new approach to data assimilation initialization problems with sparse data using multiple cost functions

This article develops a novel data assimilation methodology, addressing challenges that are common in real-world settings, such as severe sparsity of observations, lack of reliable models, and non-stationarity of the system dynamics. These challenges often cause identifiability issues and can confound model parameter initialization, both of which can lead to estimated models with unrealistic qualitative dynamics and induce deeper parameter estimation errors. The proposed methodology's objective function is constructed as a sum of components, each serving a different purpose: enforcing point-wise and distribution-wise agreement between data and model output, enforcing agreement of variables and parameters with a model provided, and penalizing unrealistic rapid parameter changes, unless they are due to external drivers or interventions. This methodology was motivated by, developed and evaluated in the context of estimating blood glucose levels in different medical settings. Both simulated and real data are used to evaluate the methodology from different perspectives, such as its ability to estimate unmeasured variables, its ability to reproduce the correct qualitative blood glucose dynamics, how it manages known non-stationarity, and how it performs when given a range of dense and severely sparse data. The results show that a multicomponent cost function can balance the minimization of point-wise errors with global properties, robustly preserving correct qualitative dynamics and managing data sparsity.

math.OC

Distributional barycenter problem through data-driven flows

A new method is proposed for the solution of the data-driven optimal transport barycenter problem and of the more general distributional barycenter problem that the article introduces. The method improves on previous approaches based on adversarial games, by slaving the discriminator to the generator, minimizing the need for parameterizations and by allowing the adoption of general cost functions. It is applied to numerical examples, which include analyzing the MNIST data set with a new cost function that penalizes non-isometric maps.

math.OC

Clustering, factor discovery and optimal transport

The clustering problem, and more generally, latent factor discovery --or latent space inference-- is formulated in terms of the Wasserstein barycenter problem from optimal transport. The objective proposed is the maximization of the variability attributable to class, further characterized as the minimization of the variance of the Wasserstein barycenter. Existing theory, which constrains the transport maps to rigid translations, is extended to affine transformations. The resulting non-parametric clustering algorithms include k-means as a special case and exhibit more robust performance. A continuous version of these algorithms discovers continuous latent variables and generalizes principal curves. The strength of these algorithms is demonstrated by tests on both artificial and real-world data sets.

math.OC

Conditional Density Estimation, Latent Variable Discovery and Optimal Transport

A framework is proposed that addresses both conditional density estimation and latent variable discovery. The objective function maximizes explanation of variability in the data, achieved through the optimal transport barycenter generalized to a collection of conditional distributions indexed by a covariate --either given or latent-- in any suitable space. Theoretical results establish the existence of barycenters, a minimax formulation of optimal transport maps, and a general characterization of variability via the optimal transport cost. This framework leads to a family of non-parametric neural network-based algorithms, the BaryNet, with a supervised version that estimates conditional distributions and an unsupervised version that assigns latent variables. The efficacy of BaryNets is demonstrated by tests on both artificial and real-world data sets. A parallel drawn between autoencoders and the barycenter framework leads to the Barycentric autoencoder algorithm (BAE).

math.OC

Adversarial Optimal Transport Through The Convolution Of Kernels With Evolving Measures

A novel algorithm is proposed to solve the sample-based optimal transport problem. An adversarial formulation of the push-forward condition uses a test function built as a convolution between an adaptive kernel and an evolving probability distribution $ν$ over a latent variable $b$. Approximating this convolution by its simulation over evolving samples $b^i(t)$ of $ν$, the parameterization of the test function reduces to determining the flow of these samples. This flow, discretized over discrete time steps $t_n$, is built from the composition of elementary maps. The optimal transport also follows a flow that, by duality, must follow the gradient of the test function. The representation of the test function as the Monte Carlo simulation of a distribution makes the algorithm robust to dimensionality, and its evolution under a memory-less flow produces rich, complex maps from simple parametric transformations. The algorithm is illustrated with numerical examples.

stat.ML

Time-Series Analysis via Low-Rank Matrix Factorization Applied to Infant-Sleep Data

We propose a nonparametric model for time series with missing data based on low-rank matrix factorization. The model expresses each instance in a set of time series as a linear combination of a small number of shared basis functions. Constraining the functions and the corresponding coefficients to be nonnegative yields an interpretable low-dimensional representation of the data. A time-smoothing regularization term ensures that the model captures meaningful trends in the data, instead of overfitting short-term fluctuations. The low-dimensional representation makes it possible to detect outliers and cluster the time series according to the interpretable features extracted by the model, and also to perform forecasting via kernel regression. We apply our methodology to a large real-world dataset of infant-sleep data gathered by caregivers with a mobile-phone app. Our analysis automatically extracts daily-sleep patterns consistent with the existing literature. This allows us to compute sleep-development trends for the cohort, which characterize the emergence of circadian sleep and different napping habits. We apply our methodology to detect anomalous individuals, to cluster the cohort into groups with different sleeping tendencies, and to obtain improved predictions of future sleep behavior.

stat.ML

Data Driven Conditional Optimal Transport

A data driven procedure is developed to compute the optimal map between two conditional probabilities $ρ(x|z_{1},...,z_{L})$ and $μ(y|z_{1},...,z_{L})$ depending on a set of covariates $z_{i}$. The procedure is tested on synthetic data from the ACIC Data Analysis Challenge 2017 and it is applied to non uniform lightness transfer between images. Exactly solvable examples and simulations are performed to highlight the differences with ordinary optimal transport.

math.OC

Adaptive Optimal Transport

An adaptive, adversarial methodology is developed for the optimal transport problem between two distributions $μ$ and $ν$, known only through a finite set of independent samples $(x_i)_{i=1..N}$ and $(y_j)_{j=1..M}$. The methodology automatically creates features that adapt to the data, thus avoiding reliance on a priori knowledge of data distribution. Specifically, instead of a discrete point-bypoint assignment, the new procedure seeks an optimal map $T(x)$ defined for all $x$, minimizing the Kullback-Leibler divergence between $(T(xi))$ and the target $(y_j)$. The relative entropy is given a sample-based, variational characterization, thereby creating an adversarial setting: as one player seeks to push forward one distribution to the other, the second player develops features that focus on those areas where the two distributions fail to match. The procedure solves local problems matching consecutive, intermediate distributions between $μ$ and $ν$. As a result, maps of arbitrary complexity can be built by composing the simple maps used for each local problem. Displaced interpolation is used to guarantee global from local optimality. The procedure is illustrated through synthetic examples in one and two dimensions.

math.OC

A data-driven linear-programming methodology for optimal transport

A data-driven formulation of the optimal transport problem is presented and solved using adaptively refined meshes to decompose the problem into a sequence of finite linear programming problems. Both the marginal distributions and their unknown optimal coupling are approximated through mixtures, which decouples the problem into the the optimal transport between the individual components of the mixtures and a classical assignment problem linking them all. A factorization of the components into products of single-variable distributions makes the first sub-problem solvable in closed form. The size of the assignment problem is addressed through an adaptive procedure: a sequence of linear programming problems which utilize at each level the solution from the previous coarser mesh to restrict the size of the function space where solutions are sought. The linear programming approach for pairwise optimal transportation, combined with an iterative scheme, gives a data driven algorithm for the Wasserstein barycenter problem, which is well suited to parallel computing.

math.NA

Prototypal Analysis and Prototypal Regression

Prototypal analysis is introduced to overcome two shortcomings of archetypal analysis: its sensitivity to outliers and its non-locality, which reduces its applicability as a learning tool. Same as archetypal analysis, prototypal analysis finds prototypes through convex combination of the data points and approximates the data through convex combination of the archetypes, but it adds a penalty for using prototypes distant from the data points for their reconstruction. Prototypal analysis can be extended---via kernel embedding---to probability distributions, since the convexity of the prototypes makes them interpretable as mixtures. Finally, prototypal regression is developed, a robust supervised procedure which allows the use of distributions as either features or labels.

stat.ML

Principal dynamical components

A new procedure is proposed for the dimensional reduction of time series. Similarly to principal components, the procedure seeks a low-dimensional manifold that minimizes information loss. Unlike principal components, however, the new procedure involves dynamical considerations, through the proposal of a predictive dynamical model in the reduced manifold. Hence the minimization of the uncertainty is not only over the choice of a reduced manifold, as in principal components, but also over the parameters of the dynamical model. Further generalizations are provided to non-autonomous and non-Markovian scenarios, which are then applied to historical sea-surface temperature data.

math.ST

Oceanic Internal Wave Field: Theory of Scale-invariant Spectra

Steady scale-invariant solutions of a kinetic equation describing the statistics of oceanic internal gravity waves based on wave turbulence theory are investigated. It is shown in the non-rotating scale-invariant limit that the collision integral in the kinetic equation diverges for almost all spectral power-law exponents. These divergences come from resonant interactions with the smallest horizontal wavenumbers and/or the largest horizontal wavenumbers with extreme scale-separations. We identify a small domain in which the scale-invariant collision integral converges and numerically find a convergent power-law solution. This numerical solution is close to the Garrett--Munk spectrum. Power-law exponents which potentially permit a balance between the infra-red and ultra-violet divergences are investigated. The balanced exponents are generalizations of an exact solution of the scale-invariant kinetic equation, the Pelinovsky--Raevsky spectrum. A balance between oppositely signed divergences states that infinity minus infinity may be approximately equal to zero. A small but finite Coriolis parameter representing the effects of rotation is introduced into the kinetic equation to determine solutions over the divergent part of the domain using rigorous asymptotic arguments. This gives rise to the induced diffusion regime. The derivation of the kinetic equation is based on an assumption of weak nonlinearity. Dominance of the nonlocal interactions puts the self-consistency of the kinetic equation at risk. Yet these weakly nonlinear stationary states are consistent with much of the observational evidence.

math-ph

Energy spectra of the ocean's internal wave field: theory and observations

The high-frequency limit of the Garrett and Munk spectrum of internal waves in the ocean and the observed deviations from it are shown to form a pattern consistent with the predictions of wave turbulence theory. In particular, the high frequency limit of the Garrett and Munk spectrum constitutes an {\it exact} steady state solution of the corresponding kinetic equation.

math-ph