SearcharxivSearch

arXiv subjects

Ronald R. Coifman

Publications and source records attributed to Ronald R. Coifman.

At least 19 recordsLinked to original sources

Regularization and compression of integral operator kernels with tensor paraproducts

We present a new perspective on integral operator kernels given by a nonlinear function of the distance. This perspective is useful since it permits one to quasilinearize the nonlinear transformations acting on the tensor Haar expansion of the distance, d(x,y). This leads to a new representation comprised of a principal and residual term with desirable features; namely, enhanced local regularity of the principal component (and in certain situations rapid off-diagonal decay leading to O(N) storage) and a residual expansion of improved approximation as a consequence of enhanced global regularity. Numerical experiments on the potential kernel are included

math.NA

Blaschke products and unwinding in higher dimensions

We give a necessary and sufficient condition for the convergence of an infinite product of rational inner functions on the polydisk, and explore generalization to the polydisk of Malmquist- Takenaka bases and various versions of unwinding

math.CV

Extracting Dual Analytic Geometries of Linear Transformations to Achieve Efficient Computation

We propose a novel framework for fast integral operations by uncovering hidden geometries in the row and column structures of the underlying operators. This is accomplished through the \texttt{Questionnaire} algorithm, an iterative procedure that constructs adaptive hierarchical partition trees, revealing latent multiscale organizations and exposing local low-rank structures within the data. Guided by these geometries, we employ two complementary techniques: (1) The \texttt{\texttt{Butterfly}} algorithm, which exploits the learned hierarchical low-rank structure; and (2) Adaptive \texttt{eGHWT}, best tilings in both space and frequency using all levels of the generalized Haar--Walsh wavelet packets. These techniques enable efficient matrix factorization and multiplication. We coin our algorithms as \texttt{Questionnaire Factorization and Fast Transform (QFFT)}. Unlike classical approaches that rely on prior knowledge of the underlying geometry, \texttt{QFFT} is fully data-driven and applicable to matrices arising from irregular or unknown distributions. Even when the rows and columns both appear mutually orthogonal, our framework identifies the intrinsic ordering of orthogonal vectors that reveal hidden sparsity of the kernel. We demonstrate the effectiveness of our approach on matrices associated with heterogeneous operators and families of orthogonal polynomials. The resulting compressed representations reduce storage complexity from $\mathcal{O}(N^2)$ to $\mathcal{O}(N \log N)$, enabling fast computation and scalable implementation.

math.NA

On the Practical Use of Blaschke Decomposition in Nonstationary Signal Analysis

The Blaschke decomposition-based algorithm, {\em Phase Dynamics Unwinding} (PDU), possesses several attractive theoretical properties, including fast convergence, effective decomposition, and multiscale analysis. However, its application to real-world signal decomposition tasks encounters notable challenges. In this work, we propose two techniques, divide-and-conquer via tapering and cumulative summation (cumsum), to handle complex trends and amplitude modulations and the mode-mixing caused by winding. The resulting method, termed {\em windowed PDU}, enhances PDU's performance in practical decomposition tasks. We validate our approach through both simulated and real-world signals, demonstrating its effectiveness across diverse scenarios.

stat.ME

Joint Hierarchical Representation Learning of Samples and Features via Informed Tree-Wasserstein Distance

High-dimensional data often exhibit hierarchical structures in both modes: samples and features. Yet, most existing approaches for hierarchical representation learning consider only one mode at a time. In this work, we propose an unsupervised method for jointly learning hierarchical representations of samples and features via Tree-Wasserstein Distance (TWD). Our method alternates between the two data modes. It first constructs a tree for one mode, then computes a TWD for the other mode based on that tree, and finally uses the resulting TWD to build the second mode's tree. By repeatedly alternating through these steps, the method gradually refines both trees and the corresponding TWDs, capturing meaningful hierarchical representations of the data. We provide a theoretical analysis showing that our method converges. We show that our method can be integrated into hyperbolic graph convolutional networks as a pre-processing technique, improving performance in link prediction and node classification tasks. In addition, our method outperforms baselines in sparse approximation and unsupervised Wasserstein distance learning tasks on word-document and single-cell RNA-sequencing datasets.

cs.LG

Intrinsic and Extrinsic Organized Attention: Softmax Invariance and Network Sparsity

We examine the intrinsic (within the attention head) and extrinsic (amongst the attention heads) structure of the self-attention mechanism in transformers. Theoretical evidence for invariance of the self-attention mechanism to softmax activation is obtained by appealing to paradifferential calculus, (and is supported by computational examples), which relies on the intrinsic organization of the attention heads. Furthermore, we use an existing methodology for hierarchical organization of tensors to examine network structure by constructing hierarchal partition trees with respect to the query, key, and head axes of network 3-tensors. Such an organization is consequential since it allows one to profitably execute common signal processing tasks on a geometry where the organized network 3-tensors exhibit regularity. We exemplify this qualitatively, by visualizing the hierarchical organization of the tree comprised of attention heads and the diffusion map embeddings, and quantitatively by investigating network sparsity with the expansion coefficients of individual attention heads and the entire network with respect to the bi and tri-haar bases (respectively) on the space of queries, keys, and heads of the network. To showcase the utility of our theoretical and methodological findings, we provide computational examples using vision and language transformers. The ramifications of these findings are two-fold: (1) a subsequent step in interpretability analysis is theoretically admitted, and can be exploited empirically for downstream interpretability tasks (2) one can use the network 3-tensor organization for empirical network applications such as model pruning (by virtue of network sparsity) and network architecture comparison.

math.NA

Tree-Wasserstein Distance for High Dimensional Data with a Latent Feature Hierarchy

Finding meaningful distances between high-dimensional data samples is an important scientific task. To this end, we propose a new tree-Wasserstein distance (TWD) for high-dimensional data with two key aspects. First, our TWD is specifically designed for data with a latent feature hierarchy, i.e., the features lie in a hierarchical space, in contrast to the usual focus on embedding samples in hyperbolic space. Second, while the conventional use of TWD is to speed up the computation of the Wasserstein distance, we use its inherent tree as a means to learn the latent feature hierarchy. The key idea of our method is to embed the features into a multi-scale hyperbolic space using diffusion geometry and then present a new tree decoding method by establishing analogies between the hyperbolic embedding and trees. We show that our TWD computed based on data observations provably recovers the TWD defined with the latent feature hierarchy and that its computation is efficient and scalable. We showcase the usefulness of the proposed TWD in applications to word-document and single-cell RNA-sequencing datasets, demonstrating its advantages over existing TWDs and methods based on pre-trained models.

cs.LG

On Learning what to Learn: heterogeneous observations of dynamics and establishing (possibly causal) relations among them

Before we attempt to learn a function between two (sets of) observables of a physical process, we must first decide what the inputs and what the outputs of the desired function are going to be. Here we demonstrate two distinct, data-driven ways of initially deciding ``the right quantities'' to relate through such a function, and then proceed to learn it. This is accomplished by processing multiple simultaneous heterogeneous data streams (ensembles of time series) from observations of a physical system: multiple observation processes of the system. We thus determine (a) what subsets of observables are common between the observation processes (and therefore observable from each other, relatable through a function); and (b) what information is unrelated to these common observables, and therefore particular to each observation process, and not contributing to the desired function. Any data-driven function approximation technique can subsequently be used to learn the input-output relation, from k-nearest neighbors and Geometric Harmonics to Gaussian Processes and Neural Networks. Two particular ``twists'' of the approach are discussed. The first has to do with the identifiability of particular quantities of interest from the measurements. We now construct mappings from a single set of observations of one process to entire level sets of measurements of the process, consistent with this single set. The second attempts to relate our framework to a form of causality: if one of the observation processes measures ``now'', while the second observation process measures ``in the future'', the function to be learned among what is common across observation processes constitutes a dynamical model for the system evolution.

cs.LG

Gappy local conformal auto-encoders for heterogeneous data fusion: in praise of rigidity

Fusing measurements from multiple, heterogeneous, partial sources, observing a common object or process, poses challenges due to the increasing availability of numbers and types of sensors. In this work we propose, implement and validate an end-to-end computational pipeline in the form of a multiple-auto-encoder neural network architecture for this task. The inputs to the pipeline are several sets of partial observations, and the result is a globally consistent latent space, harmonizing (rigidifying, fusing) all measurements. The key enabler is the availability of multiple slightly perturbed measurements of each instance:, local measurement, "bursts", that allows us to estimate the local distortion induced by each instrument. We demonstrate the approach in a sequence of examples, starting with simple two-dimensional data sets and proceeding to a Wi-Fi localization problem and to the solution of a "dynamical puzzle" arising in spatio-temporal observations of the solutions of Partial Differential Equations.

cs.LG

Hyperbolic Diffusion Embedding and Distance for Hierarchical Representation Learning

Finding meaningful representations and distances of hierarchical data is important in many fields. This paper presents a new method for hierarchical data embedding and distance. Our method relies on combining diffusion geometry, a central approach to manifold learning, and hyperbolic geometry. Specifically, using diffusion geometry, we build multi-scale densities on the data, aimed to reveal their hierarchical structure, and then embed them into a product of hyperbolic spaces. We show theoretically that our embedding and distance recover the underlying hierarchical structure. In addition, we demonstrate the efficacy of the proposed method and its advantages compared to existing methods on graph embedding benchmarks and hierarchical datasets.

cs.LG

On Complex Analytic tools, and the Holomorphic Rotation methods

We describe recent nonlinear analytic approximation tools in the classical setting of Hardy spaces in the upper half plane and show how to transfer them to the higher dimensional real setting of harmonic functions in upper half spaces. It is known [6] that all harmonic functions in higher dimensions are combinations of holomorphic functions on 2 dimensional planes, extended as, constant in normal directions. We derive representation theorems, with corresponding isometries, opening the door for applications in higher dimensions, to the processing of highly oscillatory multidimensional signals.

math.CA

Questionnaires to PDEs: From Disorganized Data to Emergent Generative Dynamic Models

Starting with sets of disorganized observations of spatially varying and temporally evolving systems, obtained at different (also disorganized) sets of parameters, we demonstrate the data-driven derivation of parameter dependent, evolutionary partial differential equation (PDE) models capable of generating the data. This tensor type of data is reminiscent of shuffled (multi-dimensional) puzzle tiles. The independent variables for the evolution equations (their "space" and "time") as well as their effective parameters are all "emergent", i.e., determined in a data-driven way from our disorganized observations of behavior in them. We use a diffusion map based "questionnaire" approach to build a parametrization of our emergent space/time/parameter space for the data. This approach iteratively processes the data by successively observing them on the "space", the "time", and the "parameter" axes of a tensor. Once the data are organized, we use machine learning (here, neural networks) to approximate the operators governing the evolution equations in this emergent space. Our illustrative example is based on a previously developed vertex-plus-signaling model of Drosophila embryonic development. This allows us to discuss features of the process like symmetry breaking, translational invariance, and autonomousness of the emergent PDE model, as well as its interpretability.

math.DS

A common variable minimax theorem for graphs

Let $\mathcal{G} = \{G_1 = (V, E_1), \dots, G_m = (V, E_m)\}$ be a collection of $m$ graphs defined on a common set of vertices $V$ but with different edge sets $E_1, \dots, E_m$. Informally, a function $f :V \rightarrow \mathbb{R}$ is smooth with respect to $G_k = (V,E_k)$ if $f(u) \sim f(v)$ whenever $(u, v) \in E_k$. We study the problem of understanding whether there exists a nonconstant function that is smooth with respect to all graphs in $\mathcal{G}$, simultaneously, and how to find it if it exists.

math.SP

Doubly-Stochastic Normalization of the Gaussian Kernel is Robust to Heteroskedastic Noise

A fundamental step in many data-analysis techniques is the construction of an affinity matrix describing similarities between data points. When the data points reside in Euclidean space, a widespread approach is to from an affinity matrix by the Gaussian kernel with pairwise distances, and to follow with a certain normalization (e.g. the row-stochastic normalization or its symmetric variant). We demonstrate that the doubly-stochastic normalization of the Gaussian kernel with zero main diagonal (i.e., no self loops) is robust to heteroskedastic noise. That is, the doubly-stochastic normalization is advantageous in that it automatically accounts for observations with different noise variances. Specifically, we prove that in a suitable high-dimensional setting where heteroskedastic noise does not concentrate too much in any particular direction in space, the resulting (doubly-stochastic) noisy affinity matrix converges to its clean counterpart with rate $m^{-1/2}$, where $m$ is the ambient dimension. We demonstrate this result numerically, and show that in contrast, the popular row-stochastic and symmetric normalizations behave unfavorably under heteroskedastic noise. Furthermore, we provide examples of simulated and experimental single-cell RNA sequence data with intrinsic heteroskedasticity, where the advantage of the doubly-stochastic normalization for exploratory analysis is evident.

stat.ML

LOCA: LOcal Conformal Autoencoder for standardized data coordinates

We propose a deep-learning based method for obtaining standardized data coordinates from scientific measurements.Data observations are modeled as samples from an unknown, non-linear deformation of an underlying Riemannian manifold, which is parametrized by a few normalized latent variables. By leveraging a repeated measurement sampling strategy, we present a method for learning an embedding in $\mathbb{R}^d$ that is isometric to the latent variables of the manifold. These data coordinates, being invariant under smooth changes of variables, enable matching between different instrumental observations of the same phenomenon. Our embedding is obtained using a LOcal Conformal Autoencoder (LOCA), an algorithm that constructs an embedding to rectify deformations by using a local z-scoring procedure while preserving relevant geometric information. We demonstrate the isometric embedding properties of LOCA on various model settings and observe that it exhibits promising interpolation and extrapolation capabilities. Finally, we apply LOCA to single-site Wi-Fi localization data, and to $3$-dimensional curved surface estimation based on a $2$-dimensional projection.

cs.LG

Multiscale decompositions of Hardy spaces

An inspiration at the origin of wavelet analysis (when Grossmann, Morlet, Meyer and collaborators were interacting and exploring versions of multiscale representations) was provided by the analysis of holomorphic signals, for which the images of the phase of Cauchy wavelets were remarkable in their ability to reveal intricate singularities or dynamic structures, such as instantaneous frequency jumps in musical recordings. Our goal is to follow their seminal work and introduce recent developments in nonlinear analysis. In particular we sketch methods extending conventional Fourier analysis, exploiting both phase and amplitudes of holomorphic functions. The Blaschke factors are a key ingredient, in building analytic tools, starting with the Malmquist Takenaka orthonormal bases of the Hardy space, continuing with "best" adapted bases obtained through phase unwinding, and concluding with relations to composition of Blaschke products and their dynamics. We also remark that the phase of a Blaschke product is a one layer neural net with arctan as an activation sigmoid and that the composition is a "Deep Neural Net" whose depth is the number of compositions. Our results provide a wealth of related library of orthonormal bases.

math.CA

A Remark on the Arcsine Distribution and the Hilbert Transform

It is known that if $(p_n)_{n \in \mathbb{N}}$ is a sequence of orthogonal polynomials in $L^2([-1,1], w(x)dx)$, then the roots are distributed according to an arcsine distribution $π^{-1} (1-x^2)^{-1}dx$ for a wide variety of weights $w(x)$. We connect this to a result of the Hilbert transform due to Tricomi: if $f(x)(1-x^2)^{1/4} \in L^2(-1,1)$ and its Hilbert transform $Hf$ vanishes on $(-1,1)$, then the function $f$ is a multiple of the arcsine distribution $$ f(x) = \frac{c}{\sqrt{1-x^2}}χ_{(-1,1)} \qquad \mbox{where}~c~\in \mathbb{R}.$$ We also prove a localized Parseval-type identity that seems to be new: if $f(x)(1-x^2)^{1/4} \in L^2(-1,1)$ and $f(x) \sqrt{1-x^2}$ has mean value 0 on $(-1,1)$, then $$ \int_{-1}^{1}{ (Hf)(x)^2 \sqrt{1-x^2} dx} = \int_{-1}^{1}{ f(x)^2 \sqrt{1-x^2} dx}.$$

math.CA

Co-manifold learning with missing data

Representation learning is typically applied to only one mode of a data matrix, either its rows or columns. Yet in many applications, there is an underlying geometry to both the rows and the columns. We propose utilizing this coupled structure to perform co-manifold learning: uncovering the underlying geometry of both the rows and the columns of a given matrix, where we focus on a missing data setting. Our unsupervised approach consists of three components. We first solve a family of optimization problems to estimate a complete matrix at multiple scales of smoothness. We then use this collection of smooth matrix estimates to compute pairwise distances on the rows and columns based on a new multi-scale metric that implicitly introduces a coupling between the rows and the columns. Finally, we construct row and column representations from these multi-scale metrics. We demonstrate that our approach outperforms competing methods in both data visualization and clustering.

stat.ML