Searcharxiv⌕ Search

arXiv subjects

Lukas Silvester Barth

Publications and source records attributed to Lukas Silvester Barth.

8 recordsLinked to original sources

Fuzzy simplicial sets and their application to geometric data analysis

In this article, we expand upon the concepts introduced by David Spivak about the relationship between the category $\mathbf{UM}$ of uber metric spaces and the category $\mathbf{sFuz}$ of fuzzy simplicial sets. We show that fuzzy simplicial sets can be regarded as natural combinatorial generalizations of metric relations. Furthermore, we take inspiration from UMAP to apply the theory to manifold learning, dimension reduction and data visualization, while refining some of their constructions. We generalize the adjunction between $\mathbf{UM}$ and $\mathbf{sFuz}$, derive an explicit description of colimits in $\mathbf{UM}$, and show that $\mathbf{UM}$ can be embedded into $\mathbf{sFuz}$. Furthermore, we prove analogous results for the category of extended-pseudo metric spaces $\mathbf{EPMet}$. We also provide rigorous definitions of functors that make it possible to recursively merge sets of fuzzy simplicial sets and provide a description of the adjunctions between the category of truncated fuzzy simplicial sets and $\mathbf{sFuz}$, which we relate to persistent homology. Combining those constructions, we can show a surprising connection between the well-known dimension reduction methods UMAP and Isomap and derive an alternative algorithm, which we call IsUMap, that combines some of the strengths of both methods (source code on github). Additionally, we developed a new embedding method that allows to preserve clusters detected in the original metric space that we construct from the data. The visualization of the optimization process gives the user information both about the inner-cluster distributions in the original metric space and their inter-cluster relations. We compare our new method with UMAP, Isomap and t-SNE on a series of low- and high-dimensional datasets, demonstrate how our method improves upon them and provide explanations for observed differences.

math.AT↗

Efficient compression of neural networks and datasets

Compression and generalization are fundamentally related through Solomonoff induction and the minimum description length principle (MDL), which predict that simpler models generalize better when data arises from low-complexity distributions. In this article, we combine insights from algorithmic information theory and techniques from neural network pruning to improve model generalization by identifying the most effective data compression method. Since exact MDL optimization is intractable, we cast it as $\ell_0$ regularized learning and explain why parameter sparsity provides an effective computable approximation of model description length. To identify the best practical approach, we systematically compare and refine complementary sparse optimization methods. In particular, we improve probabilistic pruning through a procedure that does not require Monte Carlo sampling and refine smooth $\ell_0$ approximations with a binary search routine that reduces hyperparameter complexity. Across convolutional networks and transformers evaluated on image and text datasets, our refined methods improve upon their predecessors, achieve substantial model compression with minimal accuracy loss, and yield short data description lengths. Finally, we use these methods in a controlled teacher-student setting to empirically verify the prediction of Solomonoff induction that compressed models learn more sample-efficiently and generalize better.

cs.LG↗

Geometry and Dress groups with non-symmetric cost functions

A metric relation by definition is symmetric. Since many data sets are non-symmetric, in this paper we develop a systematic theory of non-symmetric cost functions. Betweenness relations play an important role. We also introduce the notion of a Dress group in the non-symmetric setting and indicate a notion of curvature.

math.MG↗

Probabilistic Foundations of Fuzzy Simplicial Sets for Nonlinear Dimensionality Reduction

Fuzzy simplicial sets have become an object of interest in dimensionality reduction and manifold learning, most prominently through their role in UMAP. However, their definition through tools from algebraic topology without a clear probabilistic interpretation detaches them from commonly used theoretical frameworks in those areas. In this work we introduce a framework that explains fuzzy simplicial sets as marginals of probability measures on simplicial sets. In particular, this perspective shows that the fuzzy weights of UMAP arise from a generative model that samples Vietoris-Rips filtrations at random scales, yielding cumulative distribution functions of pairwise distances. More generally, the framework connects fuzzy simplicial sets to probabilistic models on the face poset, clarifies the relation between Kullback-Leibler divergence and fuzzy cross-entropy in this setting, and recovers standard t-norms and t-conorms via Boolean operations on the underlying simplicial sets. We then show how new embedding methods may be derived from this framework and illustrate this on an example where we generalize UMAP using Čech filtrations with triplet sampling. In summary, this probabilistic viewpoint provides a unified probabilistic theoretical foundation for fuzzy simplicial sets, clarifies the role of UMAP within this framework, and enables the systematic derivation of new dimensionality reduction methods.

cs.LG↗

Probabilistic and nonlinear compressive sensing

We present a smooth probabilistic reformulation of $\ell_0$ regularized regression that does not require Monte Carlo sampling and allows for the computation of exact gradients, facilitating rapid convergence to local optima of the best subset selection problem. The method drastically improves convergence speed compared to similar Monte Carlo based approaches. Furthermore, we empirically demonstrate that it outperforms compressive sensing algorithms such as IHT and (Relaxed-) Lasso across a wide range of settings and signal-to-noise ratios. The implementation runs efficiently on both CPUs and GPUs and is freely available at https://github.com/L0-and-behold/probabilistic-nonlinear-cs. We also contribute to research on nonlinear generalizations of compressive sensing by investigating when parameter recovery of a nonlinear teacher network is possible through compression of a student network. Building upon theorems of Fefferman and Markel, we show theoretically that the global optimum in the infinite-data limit enforces recovery up to certain symmetries. For empirical validation, we implement a normal-form algorithm that selects a canonical representative within each symmetry class. However, while compression can help to improve test loss, we find that exact parameter recovery is not even possible up to symmetries. In particular, we observe a surprising rebound effect where teacher and student configurations initially converge but subsequently diverge despite continuous decrease in test loss. These findings indicate fundamental differences between linear and nonlinear compressive sensing.

cs.LG↗

Merging Hazy Sets with m-Schemes: A Geometric Approach to Data Visualization

Many machine learning algorithms try to visualize high dimensional metric data in 2D in such a way that the essential geometric and topological features of the data are highlighted. In this paper, we introduce a framework for aggregating dissimilarity functions that arise from locally adjusting a metric through density-aware normalization, as employed in the IsUMap method. We formalize these approaches as m-schemes, a class of methods closely related to t-norms and t-conorms in probabilistic metrics, as well as to composition laws in information theory. These m-schemes provide a flexible and theoretically grounded approach to refining distance-based embeddings.

cs.LG↗

IsUMap: Manifold Learning and Data Visualization leveraging Vietoris-Rips filtrations

This work introduces IsUMap, a novel manifold learning technique that enhances data representation by integrating aspects of UMAP and Isomap with Vietoris-Rips filtrations. We present a systematic and detailed construction of a metric representation for locally distorted metric spaces that captures complex data structures more accurately than the previous schemes. Our approach addresses limitations in existing methods by accommodating non-uniform data distributions and intricate local geometries. We validate its performance through extensive experiments on examples of various geometric objects and benchmark real-world datasets, demonstrating significant improvements in representation quality.

cs.LG↗

A geometric framework to compare classical field theories and to transfer solutions between PDEs

In this contribution, a mathematical framework is constructed to relate and compare non-linear partial differential equations (PDEs) in the category of smooth manifolds. In particular, it can be used to compare those aspects of field theories (e.g. of classical (Newtonian) mechanics, hydrodynamics, electrodynamics, relativity theory, classical Yang-Mills theory and so on) that are described by such equations. Employing a geometric (jet space) approach, a suitable notion of shared structure of two systems of PDEs is identified. It is proven that this shared structure can serve to transfer solutions from one theory to another and a generalization of so-called Bäcklund transformations is derived that can be used to generate non-trivial solutions of some non-linear PDEs. A procedure (based on formal integrability) is introduced with which one can explicitly compute the minimal consistency conditions that two systems of PDEs need to fulfill in order to share structure under a given correspondence. Furthermore, it is shown how symmetry groups can be used to identify useful correspondences and structure that is shared up to symmetries. Thereby, the role that Bäcklund transformations play in the theory of quotient equations is clarified. Explicit examples illustrate the general ideas throughout the text and in the last chapter, the framework is applied to systems related to electrodynamics and hydrodynamics.

math-ph↗