Searcharxiv⌕ Search

arXiv subjects

A. Martina Neuman

Publications and source records attributed to A. Martina Neuman.

17 recordsLinked to original sources

Consistency of augmentation graph and network approximability in contrastive learning

Contrastive learning leverages data augmentation to develop feature representation without relying on large labeled datasets. However, despite its empirical success, the theoretical foundations of contrastive learning remain incomplete, with many essential guarantees left unaddressed, particularly the realizability assumption concerning neural approximability of an optimal spectral contrastive loss solution. In this work, we overcome these limitations by analyzing pointwise and spectral consistency of the augmentation graph Laplacian. We establish that, under specific conditions for data generation and graph connectivity, as the augmented dataset size increases, the augmentation graph Laplacian converges to a weighted Laplace-Beltrami operator on the natural data manifold. These consistency results ensure that the graph Laplacian spectrum effectively captures the manifold geometry. Consequently, they give way to a robust framework for establishing neural approximability, directly resolving the realizability assumption in a current paradigm.

cs.LG↗

Statistical Guarantees for Reasoning Probes on Looped Boolean Circuits

We study the statistical behavior of reasoning probes in a stylized model of iterative computation inspired by neural algorithmic reasoning. The underlying computation is given by a looped Boolean circuit whose graph is a perfect $ν$-ary tree ($ν\ge 2$), with outputs recursively fed back as inputs across computation rounds. A probe observes a sampled subset of internal nodes and seeks to infer the latent operation at each node, represented as a probability distribution over a finite set of admissible Boolean gates. This partial observability induces a transductive generalization problem on a structured computation graph. We show that when the probe is parameterized by a graph convolutional network and queries $N$ nodes, the worst-case generalization error decays at the optimal rate $\mathcal{O}(\sqrt{\log(2/δ)}/\sqrt{N})$ with probability at least $1-δ$. Our analysis combines metric embedding techniques with tools from optimal transport. A key insight is that this rate is achievable independently of the size of the computation graph, enabled by a low-distortion one-dimensional snowflake embedding of the induced graph metric. These results highlight a geometric mechanism underlying statistical efficiency in probing structured, iterative computations.

stat.ML↗

Tighter Learning Guarantees on Digital Computers via Concentration of Measure on Finite Spaces

Machine learning models with inputs in a Euclidean space $\mathbb{R}^d$, when implemented on digital computers, generalize, and their generalization gap converges to $0$ at a rate of $c/N^{1/2}$ concerning the sample size $N$. However, the constant $c>0$ obtained through classical methods can be large in terms of the ambient dimension $d$ and machine precision, posing a challenge when $N$ is small to realistically large. In this paper, we derive a family of generalization bounds $\{c_m/N^{1/(2\vee m)}\}_{m=1}^{\infty}$ tailored for learning models on digital computers, which adapt to both the sample size $N$ and the so-called geometric representation dimension $m$ of the discrete learning problem. Adjusting the parameter $m$ according to $N$ results in significantly tighter generalization bounds for practical sample sizes $N$, while setting $m$ small maintains the optimal dimension-free worst-case rate of $\mathcal{O}(1/N^{1/2})$. Notably, $c_{m}\in \mathcal{O}(m^{1/2})$ for learning models on discretized Euclidean domains. Furthermore, our adaptive generalization bounds are formulated based on our new non-asymptotic result for concentration of measure in finite metric spaces, established via leveraging metric embedding arguments.

cs.LG↗

Adaptivity Under Realizability Constraints: Comparing In-Context and Agentic Learning

We compare in-context learning with fixed queries and agentic learning with adaptive queries for uniform approximation of task families. We consider two settings: an unrestricted regime, where querying and approximation are arbitrary functions, and a realizable regime, where we require these operations to be implemented by ReLU neural networks. In both settings, adaptivity never hinders approximation performance. However, this advantage can change when one passes from the unrestricted regime to the realizable regime. We identify four distinct approximation scenarios, each witnessed by an explicit task family: (a) no advantage of adaptivity; (b) an advantage in the unrestricted regime that persists under ReLU realizability; (c) an advantage that arises only under realizability; and (d) an advantage that disappears under realizability. This demonstrates that representational constraints interact profoundly with the effect of adaptivity.

cs.LG↗

Reconstruction of frequency-localized functions from pointwise samples via least squares and deep learning

Recovering frequency-localized functions from pointwise data is a fundamental task in signal processing. We examine this problem from an approximation-theoretic perspective, focusing on least squares and deep learning-based methods. First, we establish a novel recovery theorem for least squares approximations using the Slepian basis from uniform random samples in low dimensions, explicitly tracking the dependence of the bandwidth on the sampling complexity. Building on these results, we then present a recovery guarantee for approximating bandlimited functions via deep learning from pointwise data. This result, framed as a practical existence theorem, provides conditions on the network architecture, training procedure, and data acquisition sufficient for accurate approximation. To complement our theoretical findings, we perform numerical comparisons between least squares and deep learning for approximating one- and two-dimensional functions. We conclude with a discussion of the theoretical limitations and the practical gaps between theory and implementation.

math.CA↗

Learning from one graph: transductive learning guarantees via the geometry of small random worlds

Since their introduction by Kipf and Welling in $2017$, a primary use of graph convolutional networks is transductive node classification, where missing labels are inferred within a single observed graph and its feature matrix. Despite the widespread use of the network model, the statistical foundations of transductive learning remain limited, as standard inference frameworks typically rely on multiple independent samples rather than a single graph. In this work, we address these gaps by developing new concentration-of-measure tools that leverage the geometric regularities of large graphs via low-dimensional metric embeddings. The emergent regularities are captured using a random graph model; however, the methods remain applicable to deterministic graphs once observed. We establish two principal learning results. The first concerns arbitrary deterministic $k$-vertex graphs, and the second addresses random graphs that share key geometric properties with an Erdős-Rényi graph $\mathbf{G}=\mathbf{G}(k,p)$ in the regime $p \in \mathcal{O}((\log (k)/k)^{1/2})$. The first result serves as the basis for and illuminates the second. We then extend these results to the graph convolutional network setting, where additional challenges arise. Lastly, our learning guarantees remain informative even with a few labelled nodes $N$ and achieve the optimal nonparametric rate $\mathcal{O}(N^{-1/2})$ as $N$ grows.

stat.ML↗

Stable Learning Using Spiking Neural Networks Equipped With Affine Encoders and Decoders

We study the learning problem associated with spiking neural networks. Specifically, we focus on spiking neural networks composed of simple spiking neurons having only positive synaptic weights, equipped with an affine encoder and decoder; we refer to these as affine spiking neural networks. These neural networks are shown to depend continuously on their parameters, which facilitates classical covering number-based generalization statements and supports stable gradient-based training. We demonstrate that the positivity of the weights enables a wide range of expressivity results, including rate-optimal approximation of smooth functions and dimension-independent approximation of Barron regular functions. In particular, we show in theory and simulations that affine spiking neural networks are capable of approximating shallow ReLU neural networks. Furthermore, we apply these affine spiking neural networks to standard machine learning benchmarks and reach competitive results. Finally, we observe that from a generalization perspective, contrary to feedforward neural networks or previous results for general spiking neural networks, the depth has little to no adverse effect on the generalization capabilities.

cs.NE↗

Theoretical guarantees for the advantage of GNNs over NNs in generalizing bandlimited functions on Euclidean cubes

Graph Neural Networks (GNNs) have emerged as formidable resources for processing graph-based information across diverse applications. While the expressive power of GNNs has traditionally been examined in the context of graph-level tasks, their potential for node-level tasks, such as node classification, where the goal is to interpolate missing node labels from the observed ones, remains relatively unexplored. In this study, we investigate the proficiency of GNNs for such classifications, which can also be cast as a function interpolation problem. Explicitly, we focus on ascertaining the optimal configuration of weights and layers required for a GNN to successfully interpolate a band-limited function over Euclidean cubes. Our findings highlight a pronounced efficiency in utilizing GNNs to generalize a bandlimited function within an $\varepsilon$-error margin. Remarkably, achieving this task necessitates only $O_d((\log\varepsilon^{-1})^d)$ weights and $O_d((\log\varepsilon^{-1})^d)$ training samples. We explore how this criterion stacks up against the explicit constructions of currently available Neural Networks (NNs) designed for similar tasks. Significantly, our result is obtained by drawing an innovative connection between the GNN structures and classical sampling theorems. In essence, our pioneering work marks a meaningful contribution to the research domain, advancing our understanding of the practical GNN applications.

cs.LG↗

Convolutional dynamical sampling and some new results

In this work, we explore the dynamical sampling problem on $\ell^2(\mathbb{Z})$ driven by a convolution operator defined by a convolution kernel. This problem is inspired by the need to recover a bandlimited heat diffusion field from space-time samples and its discrete analogue. In this book chapter, we review recent results in the finite-dimensional case and extend these findings to the infinite-dimensional case, focusing on the study of the density of space-time sampling sets.

cs.IT↗

Transferability of Graph Neural Networks using Graphon and Sampling Theories

Graph neural networks (GNNs) have become powerful tools for processing graph-based information in various domains. A desirable property of GNNs is transferability, where a trained network can swap in information from a different graph without retraining and retain its accuracy. A recent method of capturing transferability of GNNs is through the use of graphons, which are symmetric, measurable functions representing the limit of large dense graphs. In this work, we contribute to the application of graphons to GNNs by presenting an explicit two-layer graphon neural network (WNN) architecture. We prove its ability to approximate bandlimited graphon signals within a specified error tolerance using a minimal number of network weights. We then leverage this result, to establish the transferability of an explicit two-layer GNN over all sufficiently large graphs in a convergent sequence. Our work addresses transferability between both deterministic weighted graphs and simple random graphs and overcomes issues related to the curse of dimensionality that arise in other GNN results. The proposed WNN and GNN architectures offer practical solutions for handling graph data of varying sizes while maintaining performance guarantees without extensive retraining.

cs.LG↗

Functions of nearly maximal Gowers-Host-Kra norms on Euclidean spaces

Let $k\geq 2, n\geq 1$ be integers. Let $f: \mathbb{R}^{n} \to \mathbb{C}$. The $k$th Gowers-Host-Kra norm of $f$ is defined recursively by \begin{equation*} \| f\|_{U^{k}}^{2^{k}} =\int_{\mathbb{R}^{n}} \| T^{h}f \cdot \bar{f} \|_{U^{k-1}}^{2^{k-1}} \, dh \end{equation*} with $T^{h}f(x) = f(x+h)$ and $\|f\|_{U^1} = | \int_{\mathbb{R}^{n}} f(x)\, dx |$. These norms were introduced by Gowers in his work on Szemerédi's theorem, and by Host-Kra in ergodic setting. It's shown by Eisner and Tao that for every $k\geq 2$ there exist $A(k,n)< \infty$ and $p_{k} = 2^{k}/(k+1)$ such that $\| f\|_{U^{k}} \leq A(k,n)\|f\|_{p_{k}}$, for all $f \in L^{p_{k}}(\mathbb{R}^{n})$. The optimal constant $A(k,n)$ and the extremizers for this inequality are known. In this exposition, it is shown that if the ratio $\| f \|_{U^{k}}/\|f\|_{p_{k}}$ is nearly maximal, then $f$ is close in $L^{p_{k}}$ norm to an extremizer.

math.CA↗

Restricted Riemannian geometry for positive semidefinite matrices

We introduce the manifold of {\it restricted} $n\times n$ positive semidefinite matrices of fixed rank $p$, denoted $S(n,p)^{*}$. The manifold itself is an open and dense submanifold of $S(n,p)$, the manifold of $n\times n$ positive semidefinite matrices of the same rank $p$, when both are viewed as manifolds in $\mathbb{R}^{n\times n}$. This density is the key fact that makes the consideration of $S(n,p)^{*}$ statistically meaningful. We furnish $S(n,p)^{*}$ with a convenient, and geodesically complete, Riemannian geometry, as well as a Lie group structure, that permits analytical closed forms for endpoint geodesics, parallel transports, Fréchet means, exponential and logarithmic maps. This task is done partly through utilizing a {\it reduced} Cholesky decomposition, whose algorithm is also provided. We produce a second algorithm from this framework to estimate principal eigenspaces and demonstrate its superior performance over other existing algorithms.

math.DG↗

Graph Laplacians on Shared Nearest Neighbor graphs and graph Laplacians on $k$-Nearest Neighbor graphs having the same limit

A Shared Nearest Neighbor (SNN) graph is a type of graph construction using shared nearest neighbor information, which is a secondary similarity measure based on the rankings induced by a primary $k$-nearest neighbor ($k$-NN) measure. SNN measures have been touted as being less prone to the curse of dimensionality than conventional distance measures, and thus methods using SNN graphs have been widely used in applications, particularly in clustering high-dimensional data sets and in finding outliers in subspaces of high dimensional data. Despite this, the theoretical study of SNN graphs and graph Laplacians remains unexplored. In this pioneering work, we make the first contribution in this direction. We show that large scale asymptotics of an SNN graph Laplacian reach a consistent continuum limit; this limit is the same as that of a $k$-NN graph Laplacian. Moreover, we show that the pointwise convergence rate of the graph Laplacian is linear with respect to $(k/n)^{1/m}$ with high probability.

stat.ML↗

The pyramid operator

This paper gives a concept of an integral operator defined on a manifold $M$ consisting of triple of points in $\mathbb{R}^{d}$ making up a regular $3$-simplex with the origin. The boundedness of such operator is investigated. The boundedness region contains more than the Banach range - a fact that mirrors the spherical $L^{p}$-improving estimate. The purpose of this paper is two-fold: one is to investigate into an integral operator over a manifold created from high-dimensional regular simplices, two is to start a maximal operator theory for such integral operator.

math.CA↗

$L^2\times L^2\times L^2\to L^{2/3}$ boundedness for trilinear multiplier operator

This paper discusses the boundedness of the trilinear multiplier operator $T_{m}(f_1,f_2,f_3)$, when the multiplier satisfies a certain degree of smoothness but with no decaying condition and is $L^{q}$-integrable with an admissible range of $q$. The boundedness is stated in the terms of $\|m\|_{L^{q}}$. In particular, \begin{equation*}\|T_{m}\|_{L^2\times L^2\times L^2\to L^{2/3}}\lesssim\|m\|_{L^{q}}^{q/3}.\end{equation*}

math.CA↗

Anti-uniformity norms, anti-uniformity functions and their algebras on Euclidean spaces

Let $k\geq 2$ be an integer. Given a uniform function $f$ - one that satisfies $\|f\|_{U(k)}<\infty$, there is an associated anti-uniform function $g$ - one that satisfied $\|g\|_{U(k)}^{*}$. The question is, can one approximate $g$ with the Gowers-Host-Kra dual function $D_{k}f$ of $f$? Moreover, given the generalized cubic convolution products $D_{k}(f_α:α\in\tilde{V}_{k})$, what sorts of algebras can they form? In short, this paper explores possible structures of anti-uniformity on Euclidean spaces.

math.CA↗

Sparse bounds on variational norms along monomial curves

Consider a monomial curve $γ:\mathbb{R}\to\mathbb{R}^{d}$ and a family of truncated Hilbert transforms along $γ$, $\mathcal{H}^γ$. This paper addresses the possibility of the pointwise sparse domination of the $r$-variation of $\mathcal{H}^γ$ - namely, whether the following is true: \begin{equation*}V^{r}\circ\mathcal{H}^γf(x)\lesssim \mathcal{S}f(x)\end{equation*} where $f$ is a nonnegative measurable function, $r>2$ and $\mathcal{S}f(x) = \sum_{Q\in\mathcal{Q}}\langle f\rangle_{Q,p}χ_{Q}(x)$ for some $p$ and some sparse collection $\mathcal{Q}$ depending on $f,p$.

math.CA↗