SearcharxivSearch

arXiv subjects

Emanuele Zappala

Publications and source records attributed to Emanuele Zappala.

At least 19 recordsLinked to original sources

Universal approximation of continuous maps between topological vector spaces

We study the problem of universal approximation of continuous maps between topological vector spaces by neural networks. We provide a universal approximation theorem for maps between locally convex topological vector spaces with and without paracompactness assumptions. We extend this result to continuous maps between non-locally convex topological vector spaces on compact sets under the assumption of finite topological dimension.

math.GN

Quantum Cocycle Invariants of Knots from Yang-Baxter Cohomology

Yang-Baxter operators (YBOs) have been employed to construct quantum knot invariants. More recently, cohomology theories for YBOs have been independently developed, drawing inspiration from analogous theories for quandles and other discrete algebraic structures. Quandle cohomology, in particular, gives rise to cocycle invariants for knots via $2$-cocycles, which are closely related to quandle extensions, and knotted surfaces via $3$-cocycles. These quandle cocycle knot invariants have also been shown to admit interpretations as quantum invariants. Similarly, $2$-cocycles in Yang-Baxter cohomology can be interpreted in terms of deformations of YBOs. Building on these parallels, we introduce quantum cocycle invariants of knots using $2$-cocycles of Yang-Baxter cohomology, from the perspective of deformation theory. In particular, we demonstrate that the quandle cocycle invariant can be interpreted in this framework, while the quantum version yields stronger invariants in certain examples. We develop our theory along two primary approaches to quantum invariants: via trace constructions and through (co)pairings. Explicit examples are provided, including those based on the Kauffman bracket. Furthermore, we show that both the Jones and Alexander polynomials can be derived within this framework as invariants arising from higher-order formal Laurent polynomial deformations via Yang-Baxter cohomology.

math.GT

Amplituhedra for generic quantum processes via the TQNN representation of UQC

We study the relationship between computation and scattering both operationally (hence phenomenologically) and formally. We develop a representation of universal quantum computation (UQC) within the formalism of topological quantum neural networks (TQNNs), using the Reshetikhin-Turaev and Turaev-Viro models to show how TQNNs implement quantum error-correcting codes. We then exhibit a formal correspondence between TQNNs and amplituhedra to support the existence of amplituhedra for representing generic quantum processes. This construction shows how amplituhedra are geometric representations of underlying topological structures. We conclude by pointing to applications areas enabled by these results.

quant-ph

Projection Methods for Operator Learning and Universal Approximation

We obtain a new universal approximation theorem for continuous (possibly nonlinear) operators on arbitrary Banach spaces using the Leray-Schauder mapping. Moreover, we introduce and study a method for operator learning in Banach spaces $L^p$ of functions with multiple variables, based on orthogonal projections on polynomial bases. We derive a universal approximation result for operators where we learn a linear projection and a finite dimensional mapping under some additional assumptions. For the case of $p=2$, we give some sufficient conditions for the approximation results to hold. This article serves as the theoretical framework for a deep learning methodology in operator learning.

math.NA

Neural Integral Operators for Inverse Problems: An Operator-Learning Framework for Small-Sample Spectroscopic Classification

Learning maps between function spaces with a strong inductive bias is a central challenge in soft computing, especially when training data are scarce and standard deep architectures overfit. We introduce a \emph{neural integral operator} (NIO) framework based on integral equations of the first kind, in which the Urysohn kernel of the operator is parameterized by a feed-forward network~$G_{θ_G}$ and the latent function is produced by a convolutional encoder~$E_{ϕ_E}$, both trained jointly end-to-end via cross-entropy loss. The integral defining the learned operator is approximated by Monte Carlo sampling, which we argue acts as an implicit stochastic regularizer operating at the level of the integrand and complementing parameter-level regularizers such as weight decay and dropout. We benchmark the framework on three real-world spectroscopic classification tasks (FT-IR fruit purees, NIR meat, NIR textiles) of varying size and complexity, against traditional machine learning (decision tree, support vector machine, with and without UMAP) and modern deep learning baselines (FFNN, CNN+FFNN, shallow CNN, transformer). The proposed NIO is consistently among the top two performing models across all datasets and metrics, achieves the best results on the most challenging small-and-complex dataset (Textile), and yields lower performance variance than competing deep models in the small-data regime. The results suggest that operator-learning architectures with stochastic numerical integration are a viable soft-computing strategy for inverse problems in spectroscopy when conventional deep learning approaches are limited by data scarcity.

cs.LG

Nonlocal operator learning for fMRI encoding and decoding tasks

Functional MRI data exhibit high-dimensional spatiotemporal structure, making both prediction and decoding challenging. In this work, we investigate neural integral-operator-based models for encoding and decoding tasks in fMRI, with particular emphasis on the role of nonlocal spatiotemporal context. We implement a latent neural integral operator framework that performs fixed point iterations in an auxiliary space from which classification and stimuli prediction is performed via a decoder. We evaluate our model on two open-source fMRI datasets. Our experiments examine both decoding of stimuli from fMRI recordings and encoding of fMRI dynamics from stimulus representations. A main focus is the effect of spatiotemporal context: we systematically compare short and long temporal windows, as well as the use of visual cortex vs whole brain recordings, and analyze their influence on performance and latent-space geometry. Across tasks and datasets, larger temporal windows generally improve results and produce more structured learned representations. In decoding experiments, the learned latent space often provides clearer class separation than the raw data. In encoding experiments, although absolute performance remains moderate due to the difficulty of the task, longer temporal windows still yield consistent gains. These findings suggest that neural integral operators provide a promising framework for modeling fMRI dynamics and that broader spatiotemporal context can be beneficial for both prediction and representation learning. More broadly, the results indicate that exploiting distributed nonlocal structure in brain dynamics requires model architectures specifically designed to capture such dependencies.

cs.LG

Universal Approximation of Operators with Transformers and Neural Integral Operators

We study the universal approximation properties of transformers and neural integral operators for operators in Banach spaces. In particular, we show that the transformer architecture is a universal approximator of integral operators between Hölder spaces. Moreover, we show that a generalized version of neural integral operators, based on the Gavurin integral, are universal approximators of arbitrary operators between Banach spaces. Lastly, we show that a modified version of transformer, which uses Leray-Schauder mappings, is a universal approximator of operators between arbitrary Banach spaces.

cs.LG

Leray-Schauder Mappings for Operator Learning

We present an algorithm for learning operators between Banach spaces, based on the use of Leray-Schauder mappings to learn a finite-dimensional approximation of compact subspaces. We show that the resulting method is a universal approximator of (possibly nonlinear) operators. We demonstrate the efficiency of the approach on two benchmark datasets showing it achieves results comparable to state of the art models.

cs.LG

A geometric phase approach to quark confinement from stochastic gauge-geometry flows

We apply a stochastic version of the geometric (Ricci) flow, complemented with the stochastic flow of the gauge Yang--Mills sector, in order to seed the chromo-magnetic and chromo-electric vortices that source the area-law for QCD confinement. The area-law is the key signature of quark confinement in Yang--Mills gauge theories with a non-trivial center symmetry. In particular, chromo-magnetic vortices enclosed within the chromo-electric Wilson loops instantiate the area-law asymptotic behaviour of the Wilson loop vacuum expectation values. The stochastic gauge-geometry flow is responsible for the topology changes that induce the appearance of the vortices. When vortices vanish, due to topology changes in the manifolds associated to the hadronic ground states, the evaluation of the Wilson loop yields a dependence on the length of the path, hence reproducing the perimeter law of the hadronic (Higgs) phase of real QCD. Confinement, instead, is naturally achieved within this context as a by-product of the topology change of the manifold over which the dynamics of the Yang--Mills fields is defined. It is then provided by the Aharonov--Bohm effect induced by the concatenation of the compact chromo-electric and chromo-magnetic fluxes originated by the topology changes. The stochastic gauge-geometry flow naturally accomplishes a treatment of the emergence of the vortices and the generation of turbulence effects. Braiding and knotting, resulting from topology changes, namely stochastic fluctuations, stabilize the chromo-magnetic vortices. Finally, we observe that dimensional transmutation for the Yang-Mills fields can be derived from the scaling property of the geometric part of the stochastic flow. Specifically, a relation that involves the infrared equilibrium limit of the Planck constant can be derived that yields the correct order of magnitude for $Λ_{\rm QCD}$.

hep-th

Spectral methods for Neural Integral Equations

Neural integral equations are deep learning models based on the theory of integral equations, where the model consists of an integral operator and the corresponding equation (of the second kind) which is learned through an optimization procedure. This approach allows to leverage the nonlocal properties of integral operators in machine learning, but it is computationally expensive. In this article, we introduce a framework for neural integral equations based on spectral methods that allows us to learn an operator in the spectral domain, resulting in a cheaper computational cost, as well as in high interpolation accuracy. We study the properties of our methods and show various theoretical guarantees regarding the approximation capabilities of the model, and convergence to solutions of the numerical methods. We provide numerical experiments to demonstrate the practical effectiveness of the resulting model.

math.NA

Non-Markovian Discrete Diffusion with Causal Language Models

Discrete diffusion models offer a flexible, controllable approach to structured sequence generation, yet they still lag behind causal language models in expressive power. A key limitation lies in their reliance on the Markovian assumption, which restricts each step to condition only on the current state, leading to potential uncorrectable error accumulation. In this paper, we introduce CaDDi (Causal Discrete Diffusion Model), a discrete diffusion model that conditions on the entire generative trajectory, thereby lifting the Markov constraint and allowing the model to revisit and improve past states. By unifying sequential (causal) and temporal (diffusion) reasoning in a single non-Markovian transformer, CaDDi also treats standard causal language models as a special case and permits the direct reuse of pretrained LLM weights with no architectural changes. Empirically, CaDDi outperforms state-of-the-art discrete diffusion baselines on natural-language benchmarks, substantially narrowing the remaining gap to large autoregressive transformers.

cs.LG

Operational protocols cannot certify classicality

The existence and practical utility of operational protocols that certify entanglement raises the question of whether operational protocols exist that certify the absence of entanglement, i.e. that certify separability. We show, within a purely topological, interpretation-independent representation, that such protocols do not exist. Classicality is therefore, as Bohr suggested, purely a pragmatic notion.

quant-ph

Yang-Baxter Hochschild Cohomology

Braided algebras are associative algebras endowed with a Yang-Baxter operator that satisfies certain compatibility conditions involving the multiplication. Along with Hochschild cohomology of algebras, there is also a notion of Yang-Baxter cohomology, which is associated to any Yang-Baxter operator. In this article, we introduce and study a cohomology theory for braided algebras in dimensions 2 and 3, that unifies Hochschild and Yang-Baxter cohomology theories, and generalizes to all dimensions in characteristic $2$. We show that its second cohomology group classifies infinitesimal deformations of braided algebras. We provide infinite families of examples of braided algebras, including Hopf algebras, tensorized multiple conjugation quandles, and braided Frobenius algebras. Moreover, we derive the obstructions to higher deformations, which lie in the third cohomology group. Relations to Hopf algebra cohomology are also discussed.

math.QA

Intelligence at the Edge of Chaos

We explore the emergence of intelligent behavior in artificial systems by investigating how the complexity of rule-based systems influences the capabilities of models trained to predict these rules. Our study focuses on elementary cellular automata (ECA), simple yet powerful one-dimensional systems that generate behaviors ranging from trivial to highly complex. By training distinct Large Language Models (LLMs) on different ECAs, we evaluated the relationship between the complexity of the rules' behavior and the intelligence exhibited by the LLMs, as reflected in their performance on downstream tasks. Our findings reveal that rules with higher complexity lead to models exhibiting greater intelligence, as demonstrated by their performance on reasoning and chess move prediction tasks. Both uniform and periodic systems, and often also highly chaotic systems, resulted in poorer downstream performance, highlighting a sweet spot of complexity conducive to intelligence. We conjecture that intelligence arises from the ability to predict complexity and that creating intelligence may require only exposure to complexity.

cs.AI

Deformation Cohomology for Braided Commutativity

Braided algebras are algebraic structures consisting of an algebra endowed with a Yang-Baxter operator, satisfying some compatibility conditions.Yang-Baxter Hochschild cohomology was introduced by the authors to classify infinitesimal deformations of braided algebras, and determine obstructions to higher order deformations. Several examples of braided algebras satisfy a weaker version of commutativity, which is called braided commutativity and involves the Yang-Baxter operator of the algebra. We extend the theory of Yang-Baxter Hochschild cohomology to study braided commutative deformations of braided algebras. The resulting cohomology theory classifies infinitesimal deformations of braided algebras that are braided commutative, and provides obstructions for braided commutative higher order deformations. We consider braided commutativity for Hopf algebras in detail, and obtain some classes of nontrivial examples.

math.QA

Whether a quantum computation employs nonlocal resources is operationally undecidable

Computational complexity characterizes the usage of spatial and temporal resources by computational processes. In the classical theory of computation, e.g. in the Turing Machine model, computational processes employ only local space and time resources, and their resource usage can be accurately measured by us as users. General relativity and quantum theory, however, introduce the possibility of computational processes that employ nonlocal spatial or temporal resources. While the space and time complexity of classical computing can be given a clear operational meaning, this is no longer the case in any setting involving nonlocal resources. In such settings, theoretical analyses of resource usage cease to be reliable indicators of practical computational capability. We prove that the verifier (C) in a multiple interactive provers with shared entanglement (MIP*) protocol cannot operationally demonstrate that the "multiple" provers are independent, i.e. cannot operationally distinguish a MIP* machine from a monolithic quantum computer. Thus C cannot operationally distinguish a MIP* machine from a quantum TM, and hence cannot operationally demonstrate the solution to arbitrary problems in RE. Any claim that a MIP* machine has solved a TM-undecidable problem is, therefore, circular, as the problem of deciding whether a physical system is a MIP* machine is itself TM-undecidable. Consequently, despite the space and time complexity of classical computing having a clear operational meaning, this is no longer the case in any setting involving nonlocal resources. In such settings, theoretical analyses of resource usage cease to be reliable indicators of practical computational capability. This has practical consequences when assessing newly proposed computational frameworks based on quantum theories.

quant-ph

ER = EPR is an operational theorem

We show that in the operational setting of a two-agent, local operations, classical communication (LOCC) protocol, Alice and Bob cannot operationally distinguish monogamous entanglement from a topological identification of points in their respective local spacetimes, i.e. that ER = EPR can be recovered as an operational theorem. Our construction immediately implies that in this operational setting, the local topology of spacetime is observer-relative. It also provides a simple demonstration of the non-traversability of ER bridges. As our construction does not depend on an embedding geometry, it generalizes previous geometric approaches to ER = EPR.

quant-ph

Deep Neural Networks as the Semi-classical Limit of Topological Quantum Neural Networks: The problem of generalisation

Deep Neural Networks miss a principled model of their operation. A novel framework for supervised learning based on Topological Quantum Field Theory that looks particularly well suited for implementation on quantum processors has been recently explored. We propose using this framework to understand the problem of generalisation in Deep Neural Networks. More specifically, in this approach, Deep Neural Networks are viewed as the semi-classical limit of Topological Quantum Neural Networks. A framework of this kind explains the overfitting behavior of Deep Neural Networks during the training step and the corresponding generalisation capabilities. We explore the paradigmatic case of the perceptron, which we implement as the semiclassical limit of Topological Quantum Neural Networks. We apply a novel algorithm we developed, showing that it obtains similar results to standard neural networks, but without the need for training (optimisation).

quant-ph