SearcharxivSearch

arXiv subjects

Nihat Ay

Publications and source records attributed to Nihat Ay.

At least 19 recordsLinked to original sources

Integrated Information in the Active Inference Framework

The active inference framework provides a principled approach to modeling sentient behavior. In this framework perception and action selection are treated in a unified way. The resulting agents form an internal generative model of the relevant dynamics of the world in order to infer their future observations, their internal states and to select actions. We combine this modeling framework with the Integrated Information Theory of consciousness and are therefore able to analyze the active inference agents from the perspective of integrated information. The Integrated Information Theory aims at quantifying the level of consciousness of a system by assessing its capability to integrate information. Here, we define a measure of integrated information for the generative model by making an additional structural assumption. Experiments with simulated agents reveal a correlation between integrated information measures and the free energy of the active inference agents that increases with the size of the generative model.

cs.IT

Dynamics-Aligned Shared Hypernetworks for Contextual RL under Discontinuous Shifts

Zero-shot generalization in contextual reinforcement learning remains a core challenge, particularly when the context is latent and must be inferred from data. A canonical failure mode arises when latent context discontinuously changes how actions affect the environment, requiring incompatible control responses across contexts. We propose DMA*-SH, a framework where a single hypernetwork, trained solely via dynamics prediction, generates a small set of adapter weights shared across the dynamics model, policy, and action-value function. This shared modulation imparts an inductive bias matched to discontinuous context-to-dynamics shifts, while input/output normalization and random input masking stabilize context inference, promoting directionally concentrated representations. We provide theoretical support via expressivity separation results for hypernetwork modulation, and a variance decomposition with policy-gradient variance bounds that formalize how within-mode compression improves learning under non-overlapping contexts. For evaluation, we introduce the Actuator Inversion Benchmark (AIB), a suite of environments designed to isolate challenging context-to-dynamics interactions, including actuator inversion, actuator permutations, and weakly non-overlapping continuous dynamics. On AIB's held-out tasks, DMA*-SH achieves zero-shot generalization, outperforming domain randomization by 58.1% and surpassing a standard context-aware baseline by 11.5% on average.

cs.LG

Algorithmic bottlenecks in evolution: Genetic code, symbolic language, and the Great Filter hypothesis

The Great Filter hypothesis proposes that the emergence of technological societies capable of interstellar travel depends on a small number of exceptionally hard and highly improbable steps. Traditional versions of this hypothesis enumerate such "hard steps" along the trajectory from inanimate matter to complex technological societies but diverge in their explanations for why these particular steps should be so improbable. The theory of Major Evolutionary Transitions also faces challenges in identifying which steps should be considered universally "hard" across different evolutionary pathways. In contrast, we argue that two deeply structural obstacles dominate the evolutionary landscape: the coding threshold associated with the origin of genetic code, and the language threshold associated with the emergence of symbolic communication. We examine the developmental precursors of both transitions and analyze the underlying algorithmic bottlenecks: points at which evolving systems separate code from function, while entangling them within information hierarchies. Using a game-theoretic analysis of coupled signaling and coordination dynamics, we then argue that the corresponding multichannel games may exhibit saddle-type equilibria whose stable manifolds define narrow evolutionary paths, making the transitions intrinsically difficult to traverse. We conjecture that the so-called Great Filter is best understood not as a sequence of isolated improbable events, but as a nested structure of tangled information hierarchies. Under this conjecture, the rarity of advanced societies follows from the difficulty of crossing these coding thresholds in a competitive noisy environment. This hypothesis reframes the Great Filter as an algorithmic property of evolving systems, suggesting that only a small fraction of life may ever traverse the path toward technological societies capable of interstellar travel.

q-bio.PE

Beyond Silicon: Materials, Mechanisms, and Methods for Physical Neural Computing

Physical implementations of neural computation now extend far beyond silicon hardware, encompassing substrates such as memristive devices, photonic circuits, mechanical metamaterials, microfluidic networks, chemical reaction systems, and living neural tissue. By exploiting intrinsic physical processes such as charge transport, wave interference, elastic deformation, mass transport, and biochemical regulation, these substrates can realize neural inference and adaptation directly in matter. As silicon GPU-centered AI faces growing energy and data-movement constraints, physical neural computation is becoming increasingly relevant as a complementary path beyond conventional digital accelerators. This trend is driven in particular by pervasive intelligence, i.e., the deployment of on-device and edge AI across large numbers of resource-constrained systems. In such settings, co-locating computation with sensing and memory can reduce data shuttling and improve efficiency. Meanwhile, physical neural approaches have emerged across disparate disciplines, yet progress remains fragmented, with limited shared terminology and few principled ways to compare platforms. This survey unifies the field by mapping neural primitives to substrate-specific mechanisms, analyzing architectural and training paradigms, and identifying key engineering constraints including scalability, precision, programmability, and I/O interfacing overhead. To enable cross-domain comparison, we introduce a first-order benchmarking scheme based on standardized static and dynamic tasks and physically interpretable performance dimensions. We show that no single substrate dominates across the considered dimensions; instead, physical neural systems occupy complementary operating regimes, enabling applications ranging from ultrafast signal processing and in-memory inference to embodied control and in-sample biochemical decision making.

cs.NE

Learning a Latent Pulse Shape Interface for Photoinjector Laser Systems

Controlling the longitudinal laser pulse shape in photoinjectors of Free-Electron Lasers is a powerful lever for optimizing electron beam quality, but systematic exploration of the vast design space is limited by the cost of brute-force pulse propagation simulations. We present a generative modeling framework based on Wasserstein Autoencoders to learn a differentiable latent interface between pulse shaping and downstream beam dynamics. Our empirical findings show that the learned latent space is continuous and interpretable while maintaining high-fidelity reconstructions. Pulse families such as higher-order Gaussians trace coherent trajectories, while standardizing the temporal pulse lengths shows a latent organization correlated with pulse energy. Analysis via principal components and Gaussian Mixture Models reveals a well behaved latent geometry, enabling smooth transitions between distinct pulse types via linear interpolation. The model generalizes from simulated data to real experimental pulse measurements, accurately reconstructing pulses and embedding them consistently into the learned manifold. Overall, the approach reduces reliance on expensive pulse-propagation simulations and facilitates downstream beam dynamics simulation and analysis.

cs.LG

Well-Posed KL-Regularized Control via Wasserstein and Kalman-Wasserstein KL Divergences

Kullback-Leibler (KL) divergence regularization is widely used in reinforcement learning, but it becomes infinite under support mismatch and can degenerate in low-noise regimes. Using a unified information-geometric framework, we introduce KL analogs by replacing the Fisher-Rao geometry in the dynamical formulation of the KL with transport-based geometries, and derive closed-form expressions for common distribution families. Between elliptic distributions, these divergences remain finite for degenerating equal covariances and yield a geometric interpretation of regularization heuristics used in Kalman ensemble methods. We demonstrate the utility of these divergences in KL-regularized optimal control. In the fully tractable setting of linear time-invariant systems with Gaussian process noise, the classical KL reduces to a quadratic control penalty that becomes singular as process noise vanishes. Our variants remove this singularity and yield well-posed problems. In both the double integrator and cart-pole examples, the resulting controls preserve nontrivial feedback and achieve better closed-loop performance.

math.OC

Generalising thermodynamic efficiency of interactions: inferential, information-geometric and computational perspectives

Self-organizing systems consume energy to generate internal order. The concept of thermodynamic efficiency, drawing from statistical physics and information theory, has previously been proposed to characterise a change in control parameter by relating the resulting predictability gain to the required amount of work. However, previous studies have taken a system-centric perspective and considered only single control parameters. Here, we generalise thermodynamic efficiency to multiple control parameters and extend the definition of thermodynamic efficiency to protocols in arbitrary directions, by introducing directional efficiency. Taking an observer-centric perspective, we derive two novel formulations. The first, an inferential form, relates efficiency to fluctuations of macroscopic observables, interpreting thermodynamic efficiency in terms of how well the system parameters can be inferred from observable macroscopic behaviour. The second, an information-geometric form, expresses efficiency in terms of the Fisher information matrix, interpreting it with respect to how difficult it is to navigate the statistical manifold defined by the control protocol. This observer-centric perspective is contrasted with the existing system-centric view, where efficiency is considered an intrinsic property of the system.

nlin.AO

On the Natural Gradient of the Evidence Lower Bound

This article studies the Fisher-Rao gradient, also referred to as the natural gradient, of the evidence lower bound (ELBO) which plays a central role in generative machine learning. It reveals that the gap between the evidence and its lower bound, the ELBO, has essentially a vanishing natural gradient within unconstrained optimization. As a result, maximization of the ELBO is equivalent to minimization of the Kullback-Leibler divergence from a target distribution, the primary objective function of learning. Building on this insight, we derive a condition under which this equivalence persists even when optimization is constrained to a model. This condition yields a geometric characterization, which we formalize through the notion of a cylindrical model.

cs.LG

Convergence Properties of Natural Gradient Descent for Minimizing KL Divergence

The Kullback-Leibler (KL) divergence plays a central role in probabilistic machine learning, where it commonly serves as the canonical loss function. Optimization in such settings is often performed over the probability simplex, where the choice of parameterization significantly impacts convergence. In this work, we study the problem of minimizing the KL divergence and analyze the behavior of gradient-based optimization algorithms under two dual coordinate systems within the framework of information geometry$-$ the exponential family ($θ$ coordinates) and the mixture family ($η$ coordinates). We compare Euclidean gradient descent (GD) in these coordinates with the coordinate-invariant natural gradient descent (NGD), where the natural gradient is a Riemannian gradient that incorporates the intrinsic geometry of the underlying statistical model. In continuous time, we prove that the convergence rates of GD in the $θ$ and $η$ coordinates provide lower and upper bounds, respectively, on the convergence rate of NGD. Moreover, under affine reparameterizations of the dual coordinates, the convergence rates of GD in $η$ and $θ$ coordinates can be scaled to $2c$ and $\frac{2}{c}$, respectively, for any $c>0$, while NGD maintains a fixed convergence rate of $2$, remaining invariant to such transformations and sandwiched between them. Although this suggests that NGD may not exhibit uniformly superior convergence in continuous time, we demonstrate that its advantages become pronounced in discrete time, where it achieves faster convergence and greater robustness to noise, outperforming GD. Our analysis hinges on bounding the spectrum and condition number of the Hessian of the KL divergence at the optimum, which coincides with the Fisher information matrix.

cs.LG

A Concise Mathematical Description of Active Inference in Discrete Time

In this paper we present a concise mathematical description of active inference in discrete time. The main part of the paper serves as a basic introduction to the topic, including a detailed example of the action selection mechanism. The appendix discusses the more subtle mathematical details, targeting readers who have already studied the active inference literature but struggle to make sense of the mathematical details and derivations. Throughout, we emphasize precise and standard mathematical notation, ensuring consistency with existing texts and linking all equations to widely used references on active inference. Additionally, we provide Python code that implements the action selection and learning mechanisms described in this paper and is compatible with pymdp environments.

cs.LG

Information Geometry for Wasserstein KL Divergence of Gaussian Measures on $\mathbb{R}^n$

We study the Wasserstein Kullback--Leibler divergence (WKL divergence) on the manifold of nondegenerate Gaussian measures over $\mathbb R^n$. In the canonical-divergence construction, the Fisher--Rao metric recovers forward KL along intrinsic mixture geodesics and reverse KL along geodesics of the conjugate exponential connection. Replacing the Fisher--Rao metric by the Otto metric and following the latter route produces the $e_1$-connection underlying WKL divergence. We establish its geodesic completeness, classify its forward limits, prove that every ordered pair is joined by a unique $e_1$-connector generated by a quadratic potential, and derive an explicit WKL divergence formula with separate mean and covariance contributions. WKL divergence is nonnegative and separating. At equal covariances WKL divergence equals one half of the squared Euclidean mean distance. Finally, WKL divergence extends finitely and continuously to singular targets from a nondegenerate source, but diverges when the target remains nondegenerate and the source covariance becomes singular. Path-dependent joint limits at Dirac pairs preclude a continuous extension to the full positive-semidefinite covariance product, although a lower-semicontinuous extended-real extension exists.

math.ST

Torsion of $α$-connections on the density manifold

We study the torsion of the $α$-connections defined on the density manifold in terms of a regular Riemannian metric. In the case of the Fisher-Rao metric our results confirm the fact that all $α$-connections are torsion free. For the $α$-connections obtained by the Otto metric, we show that, except for $α= -1$, they are not torsion free.

cs.IT

Outsourcing Control requires Control Complexity

An embodied agent constantly influences its environment and is influenced by it. We use the sensorimotor loop to model these interactions and thereby we can quantify different information flows in the system by various information theoretic measures. This includes a measure for the interaction among the agent's body and its environment, called Morphological Computation. Additionally, we examine the controller complexity by two measures, one of which can be seen in the context of the Integrated Information Theory of consciousness. Applying this framework to an experimental setting with simulated agents allows us to analyze the interaction between an agent and its environment, as well as the complexity of its controller, the brain of the agent. Previous research reveals an antagonistic relationship between the controller complexity and Morphological Computation. A morphology adapted well to a task can reduce the necessary complexity of the controller significantly. This creates the problem that embodied intelligence is correlated with a reduced necessity of a controller, a brain. However, in order to interact well with their surroundings, the agents first have to understand the relevant dynamics of the environment. By analyzing learning agents we observe that an increased controller complexity can facilitate a better interaction between an agent's body and its environment. Hence, learning requires an increased controller complexity and the controller complexity and Morphological Computation influence each other.

cs.IT

Analyzing Multimodal Integration in the Variational Autoencoder from an Information-Theoretic Perspective

Human perception is inherently multimodal. We integrate, for instance, visual, proprioceptive and tactile information into one experience. Hence, multimodal learning is of importance for building robotic systems that aim at robustly interacting with the real world. One potential model that has been proposed for multimodal integration is the multimodal variational autoencoder. A variational autoencoder (VAE) consists of two networks, an encoder that maps the data to a stochastic latent space and a decoder that reconstruct this data from an element of this latent space. The multimodal VAE integrates inputs from different modalities at two points in time in the latent space and can thereby be used as a controller for a robotic agent. Here we use this architecture and introduce information-theoretic measures in order to analyze how important the integration of the different modalities are for the reconstruction of the input data. Therefore we calculate two different types of measures, the first type is called single modality error and assesses how important the information from a single modality is for the reconstruction of this modality or all modalities. Secondly, the measures named loss of precision calculate the impact that missing information from only one modality has on the reconstruction of this modality or the whole vector. The VAE is trained via the evidence lower bound, which can be written as a sum of two different terms, namely the reconstruction and the latent loss. The impact of the latent loss can be weighted via an additional variable, which has been introduced to combat posterior collapse. Here we train networks with four different weighting schedules and analyze them with respect to their capabilities for multimodal integration.

cs.LG

Inversion of Bayesian Networks

Variational autoencoders and Helmholtz machines use a recognition network (encoder) to approximate the posterior distribution of a generative model (decoder). In this paper we study the necessary and sufficient properties of a recognition network so that it can model the true posterior distribution exactly. These results are derived in the general context of probabilistic graphical modelling / Bayesian networks, for which the network represents a set of conditional independence statements. We derive both global conditions, in terms of d-separation, and local conditions for the recognition network to have the desired qualities. It turns out that for the local conditions the property perfectness (for every node, all parents are joined) plays an important role.

cs.LG

Natural Reweighted Wake-Sleep

Helmholtz Machines (HMs) are a class of generative models composed of two Sigmoid Belief Networks (SBNs), acting respectively as an encoder and a decoder. These models are commonly trained using a two-step optimization algorithm called Wake-Sleep (WS) and more recently by improved versions, such as Reweighted Wake-Sleep (RWS) and Bidirectional Helmholtz Machines (BiHM). The locality of the connections in an SBN induces sparsity in the Fisher Information Matrices associated to the probabilistic models, in the form of a finely-grained block-diagonal structure. In this paper we exploit this property to efficiently train SBNs and HMs using the natural gradient. We present a novel algorithm, called Natural Reweighted Wake-Sleep (NRWS), that corresponds to the geometric adaptation of its standard version. In a similar manner, we also introduce Natural Bidirectional Helmholtz Machine (NBiHM). Differently from previous work, we will show how for HMs the natural gradient can be efficiently computed without the need of introducing any approximation in the structure of the Fisher information matrix. The experiments performed on standard datasets from the literature show a consistent improvement of NRWS and NBiHM not only with respect to their non-geometric baselines but also with respect to state-of-the-art training algorithms for HMs. The improvement is quantified both in terms of speed of convergence as well as value of the log-likelihood reached after training.

cs.LG

Invariance Properties of the Natural Gradient in Overparametrised Systems

The natural gradient field is a vector field that lives on a model equipped with a distinguished Riemannian metric, e.g. the Fisher-Rao metric, and represents the direction of steepest ascent of an objective function on the model with respect to this metric. In practice, one tries to obtain the corresponding direction on the parameter space by multiplying the ordinary gradient by the inverse of the Gram matrix associated with the metric. We refer to this vector on the parameter space as the natural parameter gradient. In this paper we study when the pushforward of the natural parameter gradient is equal to the natural gradient. Furthermore we investigate the invariance properties of the natural parameter gradient. Both questions are addressed in an overparametrised setting.

cs.LG

Approaching a large deviation theory for complex systems

The standard Large Deviation Theory (LDT) is mathematically illustrated by the Boltzmann-Gibbs factor which describes the thermal equilibrium of short-range-interacting many-body Hamiltonian systems, the velocity distribution of which is Maxwellian. It is generically applicable to systems satisfying the Central Limit Theorem (CLT). When we focus instead on stationary states of typical complex systems (e.g., classical long-range-interacting many-body Hamiltonian systems, such as self-gravitating ones), the CLT, and possibly also the LDT, need to be generalised. Specifically, when the $N\to\infty$ attractor ($N$ being the number of degrees of freedom) in the space of distributions is a $Q$-Gaussian (a nonadditive $q$-entropy-based generalisation of the standard Gaussian case, which is recovered for $Q=1$) related to a $Q$-generalised CLT, we expect the LDT probability distribution to asymptotically approach a power law. Consistently with available strong numerical indications for probabilistic models, this behaviour possibly is that associated to a $q$-exponential (defined as $e_q^x\equiv\left[1+(1-q)x\right]^{1/(1-q)}$, which is the generalisation of the standard exponential form, straightforwardly recovered for $q=1$); $q$ and $Q$ are expected to be simply connected, including the particular case $q=Q=1$. The argument of such $q$-exponential would be expected to be proportional to $N$, analogously to the thermodynamical entropy of many-body Hamiltonian systems. We provide here numerical evidence supporting the asymptotic power-law by analysing the standard map, the coherent noise model for biological extinctions and earthquakes, the Ehrenfest dog-flea model, and the random-walk avalanches.

cond-mat.stat-mech