SearcharxivSearch

arXiv subjects

Takuya Isomura

Publications and source records attributed to Takuya Isomura.

13 recordsLinked to original sources

Information mechanics: conservation and assimilation

Inference and learning are commonly cast in terms of optimisation, yet the invariant constraints governing uncertainty reduction remain unclear. This work presents information mechanics (infomechanics), a first-principles framework that describes informational structure in two canonical state coordinates. Starting from the pointwise identity implied by Bayes' rule, minimal requirements of additivity, symmetry, and finite-resolution robustness select only two distinct additive projections, yielding conservation identities for Shannon entropy, governing global uncertainty, and Fisher information, encoding complementary local geometry. Because part of Fisher information is fixed by entropy, the residual structure is captured by a non-additive, coordinate-scale-invariant state function, the information potential $\Phi$. This yields a two-coordinate description that separates the entropic baseline from residual geometric complexity. $\Phi$ vanishes uniquely for isotropic Gaussian distributions and decreases under Gaussian coarse-graining. In finite-resolution multimodal landscapes, $\Phi$ asymptotically scales with the logarithm of the effective number of local optima, linking information geometry to inference difficulty. The same two-coordinate formalism extends to the Markov chain linking hidden states, observations, and internal representations, yielding assimilation inequalities that constrain faithful external-state inference. Together, these results identify invariant constraints underlying inference, learning, and computation across biological and artificial systems.

cs.IT

Triple equivalence for the emergence of biological intelligence

Characterising the intelligence of biological organisms is challenging. This work considers intelligent algorithms developed evolutionarily within neural systems. Mathematical analyses unveil a natural equivalence between canonical neural networks, variational Bayesian inference under a class of partially observable Markov decision processes, and differentiable Turing machines, by showing that they minimise the shared Helmholtz energy. Consequently, canonical neural networks can biologically plausibly equip Turing machines and conduct variational Bayesian inferences of external Turing machines in the environment. Applying Helmholtz energy minimisation at the species level facilitates deriving active Bayesian model selection inherent in natural selection, resulting in the emergence of adaptive algorithms. In particular, canonical neural networks with two mental actions can separately memorise transition mappings of multiple external Turing machines to form a universal machine. These propositions were corroborated by numerical simulations of algorithm implementation and neural network evolution. These notions offer a universal characterisation of biological intelligence emerging from evolution in terms of Bayesian model selection and belief updating.

q-bio.NC

Bayesian mechanics of self-organising systems

Bayesian mechanics provides a framework that addresses dynamical systems that can be conceptualised as Bayesian inference. However, elucidating the requisite generative models is essential for empirical applications to realistic self-organising systems. This work shows that the Hamiltonian of generic dynamical systems constitutes a class of generative models, thus rendering their Helmholtz energy equivalent to variational free energy under the identified generative model. The self-organisation that minimises the Helmholtz energy entails matching the system's Hamiltonian with that of the environment, leading to the ensuing emergence of their generalised synchrony. In essence, these self-organising systems can be read as performing variational Bayesian inference of their interacting environment. These properties have been demonstrated using coupled oscillators, simulated and living neural networks, and quantum computers. This framework offers foundational characterisations and predictions regarding asymptotic properties of self-organising systems interacting with their environment, providing insights into potential mechanisms underlying the emergence of intelligence.

q-bio.NC

On Predictive planning and counterfactual learning in active inference

Given the rapid advancement of artificial intelligence, understanding the foundations of intelligent behaviour is increasingly important. Active inference, regarded as a general theory of behaviour, offers a principled approach to probing the basis of sophistication in planning and decision-making. In this paper, we examine two decision-making schemes in active inference based on 'planning' and 'learning from experience'. Furthermore, we also introduce a mixed model that navigates the data-complexity trade-off between these strategies, leveraging the strengths of both to facilitate balanced decision-making. We evaluate our proposed model in a challenging grid-world scenario that requires adaptability from the agent. Additionally, our model provides the opportunity to analyze the evolution of various parameters, offering valuable insights and contributing to an explainable framework for intelligent decision-making.

cs.AI

Active Inference and Intentional Behaviour

Recent advances in theoretical biology suggest that basal cognition and sentient behaviour are emergent properties of in vitro cell cultures and neuronal networks, respectively. Such neuronal networks spontaneously learn structured behaviours in the absence of reward or reinforcement. In this paper, we characterise this kind of self-organisation through the lens of the free energy principle, i.e., as self-evidencing. We do this by first discussing the definitions of reactive and sentient behaviour in the setting of active inference, which describes the behaviour of agents that model the consequences of their actions. We then introduce a formal account of intentional behaviour, that describes agents as driven by a preferred endpoint or goal in latent state-spaces. We then investigate these forms of (reactive, sentient, and intentional) behaviour using simulations. First, we simulate the aforementioned in vitro experiments, in which neuronal cultures spontaneously learn to play Pong, by implementing nested, free energy minimising processes. The simulations are then used to deconstruct the ensuing predictive behaviour, leading to the distinction between merely reactive, sentient, and intentional behaviour, with the latter formalised in terms of inductive planning. This distinction is further studied using simple machine learning benchmarks (navigation in a grid world and the Tower of Hanoi problem), that show how quickly and efficiently adaptive behaviour emerges under an inductive form of active inference.

q-bio.NC

Dimensionality reduction to maximize prediction generalization capability

Generalization of time series prediction remains an important open issue in machine learning, wherein earlier methods have either large generalization error or local minima. We develop an analytically solvable, unsupervised learning scheme that extracts the most informative components for predicting future inputs, termed predictive principal component analysis (PredPCA). Our scheme can effectively remove unpredictable noise and minimize test prediction error through convex optimization. Mathematical analyses demonstrate that, provided with sufficient training samples and sufficiently high-dimensional observations, PredPCA can asymptotically identify hidden states, system parameters, and dimensionalities of canonical nonlinear generative processes, with a global convergence guarantee. We demonstrate the performance of PredPCA using sequential visual inputs comprising hand-digits, rotating 3D objects, and natural scenes. It reliably estimates distinct hidden states and predicts future outcomes of previously unseen test input data, based exclusively on noisy observations. The simple architecture and low computational cost of PredPCA are highly desirable for neuromorphic hardware.

stat.ML

Kalman filters as the steady-state solution of gradient descent on variational free energy

The Kalman filter is an algorithm for the estimation of hidden variables in dynamical systems under linear Gauss-Markov assumptions with widespread applications across different fields. Recently, its Bayesian interpretation has received a growing amount of attention especially in neuroscience, robotics and machine learning. In neuroscience, in particular, models of perception and control under the banners of predictive coding, optimal feedback control, active inference and more generally the so-called Bayesian brain hypothesis, have all heavily relied on ideas behind the Kalman filter. Active inference, an algorithmic theory based on the free energy principle, specifically builds on approximate Bayesian inference methods proposing a variational account of neural computation and behaviour in terms of gradients of variational free energy. Using this ambitious framework, several works have discussed different possible relations between free energy minimisation and standard Kalman filters. With a few exceptions, however, such relations point at a mere qualitative resemblance or are built on a set of very diverse comparisons based on purported differences between free energy minimisation and Kalman filtering. In this work, we present a straightforward derivation of Kalman filters consistent with active inference via a variational treatment of free energy minimisation in terms of gradient descent. The approach considered here offers a more direct link between models of neural dynamics as gradient descent and standard accounts of perception and decision making based on probabilistic inference, further bridging the gap between hypotheses about neural implementation and computational principles in brain and behavioural sciences.

q-bio.NC

Quadratic speedup of global search using a biased crossover of two good solutions

The minimisation of cost functions is crucial in various optimisation fields. However, identifying their global minimum remains challenging owing to the huge computational cost incurred. This work analytically expresses the computational cost to identify an approximate global minimum for a class of cost functions defined under a high-dimensional discrete state space. Then, we derive an optimal global search scheme that minimises the computational cost. Mathematical analyses demonstrate that a combination of the gradient descent algorithm and the selection and crossover algorithm--with a biased crossover weight--maximises the search efficiency. Remarkably, its computational cost is of the square root order in contrast to that of the conventional gradient descent algorithms, indicating a quadratic speedup of global search. We corroborate this proposition using numerical analyses of the travelling salesman problem. The simple computational architecture and minimal computational cost of the proposed scheme are highly desirable for biological organisms and neuromorphic hardware.

cs.NE

On the achievability of blind source separation for high-dimensional nonlinear source mixtures

For many years, a combination of principal component analysis (PCA) and independent component analysis (ICA) has been used for blind source separation (BSS). However, it remains unclear why these linear methods work well with real-world data that involve nonlinear source mixtures. This work theoretically validates that a cascade of linear PCA and ICA can solve a nonlinear BSS problem accurately -- when the sensory inputs are generated from hidden sources via nonlinear mappings with sufficient dimensionality. Our proposed theorem, termed the asymptotic linearization theorem, theoretically guarantees that applying linear PCA to the inputs can reliably extract a subspace spanned by the linear projections from every hidden source as the major components -- and thus projecting the inputs onto their major eigenspace can effectively recover a linear transformation of the hidden sources. Then, subsequent application of linear ICA can separate all the true independent hidden sources accurately. Zero-element-wise-error nonlinear BSS is asymptotically attained when the source dimensionality is large and the input dimensionality is sufficiently larger than the source dimensionality. Our proposed theorem is validated analytically and numerically. Moreover, the same computation can be performed by using Hebbian-like plasticity rules, implying the biological plausibility of this nonlinear BSS strategy. Our results highlight the utility of linear PCA and ICA for accurately and reliably recovering nonlinearly mixed sources -- and further suggest the importance of employing sensors with sufficient dimensionality to identify true hidden sources of real-world data.

stat.ML

Inferring neuronal couplings from spiking data using a systematic procedure with a statistical criterion

Recent remarkable advances in the experimental techniques have provided a background for inferring neuronal couplings from point process data that includes a great number of neurons. Here, we propose a systematic procedure for pre- and post-processing generic point process data in an objective manner, to handle data in the framework of a binary simple statistical model, the Ising or generalized McCulloch--Pitts model. The procedure involves two steps: (1) determining time-bin size for transforming the point-process data into discrete-time binary data and (2) screening relevant couplings from the estimated couplings. For the first step, we decide the optimal time-bin size by introducing the null hypothesis that all neurons would fire independently, then choosing a time-bin size so that the null hypothesis is rejected with the most strict criterion. The likelihood associated with the null hypothesis is analytically evaluated and used for the rejection process. For the second post-processing step, after a certain estimator of coupling is obtained based on the pre-processed dataset, the estimate is compared with many other estimates derived from datasets obtained by randomizing the original dataset in the time direction. We accept the original estimate as relevant only if its absolute value is sufficiently larger than them of randomized datasets. These manipulations suppress false positive couplings induced by statistical noise. We apply this inference procedure to spiking data from synthetic and in vitro neuronal networks. The results show that the proposed procedure identifies the presence/absence of synaptic couplings fairly well including their signs, for the synthetic and experimental data. In particular, the results support that we can infer the physical connections of underlying systems in favorable situations, even when using the simple statistical model.

cond-mat.dis-nn

Computational cost for determining an approximate global minimum using the selection and crossover algorithm

This work examines the expected computational cost to determine an approximate global minimum of a class of cost functions characterized by the variance of coefficients. The cost function takes $N$-dimensional binary states as arguments and has many local minima. Iterations in the order of $2^N$ are required to determine an approximate global minimum using random search. This work analytically and numerically demonstrates that the selection and crossover algorithm with random initialization can reduce the required computational cost (i.e., number of iterations) for identifying an approximate global minimum to the order of $λ^N$ with $λ$ less than 2. The two best solutions, referred to as parents, are selected from a pool of randomly sampled states. Offspring generated by crossovers of the parents' states are distributed with a mean cost lower than that of the original distribution that generated the parents. It is revealed that in contrast to the mean, the variance of the cost of the offspring is asymptotically the same as that of the original distribution. Consequently, sampling from the offspring's distribution leads to a higher chance of determining an approximate global minimum than sampling from the original distribution, thereby accelerating the global search. This feature is distinct from the distribution obtained by a mixture of a large population of favorable states, which leads to a lower variance of offspring. These findings demonstrate the advantage of the crossover between two favorable states over a mixture of many favorable states for an efficient determination of an approximate global minimum.

cs.CC

Suppression of macroscopic oscillations in mixed populations of active and inactive oscillators coupled through lattice Laplacian

We consider suppression of macroscopic synchronized oscillations in mixed populations of active and inactive oscillators with local diffusive coupling, described by a lattice complex Ginzburg-Landau model with discrete Laplacian in general dimensions. Approximate expression for the stability of the non-oscillatory stationary state is derived on the basis of the generalized free energy of the system. We show that an effective wavenumber of the system determined by the spatial arrangement of the active and inactive oscillators is an decisive factor in the suppression, in addition to the ratio of active population to inactive population and relative intensity of each population. The effectiveness of the proposed theory is illustrated with a cortico-thalamic model of epileptic seizures, where active and inactive oscillators correspond to epileptic foci and healthy cerebral cortex tissue, respectively.

nlin.AO

Objective and efficient inference for couplings in neuronal networks

Inferring directional couplings from the spike data of networks is desired in various scientific fields such as neuroscience. Here, we apply a recently proposed objective procedure to the spike data obtained from the Hodgkin--Huxley type models and in vitro neuronal networks cultured in a circular structure. As a result, we succeed in reconstructing synaptic connections accurately from the evoked activity as well as the spontaneous one. To obtain the results, we invent an analytic formula approximately implementing a method of screening relevant couplings. This significantly reduces the computational cost of the screening method employed in the proposed objective procedure, making it possible to treat large-size systems as in this study.

q-bio.NC