Searcharxiv⌕ Search

arXiv subjects

Masato Okada

Publications and source records attributed to Masato Okada.

At least 73 records · Page 4Linked to original sources

Statistical Mechanics of Nonlinear On-line Learning for Ensemble Teachers

We analyze the generalization performance of a student in a model composed of nonlinear perceptrons: a true teacher, ensemble teachers, and the student. We calculate the generalization error of the student analytically or numerically using statistical mechanics in the framework of on-line learning. We treat two well-known learning rules: Hebbian learning and perceptron learning. As a result, it is proven that the nonlinear model shows qualitatively different behaviors from the linear model. Moreover, it is clarified that Hebbian learning and perceptron learning show qualitatively different behaviors from each other. In Hebbian learning, we can analytically obtain the solutions. In this case, the generalization error monotonically decreases. The steady value of the generalization error is independent of the learning rate. The larger the number of teachers is and the more variety the ensemble teachers have, the smaller the generalization error is. In perceptron learning, we have to numerically obtain the solutions. In this case, the dynamical behaviors of the generalization error are non-monotonic. The smaller the learning rate is, the larger the number of teachers is; and the more variety the ensemble teachers have, the smaller the minimum value of the generalization error is.

cs.LG↗

Retrieval of branching sequences in associative memory model with common external input and bias input

We investigate a recurrent neural network model with common external and bias inputs that can retrieve branching sequences. Retrieval of memory sequences is one of the most important functions of the brain. A lot of research has been done on neural networks that process memory sequences. Most of it has focused on fixed memory sequences. However, many animals can remember and recall branching sequences. Therefore, we propose an associative memory model that can retrieve branching sequences. Our model has bias input and common external input. Kawamura and Okada reported that common external input enables sequential memory retrieval in an associative memory model with auto- and weak cross-correlation connections. We show that retrieval processes along branching sequences are controllable with both the bias input and the common external input. To analyze the behaviors of our model, we derived the macroscopic dynamical description as a probability density function. The results obtained by our theory agree with those obtained by computer simulations.

cond-mat.dis-nn↗

Statistical Mechanics of On-line Learning when a Moving Teacher Goes around an Unlearnable True Teacher

In the framework of on-line learning, a learning machine might move around a teacher due to the differences in structures or output functions between the teacher and the learning machine. In this paper we analyze the generalization performance of a new student supervised by a moving machine. A model composed of a fixed true teacher, a moving teacher, and a student is treated theoretically using statistical mechanics, where the true teacher is a nonmonotonic perceptron and the others are simple perceptrons. Calculating the generalization errors numerically, we show that the generalization errors of a student can temporarily become smaller than that of a moving teacher, even if the student only uses examples from the moving teacher. However, the generalization error of the student eventually becomes the same value with that of the moving teacher. This behavior is qualitatively different from that of a linear model.

cs.LG↗

Stochastic transitions of attractors in associative memory models with correlated noise

We investigate dynamics of recurrent neural networks with correlated noise to analyze the noise's effect. The mechanism of correlated firing has been analyzed in various models, but its functional roles have not been discussed in sufficient detail. Aoyagi and Aoki have shown that the state transition of a network is invoked by synchronous spikes. We introduce two types of noise to each neuron: thermal independent noise and correlated noise. Due to the effects of correlated noise, the correlation between neural inputs cannot be ignored, so the behavior of the network has sample dependence. We discuss two types of associative memory models: one with auto- and weak cross-correlation connections and one with hierarchically correlated patterns. The former is similar in structure to Aoyagi and Aoki's model. We show that stochastic transition can be presented by correlated rather than thermal noise. In the latter, we show stochastic transition from a memory state to a mixture state using correlated noise. To analyze the stochastic transitions, we derive a macroscopic dynamic description as a recurrence relation form of a probability density function when the correlated noise exists. Computer simulations agree with theoretical results.

cond-mat.dis-nn↗

Statistical Mechanics of Linear and Nonlinear Time-Domain Ensemble Learning

Conventional ensemble learning combines students in the space domain. In this paper, however, we combine students in the time domain and call it time-domain ensemble learning. We analyze, compare, and discuss the generalization performances regarding time-domain ensemble learning of both a linear model and a nonlinear model. Analyzing in the framework of online learning using a statistical mechanical method, we show the qualitatively different behaviors between the two models. In a linear model, the dynamical behaviors of the generalization error are monotonic. We analytically show that time-domain ensemble learning is twice as effective as conventional ensemble learning. Furthermore, the generalization error of a nonlinear model features nonmonotonic dynamical behaviors when the learning rate is small. We numerically show that the generalization performance can be improved remarkably by using this phenomenon and the divergence of students in the time domain.

cond-mat.dis-nn↗

Theory of Interaction of Memory Patterns in Layered Associative Networks

A synfire chain is a network that can generate repeated spike patterns with millisecond precision. Although synfire chains with only one activity propagation mode have been intensively analyzed with several neuron models, those with several stable propagation modes have not been thoroughly investigated. By using the leaky integrate-and-fire neuron model, we constructed a layered associative network embedded with memory patterns. We analyzed the network dynamics with the Fokker-Planck equation. First, we addressed the stability of one memory pattern as a propagating spike volley. We showed that memory patterns propagate as pulse packets. Second, we investigated the activity when we activated two different memory patterns. Simultaneous activation of two memory patterns with the same strength led the propagating pattern to a mixed state. In contrast, when the activations had different strengths, the pulse packet converged to a two-peak state. Finally, we studied the effect of the preceding pulse packet on the following pulse packet. The following pulse packet was modified from its original activated memory pattern, and it converged to a two-peak state, mixed state or non-spike state depending on the time interval.

q-bio.NC↗

Statistical Mechanics of Time Domain Ensemble Learning

Conventional ensemble learning combines students in the space domain. On the other hand, in this paper we combine students in the time domain and call it time domain ensemble learning. In this paper, we analyze the generalization performance of time domain ensemble learning in the framework of online learning using a statistical mechanical method. We treat a model in which both the teacher and the student are linear perceptrons with noises. Time domain ensemble learning is twice as effective as conventional space domain ensemble learning.

cond-mat.stat-mech↗

Statistical mechanics of lossy compression using multilayer perceptrons

Statistical mechanics is applied to lossy compression using multilayer perceptrons for unbiased Boolean messages. We utilize a tree-like committee machine (committee tree) and tree-like parity machine (parity tree) whose transfer functions are monotonic. For compression using committee tree, a lower bound of achievable distortion becomes small as the number of hidden units K increases. However, it cannot reach the Shannon bound even where K -> infty. For a compression using a parity tree with K >= 2 hidden units, the rate distortion function, which is known as the theoretical limit for compression, is derived where the code length becomes infinity.

cond-mat.stat-mech↗

Statistical Mechanics of Online Learning for Ensemble Teachers

We analyze the generalization performance of a student in a model composed of linear perceptrons: a true teacher, ensemble teachers, and the student. Calculating the generalization error of the student analytically using statistical mechanics in the framework of on-line learning, it is proven that when learning rate $η<1$, the larger the number $K$ and the variety of the ensemble teachers are, the smaller the generalization error is. On the other hand, when $η>1$, the properties are completely reversed. If the variety of the ensemble teachers is rich enough, the direction cosine between the true teacher and the student becomes unity in the limit of $η\to 0$ and $K \to \infty$.

physics.soc-ph↗

Theory of Recurrent Neural Network with Common Synaptic Inputs

We discuss the effects of common synaptic inputs in a recurrent neural network. Because of the effects of these common synaptic inputs, the correlation between neural inputs cannot be ignored, and thus the network exhibits sample dependence. Networks of this type do not have well-defined thermodynamic limits, and self-averaging breaks down. We therefore need to develop a suitable theory without relying on these common properties. While the effects of the common synaptic inputs have been analyzed in layered neural networks, it was apparently difficult to analyze these effects in recurrent neural networks due to feedback connections. We investigated a sequential associative memory model as an example of recurrent networks and succeeded in deriving a macroscopic dynamical description as a recurrence relation form of a probability density function.

cond-mat.dis-nn↗

Generating functional analysis of CDMA detection dynamics

We investigate the detection dynamics of the parallel interference canceller (PIC) for code-division multiple-access (CDMA) multiuser detection, applied to a randomly spread, fully syncronous base-band uncoded CDMA channel model with additive white Gaussian noise (AWGN) under perfect power control in the large-system limit. It is known that the predictions of the density evolution (DE) can fairly explain the detection dynamics only in the case where the detection dynamics converge. At transients, though, the predictions of DE systematically deviate from computer simulation results. Furthermore, when the detection dynamics fail to convergence, the deviation of the predictions of DE from the results of numerical experiments becomes large. As an alternative, generating functional analysis (GFA) can take into account the effect of the Onsager reaction term exactly and does not need the Gaussian assumption of the local field. We present GFA to evaluate the detection dynamics of PIC for CDMA multiuser detection. The predictions of GFA exhibits good consistency with the computer simulation result for any condition, even if the dynamics fail to convergence.

cond-mat.stat-mech↗

Analysis of on-line learning when a moving teacher goes around a true teacher

In the framework of on-line learning, a learning machine might move around a teacher due to the differences in structures or output functions between the teacher and the learning machine or due to noises. The generalization performance of a new student supervised by a moving machine has been analyzed. A model composed of a true teacher, a moving teacher and a student that are all linear perceptrons with noises has been treated analytically using statistical mechanics. It has been proven that the generalization errors of a student can be smaller than that of a moving teacher, even if the student only uses examples from the moving teacher.

physics.soc-ph↗

Short-term Synaptic Depression Improves Error-correcting Ability in Cortical Circuits

Synaptic connections are known to change dynamically. High-frequency presynaptic inputs induce decrease of synaptic weights. This process is known as short-term synaptic depression. The synaptic depression controls a gain for presynaptic inputs. However, it remains a controversial issue what are functional roles of this gain control. We propose a new hypothesis that one of the functional roles is to enlarge basins of attraction. To verify this hypothesis, we employ a binary discrete-time associative memory model which consists of excitatory and inhibitory neurons. It is known that the excitatory-inhibitory balance controls an overall activity of the network. The synaptic depression might incorporate an activity control mechanism. Using a mean-field theory and computer simulations, we find that the basins of attraction are enlarged whereas the storage capacity does not change. Furthermore, the excitatory-inhibitory balance and the synaptic depression work cooperatively. This result suggests that the synaptic depression works to improve an error-correcting ability in cortical circuits.

cond-mat.dis-nn↗

Correlation of Firing in Layered Associative Neural Networks

There is growing interest in a phenomenon called the ``synfire chain'', in which firings of neurons propagate from pool to pool in the chain. The mechanism of the synfire chain has been analyzed by many resarchers. Keeping the synfire chain phenomenon in mind, we investigate a layered associative memory neural network model, in which patterns are embedded in connections between neurons. In this model, we also include uniform noise in connections, which induces common input in the next layer. Such common input in layers generate correlated firings of neurons. We theoretically obtain the evolution of retrieval states in the case of infinite pattern loading. We find that the overlap between patterns and neuronal states is not given as a deterministic quantity, but is described by a probability distribution defined over the emsemble of synaptic matrices. Our simulation results are in excellent agreement with theoretical calculations.

cond-mat.other↗

Residual Energies after Slow Quantum Annealing

Features of the residual energy after the quantum annealing are investigated. The quantum annealing method exploits quantum fluctuations to search the ground state of classical disordered Hamiltonian. If the quantum fluctuation is reduced sufficiently slowly and linearly by the time, the residual energy after the quantum annealing falls as the inverse square of the annealing time. We show this feature of the residual energy by numerical calculations for small-sized systems and derive it on the basis of the quantum adiabatic theorem.

cond-mat.dis-nn↗

Theory of localized synfire chain

Neuron is a noisy information processing unit and conventional view is that information in the cortex is carried on the rate of neurons spike emission. More recent studies on the activity propagation through the homogeneous network have demonstrated that signals can be transmitted with millisecond fidelity; this model is called the Synfire chain and suggests the possibility of the spatio-temporal coding. However, the more biologically realistic, structured feedforward network generates spatially distributed inputs. It results in the difference of spike timing. This poses a question on how the spatial structure of a network effect the stability of spatio-temporal spike patterns, and the speed of a spike packet propagation. By formulating the Fokker-Planck equation for the feedforwardly coupled network with Mexican-Hat type connectivity, we show the stability of localized spike packet and existence of Multi-stable phase where both uniform and localized spike packets are stable depending on the initial input structure. The Multi-stable phase enables us to show that a spike pattern, or the information of its own, determines the propagation speed.

q-bio.NC↗

Bifurcation analysis in an associative memory model

We previously reported the chaos induced by the frustration of interaction in a non-monotonic sequential associative memory model, and showed the chaotic behaviors at absolute zero. We have now analyzed bifurcation in a stochastic system, namely a finite-temperature model of the non-monotonic sequential associative memory model. We derived order-parameter equations from the stochastic microscopic equations. Two-parameter bifurcation diagrams obtained from those equations show the coexistence of attractors, which do not appear at absolute zero, and the disappearance of chaos due to the temperature effect.

cond-mat.dis-nn↗

Analysis of ensemble learning using simple perceptrons based on online learning theory

Ensemble learning of $K$ nonlinear perceptrons, which determine their outputs by sign functions, is discussed within the framework of online learning and statistical mechanics. One purpose of statistical learning theory is to theoretically obtain the generalization error. This paper shows that ensemble generalization error can be calculated by using two order parameters, that is, the similarity between a teacher and a student, and the similarity among students. The differential equations that describe the dynamical behaviors of these order parameters are derived in the case of general learning rules. The concrete forms of these differential equations are derived analytically in the cases of three well-known rules: Hebbian learning, perceptron learning and AdaTron learning. Ensemble generalization errors of these three rules are calculated by using the results determined by solving their differential equations. As a result, these three rules show different characteristics in their affinity for ensemble learning, that is ``maintaining variety among students." Results show that AdaTron learning is superior to the other two rules with respect to that affinity.

cond-mat.dis-nn↗