SearcharxivSearch

arXiv subjects

Seiji Miyoshi

Publications and source records attributed to Seiji Miyoshi.

13 recordsLinked to original sources

Exact equality of the MSEs for two types of nonlinear adaptive systems: Saturation and dead-zone types

Adaptive signal processing systems, commonly utilized in applications such as active noise control and acoustic echo cancellation, often encompass nonlinearities due to hardware components such as loudspeakers, microphones, and amplifiers. Examining the impact of these nonlinearities on the overall performance of adaptive systems is critically important. In this study, we employ a statistical-mechanical method to investigate the behaviors of adaptive systems, each containing an unknown system with a nonlinearity in its output. We specifically address two types of nonlinearity: saturation and dead-zone types. We analyze both the dynamic and steady-state behaviors of these systems under the effect of such nonlinearities. Our findings indicate that when the saturation value is equal to the dead-zone width, the mean square errors (MSEs) in steady states are identical for both nonlinearity types. Furthermore, we derive a self-consistent equation to obtain the saturation value and dead-zone width that maximize the steady-state MSE. We theoretically clarify that these values depend on neither the step size nor the variance of background noise.

eess.SP

Statistical-mechanical analysis of adaptive filter with clipping saturation-type nonlinearity

In most practical adaptive signal processing systems, e.g., active noise control, active vibration control, and acoustic echo cancellation, substantial nonlinearities that cannot be neglected exist. In this paper, we analyze the behaviors of an adaptive system in which the output of the adaptive filter has the clipping saturation-type nonlinearity by a statistical-mechanical method. We discuss the dynamical and steady-state behaviors of the adaptive system by asymptotic analysis, steady-state analysis, and numerical calculation. As a result, it has become clear that the saturation value has the critical point at which the system's mean-square stability and instability switch. The obtained theory well explains the strange behaviors around the critical point observed in the computer simulation. Finally, the exact value of the critical point is also derived.

eess.SP

Statistical Mechanics of On-line Ensemble Teacher Learning through a Novel Perceptron Learning Rule

In ensemble teacher learning, ensemble teachers have only uncertain information about the true teacher, and this information is given by an ensemble consisting of an infinite number of ensemble teachers whose variety is sufficiently rich. In this learning, a student learns from an ensemble teacher that is iteratively selected randomly from a pool of many ensemble teachers. An interesting point of ensemble teacher learning is the asymptotic behavior of the student to approach the true teacher by learning from ensemble teachers. The student performance is improved by using the Hebbian learning rule in the learning. However, the perceptron learning rule cannot improve the student performance. On the other hand, we proposed a perceptron learning rule with a margin. This learning rule is identical to the perceptron learning rule when the margin is zero and identical to the Hebbian learning rule when the margin is infinity. Thus, this rule connects the perceptron learning rule and the Hebbian learning rule continuously through the size of the margin. Using this rule, we study changes in the learning behavior from the perceptron learning rule to the Hebbian learning rule by considering several margin sizes. From the results, we show that by setting a margin of kappa > 0, the effect of an ensemble appears and becomes significant when a larger margin kappa is used.

cond-mat.dis-nn

Effect of Slow Switching in On-line Learning for Ensemble Teachers

We have analyzed the generalization performance of a student which slowly switches ensemble teachers. By calculating the generalization error analytically using statistical mechanics in the framework of on-line learning, we show that the dynamical behaviors of generalization error have the periodicity that is synchronized with the switching period and the behaviors differ with the number of ensemble teachers. Furthermore, we show that the smaller the switching period is, the larger the difference is.

physics.soc-ph

Statistical Mechanics of Nonlinear On-line Learning for Ensemble Teachers

We analyze the generalization performance of a student in a model composed of nonlinear perceptrons: a true teacher, ensemble teachers, and the student. We calculate the generalization error of the student analytically or numerically using statistical mechanics in the framework of on-line learning. We treat two well-known learning rules: Hebbian learning and perceptron learning. As a result, it is proven that the nonlinear model shows qualitatively different behaviors from the linear model. Moreover, it is clarified that Hebbian learning and perceptron learning show qualitatively different behaviors from each other. In Hebbian learning, we can analytically obtain the solutions. In this case, the generalization error monotonically decreases. The steady value of the generalization error is independent of the learning rate. The larger the number of teachers is and the more variety the ensemble teachers have, the smaller the generalization error is. In perceptron learning, we have to numerically obtain the solutions. In this case, the dynamical behaviors of the generalization error are non-monotonic. The smaller the learning rate is, the larger the number of teachers is; and the more variety the ensemble teachers have, the smaller the minimum value of the generalization error is.

cs.LG

Statistical Mechanics of On-line Learning when a Moving Teacher Goes around an Unlearnable True Teacher

In the framework of on-line learning, a learning machine might move around a teacher due to the differences in structures or output functions between the teacher and the learning machine. In this paper we analyze the generalization performance of a new student supervised by a moving machine. A model composed of a fixed true teacher, a moving teacher, and a student is treated theoretically using statistical mechanics, where the true teacher is a nonmonotonic perceptron and the others are simple perceptrons. Calculating the generalization errors numerically, we show that the generalization errors of a student can temporarily become smaller than that of a moving teacher, even if the student only uses examples from the moving teacher. However, the generalization error of the student eventually becomes the same value with that of the moving teacher. This behavior is qualitatively different from that of a linear model.

cs.LG

Statistical Mechanics of Linear and Nonlinear Time-Domain Ensemble Learning

Conventional ensemble learning combines students in the space domain. In this paper, however, we combine students in the time domain and call it time-domain ensemble learning. We analyze, compare, and discuss the generalization performances regarding time-domain ensemble learning of both a linear model and a nonlinear model. Analyzing in the framework of online learning using a statistical mechanical method, we show the qualitatively different behaviors between the two models. In a linear model, the dynamical behaviors of the generalization error are monotonic. We analytically show that time-domain ensemble learning is twice as effective as conventional ensemble learning. Furthermore, the generalization error of a nonlinear model features nonmonotonic dynamical behaviors when the learning rate is small. We numerically show that the generalization performance can be improved remarkably by using this phenomenon and the divergence of students in the time domain.

cond-mat.dis-nn

Statistical Mechanics of Time Domain Ensemble Learning

Conventional ensemble learning combines students in the space domain. On the other hand, in this paper we combine students in the time domain and call it time domain ensemble learning. In this paper, we analyze the generalization performance of time domain ensemble learning in the framework of online learning using a statistical mechanical method. We treat a model in which both the teacher and the student are linear perceptrons with noises. Time domain ensemble learning is twice as effective as conventional space domain ensemble learning.

cond-mat.stat-mech

Statistical Mechanics of Online Learning for Ensemble Teachers

We analyze the generalization performance of a student in a model composed of linear perceptrons: a true teacher, ensemble teachers, and the student. Calculating the generalization error of the student analytically using statistical mechanics in the framework of on-line learning, it is proven that when learning rate $η<1$, the larger the number $K$ and the variety of the ensemble teachers are, the smaller the generalization error is. On the other hand, when $η>1$, the properties are completely reversed. If the variety of the ensemble teachers is rich enough, the direction cosine between the true teacher and the student becomes unity in the limit of $η\to 0$ and $K \to \infty$.

physics.soc-ph

Analysis of on-line learning when a moving teacher goes around a true teacher

In the framework of on-line learning, a learning machine might move around a teacher due to the differences in structures or output functions between the teacher and the learning machine or due to noises. The generalization performance of a new student supervised by a moving machine has been analyzed. A model composed of a true teacher, a moving teacher and a student that are all linear perceptrons with noises has been treated analytically using statistical mechanics. It has been proven that the generalization errors of a student can be smaller than that of a moving teacher, even if the student only uses examples from the moving teacher.

physics.soc-ph

Analysis of ensemble learning using simple perceptrons based on online learning theory

Ensemble learning of $K$ nonlinear perceptrons, which determine their outputs by sign functions, is discussed within the framework of online learning and statistical mechanics. One purpose of statistical learning theory is to theoretically obtain the generalization error. This paper shows that ensemble generalization error can be calculated by using two order parameters, that is, the similarity between a teacher and a student, and the similarity among students. The differential equations that describe the dynamical behaviors of these order parameters are derived in the case of general learning rules. The concrete forms of these differential equations are derived analytically in the cases of three well-known rules: Hebbian learning, perceptron learning and AdaTron learning. Ensemble generalization errors of these three rules are calculated by using the results determined by solving their differential equations. As a result, these three rules show different characteristics in their affinity for ensemble learning, that is ``maintaining variety among students." Results show that AdaTron learning is superior to the other two rules with respect to that affinity.

cond-mat.dis-nn

Storage Capacity Diverges with Synaptic Efficiency in an Associative Memory Model with Synaptic Delay and Pruning

It is known that storage capacity per synapse increases by synaptic pruning in the case of a correlation-type associative memory model. However, the storage capacity of the entire network then decreases. To overcome this difficulty, we propose decreasing the connecting rate while keeping the total number of synapses constant by introducing delayed synapses. In this paper, a discrete synchronous-type model with both delayed synapses and their prunings is discussed as a concrete example of the proposal. First, we explain the Yanai-Kim theory by employing the statistical neurodynamics. This theory involves macrodynamical equations for the dynamics of a network with serial delay elements. Next, considering the translational symmetry of the explained equations, we re-derive macroscopic steady state equations of the model by using the discrete Fourier transformation. The storage capacities are analyzed quantitatively. Furthermore, two types of synaptic prunings are treated analytically: random pruning and systematic pruning. As a result, it becomes clear that in both prunings, the storage capacity increases as the length of delay increases and the connecting rate of the synapses decreases when the total number of synapses is constant. Moreover, an interesting fact becomes clear: the storage capacity asymptotically approaches $2/π$ due to random pruning. In contrast, the storage capacity diverges in proportion to the logarithm of the length of delay by systematic pruning and the proportion constant is $4/π$. These results theoretically support the significance of pruning following an overgrowth of synapses in the brain and strongly suggest that the brain prefers to store dynamic attractors such as sequences and limit cycles rather than equilibrium states.

cond-mat.dis-nn

Associative Memory by Recurrent Neural Networks with Delay Elements

The synapses of real neural systems seem to have delays. Therefore, it is worthwhile to analyze associative memory models with delayed synapses. Thus, a sequential associative memory model with delayed synapses is discussed, where a discrete synchronous updating rule and a correlation learning rule are employed. Its dynamic properties are analyzed by the statistical neurodynamics. In this paper, we first re-derive the Yanai-Kim theory, which involves macrodynamical equations for the dynamics of the network with serial delay elements. Since their theory needs a computational complexity of $O(L^4t)$ to obtain the macroscopic state at time step t where L is the length of delay, it is intractable to discuss the macroscopic properties for a large L limit. Thus, we derive steady state equations using the discrete Fourier transformation, where the computational complexity does not formally depend on L. We show that the storage capacity $α_C$ is in proportion to the delay length L with a large L limit, and the proportion constant is 0.195, i.e., $α_C = 0.195 L$. These results are supported by computer simulations.

cond-mat.dis-nn