SearcharxivSearch

arXiv subjects

Jun Ohkubo

Publications and source records attributed to Jun Ohkubo.

At least 19 recordsLinked to original sources

Learning Koopman operators for coupled systems via information on governing equations of subsystems

Nonlinear coupled systems are ubiquitous in science and engineering. The analysis and modeling of such systems are challenging due to their high dimensionality and complex interactions among subsystems. In recent years, operator-theoretic methods based on the Koopman operator have attracted attention as a powerful tool for analyzing and modeling nonlinear dynamical systems. Extended dynamic mode decomposition (EDMD) is one of the most popular methods for approximating the Koopman operator. However, EDMD is a purely data-driven method, and it may be unstable and inaccurate for coupled systems under limited data availability. In this paper, we propose a method to construct a finite-dimensional Koopman approximation for coupled systems using the differential equations governing each subsystem. The proposed method aims to improve data efficiency by using the known subsystem dynamics as prior information and learning the coupling-induced correction from a limited number of snapshots. We also demonstrate its effectiveness through numerical experiments on coupled oscillator systems.

cs.LG

Tensor-based computation of the Koopman generator via operator logarithm

Identifying governing equations of nonlinear dynamical systems from data is challenging. While sparse identification of nonlinear dynamics (SINDy) and its extensions are widely used for system identification, operator-logarithm approaches use the logarithm to avoid time differentiation, enabling larger sampling intervals. However, they still suffer from the curse of dimensionality. Then, we propose a data-driven method to compute the Koopman generator in a low-rank tensor train (TT) format by taking logarithms of Koopman eigenvalues while preserving the TT format. Experiments on 4-dimensional Lotka-Volterra and 10-dimensional Lorenz-96 systems show accurate recovery of vector field coefficients and scalability to higher-dimensional systems.

cs.LG

Extraction of linearized models from pre-trained networks via knowledge distillation

Recent developments in hardware, such as photonic integrated circuits and optical devices, are driving demand for research on constructing machine learning architectures tailored for linear operations. Hence, it is valuable to explore methods for constructing learning machines with only linear operations after simple nonlinear preprocessing. In this study, we propose a framework to extract a linearized model from a pre-trained neural network for classification tasks by integrating Koopman operator theory with knowledge distillation. Numerical demonstrations on the MNIST and the Fashion-MNIST datasets reveal that the proposed model consistently outperforms the conventional least-squares-based Koopman approximation in both classification accuracy and numerical stability.

cs.LG

Neural network initialization with nonlinear characteristics and information on hierarchical features

Initialization of neural network parameters, such as weights and biases, has a crucial impact on learning performance; if chosen well, we can even avoid the need for additional training with backpropagation. For example, algorithms based on the ridgelet transform or the SWIM (sampling where it matters) concept have been proposed for initialization. On the other hand, some works show hierarchical features in trained neural networks; neural networks tend to learn coarse information in the early-stage hidden layers. In this work, we investigate the effects of utilizing information on the hierarchical features in the initialization of neural networks. Hence, we propose a framework that adjusts the scale factors in the SWIM algorithm to capture low-frequency components in the early-stage hidden layers and to represent high-frequency components in the late-stage hidden layers. Numerical experiments on a one-dimensional regression task and the MNIST classification task demonstrate that the proposed method outperforms the conventional initialization algorithms. This work clarifies the importance of intrinsic hierarchical features in learning neural networks, and the finding yields an effective parameter initialization strategy that enhances their training performance.

cs.LG

Koopman operator-based discussion on partial observation in stochastic systems

It is sometimes difficult to achieve a complete observation for a full set of observables, and partial observations are necessary. For deterministic systems, the Mori-Zwanzig formalism provides a theoretical framework for handling partial observations. Recently, data-driven algorithms based on the Koopman operator theory have made significant progress, and there is a discussion to connect the Mori-Zwanzig formalism with the Koopman operator theory. In this work, we discuss the effects of partial observation in stochastic systems using the Koopman operator theory. The discussion clarifies the importance of distinguishing the state space and the function space in stochastic systems. Even in stochastic systems, the delay-embedding technique is beneficial for partial observation, and several numerical experiments show a power-law behavior of error with respect to the amplitude of the additive noise. We also discuss the relation between the exponent of the power-law behavior and the effects of partial observation.

cs.LG

Stochastic modeling of deterministic laser chaos using generator extended dynamic mode decomposition

Recently, chaotic phenomena in laser dynamics have attracted much attention to its applied aspects, and a synchronization phenomenon, leader-laggard relationship, in time-delay coupled lasers has been used in reinforcement learning. In the present paper, we discuss the possibility of capturing the essential stochasticity of the leader-laggard relationship; in nonlinear science, it is known that coarse-graining allows one to derive stochastic models from deterministic systems. We derive stochastic models with the aid of the Koopman operator approach, and we clarify that the low-pass filtered data is enough to recover the essential features of the original deterministic chaos, such as peak shifts in the distribution of being the leader and a power-law behavior in the distribution of switching-time intervals. We also confirm that the derived stochastic model works well in reinforcement learning tasks, i.e., multi-armed bandit problems, as with the original laser chaos system.

physics.app-ph

Permutation of Tensor-Train Cores for Computing Moments on Stochastic Differential Equations

Tensor networks, particularly the tensor train (TT) format, have emerged as powerful tools for high-dimensional computations in physics and computer science. In solving coupled differential equations, such as those arising from stochastic differential equations (SDEs) via duality relations, ordering the TT cores significantly influences numerical accuracy. In this study, we first systematically investigate how different orderings of the TT cores affect the accuracy of computed moments using the duality relation in stochastic processes. Through numerical experiments on a two-body interaction model, we demonstrate that specific orderings of the TT cores yield lower relative errors, particularly when they align with the underlying interaction structure of the system. Motivated by these findings, we then propose a novel quantitative measure, $score$, which is defined based on an ordering of the TT cores and an SDE parameter set. While the score is independent of the accuracy of moments to compute by definition, we assess its effectiveness by evaluating the accuracy of computed moments. Our results indicate that orderings that minimize the score tend to yield higher accuracy. This study provides insights into optimizing orderings of the TT cores, which is essential for efficient and reliable high-dimensional simulations of stochastic processes.

physics.comp-ph

Integrated utilization of equations and small dataset in the Koopman operator: applications to forward and inverse problems

In recent years, there has been a growing interest in data-driven approaches in physics, such as extended dynamic mode decomposition (EDMD). The EDMD algorithm focuses on nonlinear time-evolution systems, and the constructed Koopman matrix yields the next-time prediction with only linear matrix-product operations. Note that data-driven approaches generally require a large dataset. However, assume that one has some prior knowledge, even if it may be ambiguous. Then, one could achieve sufficient learning from only a small dataset by taking advantage of the prior knowledge. This paper yields methods for incorporating ambiguous prior knowledge into the EDMD algorithm. The ambiguous prior knowledge in this paper corresponds to the underlying time-evolution equations with unknown parameters. First, we apply the proposed method to forward problems, i.e., prediction tasks. Second, we propose a scheme to apply the proposed method to inverse problems, i.e., parameter estimation tasks. We demonstrate the learning with only a small dataset using guiding examples, i.e., the Duffing and the van der Pol systems.

cs.LG

Koopman analysis of combinatorial optimization problems with replica exchange Monte Carlo method

Combinatorial optimization problems play crucial roles in real-world applications, and many studies from a physics perspective have contributed to specialized hardware for high-speed computation. However, some combinatorial optimization problems are easy to solve, and others are not. Hence, the qualification of the difficulty in problem-solving will be beneficial. In this paper, we employ the Koopman analysis for multiple time-series data from the replica exchange Monte Carlo method. After proposing a quantity that aggregates the information of the multiple time-series data, we performed numerical experiments. The results indicate a negative correlation between the proposed quantity and the ability of the solution search.

physics.app-ph

Aspects of importance sampling in parameter selection for neural networks using ridgelet transform

The choice of parameters in neural networks is crucial in the performance, and an oracle distribution derived from the ridgelet transform enables us to obtain suitable initial parameters. In other words, the distribution of parameters is connected to the integral representation of target functions. The oracle distribution allows us to avoid the conventional backpropagation learning process; only a linear regression is enough to construct the neural network in simple cases. This study provides a new look at the oracle distributions and ridgelet transforms, i.e., an aspect of importance sampling. In addition, we propose extensions of the parameter sampling methods. We demonstrate the aspect of importance sampling and the proposed sampling algorithms via one-dimensional and high-dimensional examples; the results imply that the magnitude of weight parameters could be more crucial than the intercept parameters.

cs.LG

Extraction of nonlinearity in neural networks with Koopman operator

Nonlinearity plays a crucial role in deep neural networks. In this paper, we investigate the degree to which the nonlinearity of the neural network is essential. For this purpose, we employ the Koopman operator, extended dynamic mode decomposition, and the tensor-train format. The Koopman operator approach has been recently developed in physics and nonlinear sciences; the Koopman operator deals with the time evolution in the observable space instead of the state space. Since we can replace the nonlinearity in the state space with the linearity in the observable space, it is a hopeful candidate for understanding complex behavior in nonlinear systems. Here, we analyze learned neural networks for the classification problems. As a result, the replacement of the nonlinear middle layers with the Koopman matrix yields enough accuracy in numerical experiments. In addition, we confirm that the pruning of the Koopman matrix gives sufficient accuracy even at high compression ratios. These results indicate the possibility of extracting some features in the neural networks with the Koopman operator approach.

cs.LG

Improvement of system identification of stochastic systems via Koopman generator and locally weighted expectation

The estimation of equations from data is of interest in physics. One of the famous methods is the sparse identification of nonlinear dynamics (SINDy), which utilizes sparse estimation techniques to estimate equations from data. Recently, a method based on the Koopman operator has been developed; the generator extended dynamic mode decomposition (gEDMD) estimates a time evolution generator of dynamical and stochastic systems. However, a naive application of the gEDMD algorithm cannot work well for stochastic differential equations because of the noise effects in the data. Hence, the estimation based on conditional expectation values, in which we approximate the first and second derivatives on each coordinate, is practical. A naive approach is the usage of locally weighted expectations. We show that the naive locally weighted expectation is insufficient because of the nonlinear behavior of the underlying system. For improvement, we apply the clustering method in two ways; one is to reduce the effective number of data, and the other is to capture local information more accurately. We demonstrate the improvement of the proposed method for the double-well potential system with state-dependent noise.

math.DS

Attention-Enhanced Reservoir Computing

Photonic reservoir computing has been successfully utilized in time-series prediction as the need for hardware implementations has increased. Prediction of chaotic time series remains a significant challenge, an area where the conventional reservoir computing framework encounters limitations of prediction accuracy. We introduce an attention mechanism to the reservoir computing model in the output stage. This attention layer is designed to prioritize distinct features and temporal sequences, thereby substantially enhancing the prediction accuracy. Our results show that a photonic reservoir computer enhanced with the attention mechanism exhibits improved prediction capabilities for smaller reservoirs. These advancements highlight the transformative possibilities of reservoir computing for practical applications where accurate prediction of chaotic time series is crucial.

cs.ET

Compression of the Koopman matrix for nonlinear physical models via hierarchical clustering

Machine learning methods allow the prediction of nonlinear dynamical systems from data alone. The Koopman operator is one of them, which enables us to employ linear analysis for nonlinear dynamical systems. The linear characteristics of the Koopman operator are hopeful to understand the nonlinear dynamics and perform rapid predictions. The extended dynamic mode decomposition (EDMD) is one of the methods to approximate the Koopman operator as a finite-dimensional matrix. In this work, we propose a method to compress the Koopman matrix using hierarchical clustering. Numerical demonstrations for the cart-pole model and comparisons with the conventional singular value decomposition (SVD) are shown; the results indicate that the hierarchical clustering performs better than the naive SVD compressions.

cs.LG

Characterization of Locality in Spin States and Forced Moves for Optimizations

Ising formulations are widely utilized to solve combinatorial optimization problems, and a variety of quantum or semiconductor-based hardware has recently been made available. In combinatorial optimization problems, the existence of local minima in energy landscapes is problematic to use to seek the global minimum. We note that the aim of the optimization is not to obtain exact samplings from the Boltzmann distribution, and there is thus no need to satisfy detailed balance conditions. In light of this fact, we develop an algorithm to get out of the local minima efficiently while it does not yield the exact samplings. For this purpose, we utilize a feature that characterizes locality in the current state, which is easy to obtain with a type of specialized hardware. Furthermore, as the proposed algorithm is based on a rejection-free algorithm, the computational cost is low. In this work, after presenting the details of the proposed algorithm, we report the results of numerical experiments that demonstrate the effectiveness of the proposed feature and algorithm.

physics.app-ph

Redundant basis interpretation of Doi-Peliti method and an application

The Doi-Peliti method is effective for investigating classical stochastic processes, and it has wide applications, including field theoretic approaches. Furthermore, it is applicable not only to master equations but also to stochastic differential equations; one can derive a kind of discrete process from stochastic differential equations. A remarkable fact is that the Doi-Peliti method is related to a different analytical approach, i.e., generating function. The connection with the generating function approach helps to understand the derivation of discrete processes from stochastic differential equations. Here, a redundant basis interpretation for the Doi-Peliti method is proposed, which enables us to derive different types of discrete processes. The conventional correspondence with the generating function approach is also extended. The proposed extensions give us a new tool to study stochastic differential equations. As an application of the proposed interpretation, we perform numerical experiments for a finite-state approximation of the derived discrete process from the noisy van der Pol system; the redundant basis yields reasonable results compared with the conventional discrete process with the same number of states.

cond-mat.stat-mech

Embedding stochastic differential equations into neural networks via dual processes

We propose a new approach to constructing a neural network for predicting expectations of stochastic differential equations. The proposed method does not need data sets of inputs and outputs; instead, the information obtained from the time-evolution equations, i.e., the corresponding dual process, is directly compared with the weights in the neural network. As a demonstration, we construct neural networks for the Ornstein-Uhlenbeck process and the noisy van der Pol system. The remarkable feature of learned networks with the proposed method is the accuracy of inputs near the origin. Hence, it would be possible to avoid the overfitting problem because the learned network does not depend on training data sets.

cs.LG

Statistics for stochastic differential equations and approximations of resolvent

The numerical evaluation of statistics plays a crucial role in statistical physics and its applied fields. It is possible to evaluate the statistics for a stochastic differential equation with Gaussian white noise via the corresponding backward Kolmogorov equation. The important notice is that there is no need to obtain the solution of the backward Kolmogorov equation on the whole domain; it is enough to evaluate a value of the solution at a certain point that corresponds to the initial coordinate for the stochastic differential equation. For this aim, an algorithm based on combinatorics has recently been developed. In this paper, we discuss a higher-order approximation of resolvent, and an algorithm based on a second-order approximation is proposed. The proposed algorithm shows a second-order convergence. Furthermore, the convergence property of the naive algorithms naturally leads to extrapolation methods; they work well to calculate a more accurate value with fewer computational costs. The proposed method is demonstrated with the Ornstein-Uhlenbeck process and the noisy van der Pol system.

math.NA