Searcharxiv⌕ Search

arXiv subjects

Masato Okada

Publications and source records attributed to Masato Okada.

At least 37 records · Page 2Linked to original sources

Effective implementation of $l_0$-Regularised Compressed Sensing with Chaotic-Amplitude-Controlled Coherent Ising Machines

Coherent Ising Machine (CIM) is a network of optical parametric oscillators that can solve large-scale combinatorial optimisation problems by finding the ground state of an Ising Hamiltonian. As a practical application of CIM, Aonishi et al., proposed a quantum-classical hybrid system to solve optimisation problems of $l_0$-regularisation-based compressed sensing. In the hybrid system, the CIM was an open-loop system without an amplitude control feedback loop. In this case, the hybrid system is enhanced by using a closed-loop CIM to achieve chaotic behaviour around the target amplitude, which would enable escaping from local minima in the energy landscape. Both artificial and magnetic resonance image data were used for the testing of our proposed closed-loop system. Compared with the open-loop system, the results of this study demonstrate an improved degree of accuracy and a wider range of effectiveness.

quant-ph↗

Intrinsic regularization effect in Bayesian nonlinear regression scaled by observed data

Occam's razor is a guiding principle that models should be simple enough to describe observed data. While Bayesian model selection (BMS) embodies it by the intrinsic regularization effect (IRE), how observed data scale the IRE has not been fully understood. In the nonlinear regression with conditionally independent observations, we show that the IRE is scaled by observations' fineness, defined by the amount and quality of observed data. We introduce an observable that quantifies the IRE, referred to as the Bayes specific heat, inspired by the correspondence between statistical inference and statistical physics. We derive its scaling relation to observations' fineness. We demonstrate that the optimal model chosen by the BMS changes at critical values of observations' fineness, accompanying the IRE's variation. The changes are from choosing a coarse-grained model to a fine-grained one as observations' fineness increases. Our findings expand an understanding of BMS's typicality when observed data are insufficient.

physics.data-an↗

Bayesian Spectral Deconvolution of X-Ray Absorption Near Edge Structure Discriminating High- and Low-Energy Domains

In this paper, we propose a Bayesian spectral deconvolution considering the properties of peaks in different energy domains. Bayesian spectral deconvolution regresses spectral data into the sum of multiple basis functions. Conventional methods use a model that treats all peaks equally. However, in X-ray absorption near edge structure (XANES) spectra, the properties of the peaks differ depending on the energy domain, and the specific energy domain of XANES is essential in condensed matter physics. We propose a model that discriminates between the low- and high-energy domains. We also propose a prior distribution that reflects the physical properties. We compare the conventional and proposed models in terms of computational efficiency, estimation accuracy, and model evidence. We demonstrate that our method effectively estimates the number of transition components in the important energy domain, on which the material scientists focus for mapping the electronic transition analysis by first-principles simulation.

stat.ME↗

L0 regularization-based compressed sensing with quantum-classical hybrid approach

L0-regularization-based compressed sensing (L0-RBCS) has the potential to outperform L1-regularization-based compressed sensing (L1-RBCS), but the optimization in L0-RBCS is difficult because it is a combinatorial optimization problem. To perform optimization in L0-RBCS, we propose a quantum-classical hybrid system consisting of a quantum machine and a classical digital processor. The coherent Ising machine (CIM) is a suitable quantum machine for this system because this optimization problem can only be solved with a densely connected network. To evaluate the performance of the CIM-classical hybrid system theoretically, a truncated Wigner stochastic differential equation (W-SDE) is introduced as a model for the network of degenerate optical parametric oscillators, and macroscopic equations are derived by applying statistical mechanics to the W-SDE. We show that the system performance in principle approaches the theoretical limit of compressed sensing and this hybrid system may exceed the estimation accuracy of L1-RBCS in actual situations, such as in magnetic resonance imaging data analysis.

quant-ph↗

Bayesian Inference on Hamiltonian Selections for Mössbauer Spectroscopy

Mössbauer spectroscopy, which provides knowledge related to electronic states in materials, has been applied to various fields such as condensed matter physics and material sciences. In conventional spectral analyses based on least-square fitting, hyperfine interactions in materials have been determined from the shape of observed spectra. In conventional spectral analyses, it is difficult to discuss the validity of the hyperfine interactions and the estimated values. We propose a spectral analysis method based on Bayesian inference for the selection of hyperfine interactions and the estimation of Mössbauer parameters. An appropriate Hamiltonian has been selected by comparing Bayesian free energy among possible Hamiltonians. We have estimated the Mössbauer parameters and evaluated their estimated values by calculating the posterior distribution of each Mössbauer parameter with confidence intervals. We have also discussed the accuracy of the spectral analyses to elucidate the noise intensity dependence of numerical experiments.

physics.comp-ph↗

Statistical Mechanical Analysis of Catastrophic Forgetting in Continual Learning with Teacher and Student Networks

When a computational system continuously learns from an ever-changing environment, it rapidly forgets its past experiences. This phenomenon is called catastrophic forgetting. While a line of studies has been proposed with respect to avoiding catastrophic forgetting, most of the methods are based on intuitive insights into the phenomenon, and their performances have been evaluated by numerical experiments using benchmark datasets. Therefore, in this study, we provide the theoretical framework for analyzing catastrophic forgetting by using teacher-student learning. Teacher-student learning is a framework in which we introduce two neural networks: one neural network is a target function in supervised learning, and the other is a learning neural network. To analyze continual learning in the teacher-student framework, we introduce the similarity of the input distribution and the input-output relationship of the target functions as the similarity of tasks. In this theoretical framework, we also provide a qualitative understanding of how a single-layer linear learning neural network forgets tasks. Based on the analysis, we find that the network can avoid catastrophic forgetting when the similarity among input distributions is small and that of the input-output relationship of the target functions is large. The analysis also suggests that a system often exhibits a characteristic phenomenon called overshoot, which means that even if the learning network has once undergone catastrophic forgetting, it is possible that the network may perform reasonably well after further learning of the current task.

stat.ML↗

Appropriate basis selection based on Bayesian inference for analyzing measured data reflecting photoelectron wave interference

In this study, we applied Bayesian inference for extended X-ray absorption fine structure (EXAFS) to select an appropriate basis from among Fourier, wavelet and advanced Fourier bases, and we extracted a radial distribution function (RDF) and physical parameters from only EXAFS signals using physical prior knowledge, which is to be realized in general in condensed systems. To evaluate our method, the well-known EXAFS spectrum of copper was used for the EXAFS data analysis. We found that the advanced Fourier basis is selected as an appropriate basis for the regression of the EXAFS signal in a quantitative way and that the estimation of the Debye-Waller factor can be robustly realized only by using the advanced Fourier basis. Bayesian inference based on minimal restrictions allows us to not only eliminate some unphysical results but also select an appropriate basis. Generally, FEFF analysis is used for estimating physical parameters such as Deby-Waller and extracting RDF. Bayesian inference enables us to simultaneously select an appropriate basis and optimized physical parameters without FEFF analysis, which results in extracting RDF from only EXAFS signals. These advantages lead to the general usage of Bayesian inference for EXAFS data analysis.

physics.data-an↗

Sparse Modeling analysis of Extended X-ray Absorption Fine Structure data using two-body expansion

Analysis of extended X-ray absorption fine structure (EXAFS) data by the use of sparse modeling is presented. We consider the two-body term in the n-body expansion of the EXAFS signal to implement the method, together with calculations of amplitudes and phase shifts to distinguish between different back-scattering elements. Within this approach no a priori assumption about the structure is used, other than the elements present inside the material. We apply the method to the experimental EXAFS signal of metals and oxides, for which we were able to extract the radial distribution function peak positions, and the Debye-Waller factor for first neighbors.

physics.data-an↗

Fast Bayesian Deconvolution using Simple Reversible Jump Moves

We propose a Markov chain Monte Carlo-based deconvolution method designed to estimate the number of peaks in spectral data, along with the optimal parameters of each radial basis function. Assuming cases where the number of peaks is unknown, and a sweep simulation on all candidate models is computationally unrealistic, the proposed method efficiently searches over the probable candidates via trans-dimensional moves assisted by annealing effects from replica exchange Monte Carlo moves. Through simulation using synthetic data, the proposed method demonstrates its advantages over conventional sweep simulations, particularly in model selection problems. Application to a set of olivine reflectance spectral data with varying forsterite and fayalite mixture ratios reproduced results obtained from previous mineralogical research, indicating that our method is applicable to deconvolution on real data sets.

stat.ME↗

A Phase Prediction Method for Pattern Formation in Time-Dependent Ginzburg-Landau Dynamics for Kinetic Ising Model without a priori Assumptions on Domain Patterns

We propose a phase prediction method for the pattern formation in the uniaxial two-dimensional kinetic Ising model with the dipole-dipole interactions under the time-dependent Ginzburg-Landau dynamics. Taking the effects of the material thickness into account by assuming the uniformness along the magnetization axis, the model corresponds to thin magnetic materials with long-range repulsive interactions. We propose a new theoretical basis to understand the effects of the material parameters on the formation of the magnetic domain patterns in terms of the equation of balance governing the balance between the linear- and nonlinear forces in the equilibrium state. Based on this theoretical basis, we propose a new method to predict the phase in the equilibrium state reached after the time-evolution under the dynamics with a given set of parameters, by approximating the third-order term using the restricted phase-space approximation [R. Anzaki, K. Fukushima, Y. Hidaka, and T. Oka, Ann. Phys. 353, 107 (2015)] for the $ϕ^4$-models. Although the proposed method does not have the perfect concordance with the actual numerical results, it has no arbitrary parameters and functions to tune the prediction. In other words, it is a method with no a priori assumptions on domain patterns.

cond-mat.stat-mech↗

Data-Dependence of Plateau Phenomenon in Learning with Neural Network --- Statistical Mechanical Analysis

The plateau phenomenon, wherein the loss value stops decreasing during the process of learning, has been reported by various researchers. The phenomenon is actively inspected in the 1990s and found to be due to the fundamental hierarchical structure of neural network models. Then the phenomenon has been thought as inevitable. However, the phenomenon seldom occurs in the context of recent deep learning. There is a gap between theory and reality. In this paper, using statistical mechanical formulation, we clarified the relationship between the plateau phenomenon and the statistical property of the data learned. It is shown that the data whose covariance has small and dispersed eigenvalues tend to make the plateau phenomenon inconspicuous.

stat.ML↗

Statistical mechanical evaluation of spread spectrum watermarking model with image restoration

In cases in which an original image is blind, a decoding method where both the image and the messages can be estimated simultaneously is desirable. We propose a spread spectrum watermarking model with image restoration based on Bayes estimation. We therefore need to assume some prior probabilities. The probability for estimating the messages is given by the uniform distribution, and the ones for the image are given by the infinite range model and 2D Ising model. Any attacks from unauthorized users can be represented by channel models. We can obtain the estimated messages and image by maximizing the posterior probability. We analyzed the performance of the proposed method by the replica method in the case of the infinite range model. We first calculated the theoretical values of the bit error rate from obtained saddle point equations and then verified them by computer simulations. For this purpose, we assumed that the image is binary and is generated from a given prior probability. We also assume that attacks can be represented by the Gaussian channel. The computer simulation retults agreed with the theoretical values. In the case of prior probability given by the 2D Ising model, in which each pixel is statically connected with four-neighbors, we evaluated the decoding performance by computer simulations, since the replica theory could not be applied. Results using the 2D Ising model showed that the proposed method with image restoration is as effective as the infinite range model for decoding messages. We compared the performances in a case in which the image was blind and one in which it was informed. The difference between these cases was small as long as the embedding and attack rates were small. This demonstrates that the proposed method with simultaneous estimation is effective as a watermarking decoder.

cond-mat.stat-mech↗

Bayesian Spectral Deconvolution Based on Poisson Distribution: Bayesian Measurement and Virtual Measurement Analytics (VMA)

In this paper, we propose a new method of Bayesian measurement for spectral deconvolution, which regresses spectral data into the sum of unimodal basis function such as Gaussian or Lorentzian functions. Bayesian measurement is a framework for considering not only the target physical model but also the measurement model as a probabilistic model, and enables us to estimate the parameter of a physical model with its confidence interval through a Bayesian posterior distribution given a measurement data set. The measurement with Poisson noise is one of the most effective system to apply our proposed method. Since the measurement time is strongly related to the signal-to-noise ratio for the Poisson noise model, Bayesian measurement with Poisson noise model enables us to clarify the relationship between the measurement time and the limit of estimation. In this study, we establish the probabilistic model with Poisson noise for spectral deconvolution. Bayesian measurement enables us to perform virtual and computer simulation for a certain measurement through the established probabilistic model. This property is called "Virtual Measurement Analytics(VMA)" in this paper. We also show that the relationship between the measurement time and the limit of estimation can be extracted by using the proposed method in a simulation of synthetic data and real data for XPS measurement of MoS$_2$.

eess.SP↗

Bayesian Hamiltonian Selection in X-ray Photoelectron Spectroscopy

Core-level X-ray photoelectron spectroscopy (XPS) is a useful measurement technique for investigating the electronic states of a strongly correlated electron system. Usually, to extract physical information of a target object from a core-level XPS spectrum, we need to set an effective Hamiltonian by physical consideration so as to express complicated electron-to-electron interactions in the transition of core-level XPS, and manually tune the physical parameters of the effective Hamiltonian so as to represent the XPS spectrum. Then, we can extract physical information from the tuned parameters. In this paper, we propose an automated method for analyzing core-level XPS spectra based on the Bayesian model selection framework, which selects the effective Hamiltonian and estimates its parameters automatically. The Bayesian model selection, which often has a large computational cost, was carried out by the exchange Monte Carlo sampling method. By applying our proposed method to the 3$d$ core-level XPS spectra of Ce and La compounds, we confirmed that our proposed method selected an effective Hamiltonian and estimated its parameters appropriately; these results were consistent with conventional knowledge obtained from physical studies. Moreover, using our proposed method, we can also evaluate the uncertainty of its estimation values and clarify why the effective Hamiltonian was selected. Such information is difficult to obtain by the conventional analysis method.

cond-mat.str-el↗

Statistical mechanical analysis of sparse linear regression as a variable selection problem

An algorithmic limit of compressed sensing or related variable-selection problems is analytically evaluated when a design matrix is given by an overcomplete random matrix. The replica method from statistical mechanics is employed to derive the result. The analysis is conducted through evaluation of the entropy, an exponential rate of the number of combinations of variables giving a specific value of fit error to given data which is assumed to be generated from a linear process using the design matrix. This yields the typical achievable limit of the fit error when solving a representative $\ell_0$ problem and includes the presence of unfavourable phase transitions preventing local search algorithms from reaching the minimum-error configuration. The associated phase diagrams are presented. A noteworthy outcome of the phase diagrams is that there exists a wide parameter region where any phase transition is absent from the high temperature to the lowest temperature at which the minimum-error configuration or the ground state is reached. This implies that certain local search algorithms can find the ground state with moderate computational costs in that region. Another noteworthy result is the presence of the random first-order transition in the strong noise case. The theoretical evaluation of the entropy is confirmed by extensive numerical methods using the exchange Monte Carlo and the multi-histogram methods. Another numerical test based on a metaheuristic optimisation algorithm called simulated annealing is conducted, which well supports the theoretical predictions on the local search algorithms. In the successful region with no phase transition, the computational cost of the simulated annealing to reach the ground state is estimated as the third order polynomial of the model dimensionality.

cond-mat.dis-nn↗

Concept Formation and Dynamics of Repeated Inference in Deep Generative Models

Deep generative models are reported to be useful in broad applications including image generation. Repeated inference between data space and latent space in these models can denoise cluttered images and improve the quality of inferred results. However, previous studies only qualitatively evaluated image outputs in data space, and the mechanism behind the inference has not been investigated. The purpose of the current study is to numerically analyze changes in activity patterns of neurons in the latent space of a deep generative model called a "variational auto-encoder" (VAE). What kinds of inference dynamics the VAE demonstrates when noise is added to the input data are identified. The VAE embeds a dataset with clear cluster structures in the latent space and the center of each cluster of multiple correlated data points (memories) is referred as the concept. Our study demonstrated that transient dynamics of inference first approaches a concept, and then moves close to a memory. Moreover, the VAE revealed that the inference dynamics approaches a more abstract concept to the extent that the uncertainty of input data increases due to noise. It was demonstrated that by increasing the number of the latent variables, the trend of the inference dynamics to approach a concept can be enhanced, and the generalization ability of the VAE can be improved.

stat.ML↗

Exhaustive search for sparse variable selection in linear regression

We propose a K-sparse exhaustive search (ES-K) method and a K-sparse approximate exhaustive search method (AES-K) for selecting variables in linear regression. With these methods, K-sparse combinations of variables are tested exhaustively assuming that the optimal combination of explanatory variables is K-sparse. By collecting the results of exhaustively computing ES-K, various approximate methods for selecting sparse variables can be summarized as density of states. With this density of states, we can compare different methods for selecting sparse variables such as relaxation and sampling. For large problems where the combinatorial explosion of explanatory variables is crucial, the AES-K method enables density of states to be effectively reconstructed by using the replica-exchange Monte Carlo method and the multiple histogram method. Applying the ES-K and AES-K methods to type Ia supernova data, we confirmed the conventional understanding in astronomy when an appropriate K is given beforehand. However, we found the difficulty to determine K from the data. Using virtual measurement and analysis, we argue that this is caused by data shortage.

stat.ML↗

Statistical Mechanics of Node-perturbation Learning with Noisy Baseline

Node-perturbation learning is a type of statistical gradient descent algorithm that can be applied to problems where the objective function is not explicitly formulated, including reinforcement learning. It estimates the gradient of an objective function by using the change in the object function in response to the perturbation. The value of the objective function for an unperturbed output is called a baseline. Cho et al. proposed node-perturbation learning with a noisy baseline. In this paper, we report on building the statistical mechanics of Cho's model and on deriving coupled differential equations of order parameters that depict learning dynamics. We also show how to derive the generalization error by solving the differential equations of order parameters. On the basis of the results, we show that Cho's results are also apply in general cases and show some general performances of Cho's model.

stat.ML↗