SearcharxivSearch

arXiv subjects

Shuhei Kashiwamura

Publications and source records attributed to Shuhei Kashiwamura.

8 recordsLinked to original sources

Improving Generalization by Permutation Routing Across Model Copies

We introduce a use of the \(M\)-cover (or \(M\)-layer) transform for machine learning. The method replicates a model \(M\) times, but instead of coupling the copies through parameter averaging or an explicit attractive force, as in replicated SGD or Elastic SGD, it rewires the contexts in which local learning messages are computed. Each local loss is evaluated on a routed model whose parameters are drawn from different copies according to permutations sampled from a structured mixing kernel \(Q\). Training then uses the original local update rule, while the resulting learning messages are redistributed across the copies through these routed computational paths. Thus \(Q\) defines a topology for message transport and controls the long-loop structure of the lifted factor graph. We formulate this construction for perceptrons, committee machines, and multilayer perceptrons, showing that the same principle applies from discrete models to differentiable neural networks. The resulting framework provides a mechanism for improving generalization through structured message sharing rather than replica collapse or parameter-space coupling.

cs.LG

Uncertainty-Aware Sparse Identification of Dynamical Systems via Bayesian Model Averaging

In many problems of data-driven modeling for dynamical systems, the governing equations are not known a priori and must be selected phenomenologically from a large set of candidate interactions and basis functions. In such situations, point estimates alone can be misleading, because multiple model components may explain the observed data comparably well, especially when the data are limited or the dynamics exhibit poor identifiability. Quantifying the uncertainty associated with model selection is therefore essential for constructing reliable dynamical models from data. In this work, we develop a Bayesian sparse identification framework for dynamical systems with coupled components, aimed at inferring both interaction structure and functional form together with principled uncertainty quantification. The proposed method combines sparse modeling with Bayesian model averaging, yielding posterior inclusion probabilities that quantify the credibility of each candidate interaction and basis component. Through numerical experiments on oscillator networks, we show that the framework accurately recovers sparse interaction structures with quantified uncertainty, including higher-order harmonic components, phase-lag effects, and multi-body interactions. We also demonstrate that, even in a phenomenological setting where the true governing equations are not contained in the assumed model class, the method can identify effective functional components with quantified uncertainty. These results highlight the importance of Bayesian uncertainty quantification in data-driven discovery of dynamical models.

stat.AP

High-Dimensional Learning Dynamics of Quantized Models with Straight-Through Estimator

Quantized neural network training optimizes a discrete, non-differentiable objective. The straight-through estimator (STE) enables backpropagation through surrogate gradients and is widely used. While previous studies have primarily focused on the properties of surrogate gradients and their convergence, the influence of quantization hyperparameters, such as bit width and quantization range, on learning dynamics remains largely unexplored. We theoretically show that in the high-dimensional limit, STE dynamics converge to a deterministic ordinary differential equation. This reveals that STE training exhibits a plateau followed by a sharp drop in generalization error, with plateau length depending on the quantization range. A fixed-point analysis quantifies the asymptotic deviation from the unquantized linear model. We also extend analytical techniques for stochastic gradient descent to nonlinear transformations of weights and inputs.

stat.ML

Bayesian estimation of coupling strength and heterogeneity in a coupled oscillator model from macroscopic quantities

Various macroscopic oscillations, such as the heartbeat and the flashing of fireflies, are created by synchronizing oscillatory units (oscillators). To elucidate the mechanism of synchronization, several coupled oscillator models have been devised and extensively analyzed. Although parameter estimation of these models has also been actively investigated, most of the proposed methods are based on the data from individual oscillators, not from macroscopic quantities. In the present study, we propose a Bayesian framework to estimate the model parameters of coupled oscillator models, using the time series data of the Kuramoto order parameter as the only given data. We adopt the exchange Monte Carlo method for the efficient estimation of the posterior distribution and marginal likelihood. Numerical experiments are performed to confirm the validity of our method and examine the dependence of the estimation error on the observational noise and system size.

physics.data-an

Mesoscopic Bayesian Inference by Solvable Models

The rapid advancement of data science and artificial intelligence has affected physics in numerous ways, including the application of Bayesian inference, setting the stage for a revolution in research methodology. Our group has proposed Bayesian measurement, a framework that applies Bayesian inference to measurement science with broad applicability across various natural sciences. This framework enables the determination of posterior probability distributions of system parameters, model selection, and the integration of multiple measurement datasets. However, applying Bayesian measurement to real data analysis requires a more sophisticated approach than traditional statistical methods like Akaike information criterion (AIC) and Bayesian information criterion (BIC), which are designed for an infinite number of measurements $N$. Therefore, in this paper, we propose an analytical theory that explicitly addresses the case where $N$ is finite in the linear regression model. We introduce $O(1)$ mesoscopic variables for $N$ observation noises. Using this mesoscopic theory, we analyze the three core principles of Bayesian measurement: parameter estimation, model selection, and measurement integration. Furthermore, by introducing these mesoscopic variables, we demonstrate that the difference in free energies, critical for both model selection and measurement integration, can be analytically reduced by two mesoscopic variables of $N$ observation noises. This provides a deeper qualitative understanding of model selection and measurement integration and further provides deeper insights into actual measurements for nonlinear models. Our framework presents a novel approach to understanding Bayesian measurement results.

physics.data-an

Effect of Weight Quantization on Learning Models by Typical Case Analysis

This paper examines the quantization methods used in large-scale data analysis models and their hyperparameter choices. The recent surge in data analysis scale has significantly increased computational resource requirements. To address this, quantizing model weights has become a prevalent practice in data analysis applications such as deep learning. Quantization is particularly vital for deploying large models on devices with limited computational resources. However, the selection of quantization hyperparameters, like the number of bits and value range for weight quantization, remains an underexplored area. In this study, we employ the typical case analysis from statistical physics, specifically the replica method, to explore the impact of hyperparameters on the quantization of simple learning models. Our analysis yields three key findings: (i) an unstable hyperparameter phase, known as replica symmetry breaking, occurs with a small number of bits and a large quantization width; (ii) there is an optimal quantization width that minimizes error; and (iii) quantization delays the onset of overparameterization, helping to mitigate overfitting as indicated by the double descent phenomenon. We also discover that non-uniform quantization can enhance stability. Additionally, we develop an approximate message-passing algorithm to validate our theoretical results.

stat.ML

Magnetic field generation by charge exchange in a supernova remnant in the early universe

We present new generation mechanisms of magnetic fields in supernova remnant shocks propagating to partially ionized plasmas in the early universe. Upstream plasmas are dissipated at the collisionless shock, but hydrogen atoms are not dissipated because they do not interact with electromagnetic fields. After the hydrogen atoms are ionized in the shock downstream region, they become cold proton beams that induce the electron return current. The injection of the beam protons can be interpreted as an external force acting on the downstream proton plasma. We show that the effective external force and the electron return current can generate magnetic fields without any seed magnetic fields. The magnetic field strength is estimated to be $B\sim 10^{-14}-10^{-11}~{\rm G}$, where the characteristic lengthscale is the mean free path of charge exchange, $\sim 10^{15}~{\rm cm}$. Since protons are marginally magnetized by the generated magnetic field in the downstream region, the magnetic field could be amplified to larger values and stretched to larger scales by turbulent dynamo and expansion.

astro-ph.HE

Bayesian Spectral Deconvolution of X-Ray Absorption Near Edge Structure Discriminating High- and Low-Energy Domains

In this paper, we propose a Bayesian spectral deconvolution considering the properties of peaks in different energy domains. Bayesian spectral deconvolution regresses spectral data into the sum of multiple basis functions. Conventional methods use a model that treats all peaks equally. However, in X-ray absorption near edge structure (XANES) spectra, the properties of the peaks differ depending on the energy domain, and the specific energy domain of XANES is essential in condensed matter physics. We propose a model that discriminates between the low- and high-energy domains. We also propose a prior distribution that reflects the physical properties. We compare the conventional and proposed models in terms of computational efficiency, estimation accuracy, and model evidence. We demonstrate that our method effectively estimates the number of transition components in the important energy domain, on which the material scientists focus for mapping the electronic transition analysis by first-principles simulation.

stat.ME