Searcharxiv⌕ Search

arXiv subjects

Eiji Kawasaki

Publications and source records attributed to Eiji Kawasaki.

15 recordsLinked to original sources

Bio-inspired Learning and Decision-Making with Probabilistic In-Memory Computing Hardware: Part 2

This report extends our previous work (Part 1), which introduced an energy-based model for learning and decision-making under uncertainty. The model leverages stochastic Langevin dynamics to continuously evolve approximate probability distributions over neuron states and model weights. However, as noted in Part 1 and confirmed through GPU-based implementations, large-scale probabilistic energy-based models of this nature face significant scalability challenges due to excessive execution latency. This latency stems from a fundamental mismatch: massively parallel models with low arithmetic intensity (such as energy-based models) are being executed on processor architectures like GPUs that rely on high-bandwidth memory (HBM) interfaces. The HBM imposes brutally sequential execution constraints on inherently parallelizable models, creating the false impression that such models are unscalable. In reality, it is the GPU architecture itself, with its dependence on HBM interfaces, that is not a scalable processor architecture for this class of AI model. In this report, we demonstrate using a detailed transaction-level model (TLM) of a probabilistic analogue in-memory computing (AIMC) processor that the same energy-based model can execute well over 1000x faster than data-center-grade hardware by eliminating the HBM interface and performing computation directly within on-chip memory.

cs.AR↗

Bio-inspired Learning and Decision-Making with Probabilistic In-Memory Computing Hardware: Part 1

Learning and decision-making in animals are often modeled as Bayesian processes, where sensory evidence is integrated with prior beliefs to guide behavior in the face of uncertainty. But what are the inherent neural dynamics that give rise to this ability, and how could they be replicated in computing systems? This abstract discusses a biologically grounded framework in which noisy neural and synaptic dynamics perform inference and learning via stochastic sampling from an internal energy function, capturing uncertainty over latent states and model parameters through neural and synaptic variability, respectively. This enables approaches such as predictive coding networks to account for epistemic uncertainty via Markov chain Monte Carlo sampling. Drawing a parallel between intrinsic noise in biological systems and electrical noise in emerging probabilistic analogue memory technologies, we highlight how analogue in-memory computing hardware naturally emerges as the solution for massively scalable and energy-efficient probabilistic inference.

cs.AI↗

Contrastive Regularization of Machine Learning Potentials

Machine learning interatomic potentials are trained to predict energies and forces but built to be sampled: their purpose is to drive molecular simulations whose observables average over the equilibrium distribution the potential defines. They exemplify a broader AI problem -- learned regressors deployed as generators -- where pointwise accuracy does not guarantee a correct distribution. We show that potentials trained by standard Mean Squared Error (MSE) minimization on Density Functional Theory (DFT) data can reach chemical accuracy on held-out data, yet still fail as samplers: their trajectories drift into spurious low-energy minima and return thermodynamic observables that depart sharply from the reference. To correct this, we introduce Contrastive Regularized MSE (CRMSE), a post-training step that augments the MSE with a contrastive term derived from the Kullback--Leibler divergence between the potential's implicit Boltzmann distribution and the target. The network serves as its own energy-based model: persistent Langevin chains expose the configurations it drifts into and raise their energy, adding no new ab initio data. On the ethanol and aspirin molecules of the MD17 dataset, CRMSE confines the sampler to the physical basin and recovers the energy distribution, interatomic-distance distributions, and dihedral free-energy profiles to near-quantitative agreement with DFT, while preserving force accuracy and keeping energy errors within chemical accuracy; it remains effective when the training set is sharply reduced. That MSE training fails this way on MD17 -- one of the most widely used benchmarks -- while a minimal contrastive correction repairs it suggests that reliable sampling depends less on data volume than on training the model against the distribution it produces: distribution-level training is not a refinement of regression accuracy, but a distinct requirement.

physics.chem-ph↗

Uncertainty in AI-driven Monte Carlo simulations

In the study of complex systems, evaluating physical observables often requires sampling representative configurations via Monte Carlo techniques. These methods rely on repeated evaluations of the system's energy and force fields, which can become computationally expensive. To accelerate these simulations, deep learning models are increasingly employed as surrogate functions to approximate the energy landscape or force fields. However, such models introduce epistemic uncertainty in their predictions, which may propagate through the sampling process and affect the simulation's macroscopic behavior. In our work, we present the Penalty Ensemble Method (PEM) to quantify epistemic uncertainty and mitigate its impact on Monte Carlo sampling. Our approach introduces an uncertainty-aware modification of the Metropolis acceptance rule, which increases the rejection probability in regions of high uncertainty, thereby enhancing the reliability of the simulation outcomes.

cond-mat.dis-nn↗

Thermodynamic properties of chemically disordered compounds via AI-driven estimation of partition function with the PULSE method

In this article, we present an improved version of the PULSE method (Partition function Unsupervised Learning Sampling and Evaluation) for estimating the thermodynamic properties of chemically disordered compounds. The aim is to reduce the computational cost of Monte Carlo approaches for this type of material and to demonstrate that this generative tool can estimate thermodynamic properties by sampling and estimating the partition function of the system. To validate this innovative approach, we use the 2D Ising model as a benchmark. We demonstrate that our method accurately reproduces average properties with high precision and efficiency compared to traditional Monte Carlo sampling methods. Our results highlight the efficiency and adaptability of the PULSE method, making it a valuable tool for studying materials for which conventional methods are too inefficient to compute properties affected by chemical disorder at low cost.

cond-mat.stat-mech↗

Multivariate Bayesian Last Layer for Regression with Uncertainty Quantification and Decomposition

We present new Bayesian Last Layer neural network models in the setting of multivariate regression under heteroscedastic noise, and propose EM algorithms for parameter learning. Bayesian modeling of a neural network's final layer has the attractive property of uncertainty quantification with a single forward pass. The proposed framework is capable of disentangling the aleatoric and epistemic uncertainty, and can be used to enhance a canonically trained deep neural network with uncertainty-aware capabilities.

stat.ML↗

Classifier Weighted Mixture models

This paper proposes an extension of standard mixture stochastic models, by replacing the constant mixture weights with functional weights defined using a classifier. Classifier Weighted Mixtures enable straightforward density evaluation, explicit sampling, and enhanced expressivity in variational estimation problems, without increasing the number of components nor the complexity of the mixture components.

stat.ML↗

Data Subsampling for Bayesian Neural Networks

Markov Chain Monte Carlo (MCMC) algorithms do not scale well for large datasets leading to difficulties in Neural Network posterior sampling. In this paper, we propose Penalty Bayesian Neural Networks - PBNNs, as a new algorithm that allows the evaluation of the likelihood using subsampled batch data (mini-batches) in a Bayesian inference context towards addressing scalability. PBNN avoids the biases inherent in other naive subsampling techniques by incorporating a penalty term as part of a generalization of the Metropolis Hastings algorithm. We show that it is straightforward to integrate PBNN with existing MCMC frameworks, as the variance of the loss function merely reduces the acceptance probability. By comparing with alternative sampling strategies on both synthetic data and the MNIST dataset, we demonstrate that PBNN achieves good predictive performance even for small mini-batch sizes of data. We show that PBNN provides a novel approach for calibrating the predictive distribution by varying the mini-batch size, significantly reducing predictive overconfidence.

stat.ML↗

Targeting the partition function of chemically disordered materials with a generative approach based on inverse variational autoencoders

Computing atomic-scale properties of chemically disordered materials requires an efficient exploration of their vast configuration space. Traditional approaches such as Monte Carlo or Special Quasirandom Structures either entail sampling an excessive amount of configurations or do not ensure that the configuration space has been properly covered. In this work, we propose a novel approach where generative machine learning is used to yield a representative set of configurations for accurate property evaluation and provide accurate estimations of atomic-scale properties with minimal computational cost. Our method employs a specific type of variational autoencoder with inverse roles for the encoder and decoder, enabling the application of an unsupervised active learning scheme that does not require any initial training database. The model iteratively generates configuration batches, whose properties are computed with conventional atomic-scale methods. These results are then fed back into the model to estimate the partition function, repeating the process until convergence. We illustrate our approach by computing point-defect formation energies and concentrations in (U, Pu)O2 mixed-oxide fuels. In addition, the ML model provides valuable insights into the physical factors influencing the target property. Our method is generally applicable to explore other properties, such as atomic-scale diffusion coefficients, in ideally or non-ideally disordered materials like high-entropy alloys.

cond-mat.mtrl-sci↗

Generative vs. Discriminative modeling under the lens of uncertainty quantification

Learning a parametric model from a given dataset indeed enables to capture intrinsic dependencies between random variables via a parametric conditional probability distribution and in turn predict the value of a label variable given observed variables. In this paper, we undertake a comparative analysis of generative and discriminative approaches which differ in their construction and the structure of the underlying inference problem. Our objective is to compare the ability of both approaches to leverage information from various sources in an epistemic uncertainty aware inference via the posterior predictive distribution. We assess the role of a prior distribution, explicit in the generative case and implicit in the discriminative case, leading to a discussion about discriminative models suffering from imbalanced dataset. We next examine the double role played by the observed variables in the generative case, and discuss the compatibility of both approaches with semi-supervised learning. We also provide with practical insights and we examine how the modeling choice impacts the sampling from the posterior predictive distribution. With regard to this, we propose a general sampling scheme enabling supervised learning for both approaches, as well as semi-supervised learning when compatible with the considered modeling approach. Throughout this paper, we illustrate our arguments and conclusions using the example of affine regression, and validate our comparative analysis through classification simulations using neural network based models.

stat.ML↗

Scaling-up Memristor Monte Carlo with magnetic domain-wall physics

By exploiting the intrinsic random nature of nanoscale devices, Memristor Monte Carlo (MMC) is a promising enabler of edge learning systems. However, due to multiple algorithmic and device-level limitations, existing demonstrations have been restricted to very small neural network models and datasets. We discuss these limitations, and describe how they can be overcome, by mapping the stochastic gradient Langevin dynamics (SGLD) algorithm onto the physics of magnetic domain-wall Memristors to scale-up MMC models by five orders of magnitude. We propose the push-pull pulse programming method that realises SGLD in-physics, and use it to train a domain-wall based ResNet18 on the CIFAR-10 dataset. On this task, we observe no performance degradation relative to a floating point model down to an update precision of between 6 and 7-bits, indicating we have made a step towards a large-scale edge learning system leveraging noisy analogue devices.

cs.ET↗

Semi-supervised generative approach to point-defect formation in chemically disordered compounds: application to uranium-plutonium mixed oxides

Machine-learning methods are nowadays of common use in the field of material science. For example, they can aid in optimizing the physicochemical properties of new materials, or help in the characterization of highly complex chemical compounds. An especially challenging problem arises in the modeling of chemically disordered solid solutions, for which some properties depend on the distribution of chemical species in the crystal lattice. This is the case of defect properties of uranium-plutonium mixed oxides nuclear fuels. The number of possible configurations is so high that the problem becomes intractable if treated with direct sampling. We thus propose a machine learning approach, based on generative modeling, to optimize the exploration of this large configuration space. A probabilistic, semi-supervised approach using Mixture Density Network is applied to estimate the concentration of thermal defects in (U, Pu)O2. We show that this method, based on the prediction of the density of states of formation energy of a defect, is computationally much more cost-efficient compared to other approaches available in the literature.

cond-mat.dis-nn↗

Discretely Indexed Flows

In this paper we propose Discretely Indexed flows (DIF) as a new tool for solving variational estimation problems. Roughly speaking, DIF are built as an extension of Normalizing Flows (NF), in which the deterministic transport becomes stochastic, and more precisely discretely indexed. Due to the discrete nature of the underlying additional latent variable, DIF inherit the good computational behavior of NF: they benefit from both a tractable density as well as a straightforward sampling scheme, and can thus be used for the dual problems of Variational Inference (VI) and of Variational density estimation (VDE). On the other hand, DIF can also be understood as an extension of mixture density models, in which the constant mixture weights are replaced by flexible functions. As a consequence, DIF are better suited for capturing distributions with discontinuities, sharp edges and fine details, which is a main advantage of this construction. Finally we propose a methodology for constructiong DIF in practice, and see that DIF can be sequentially cascaded, and cascaded with NF.

stat.ML↗

Finite Temperature Phases of Two Dimensional Spin-Orbit Coupled Bosons

We determine the finite temperature phase diagram of two dimensional bosons with two hyperfine (pseudo-spin) states coupled via Rashba-Dresselhaus spin-orbit interaction using classical field Monte Carlo calculations. For anisotropic spin-orbit coupling, we find a transition to a Berenzinskii-Kosterlitz-Thouless superfluid phase with quasi-long range order. We show that the spin-order of the quasi-condensate is driven by the anisotropy of interparticle interaction, favoring either a homogeneous plane wave state or stripe phase with broken translational symmetry. Both phases show characteristic behavior in the algebraically decaying spin density correlation function. For fully isotropic interparticle interaction, our calculations indicate a fractionalized quasi-condensate where the mean-field degeneracy of plane wave and stripe phase remains robust against critical fluctuations. In the case of fully isotropic spin-orbit coupling, the circular degeneracy of the single particle ground state destroys the algebraic ordered phase in the thermodynamic limit, but a cross-over remains for finite size systems.

cond-mat.quant-gas↗

Dynamical depinning of a Tonks Girardeau gas

We study the dynamical depinning following a sudden turn off of an optical lattice for a gas of impenetrable bosons in a tight atomic waveguide. We use a Bose-Fermi mapping to infer the exact quantum dynamical evolution. At long times, in the thermodynamic limit, we observe the approach to a non-equilibrium steady state, characterized by the absence of quasi-long-range order and a reduced visibility in the momentum distribution. Similar features are found in a finite-size system at times corresponding to half the revival time, where we find that the system approaches a quasi-steady state with a power-law behaviour.

cond-mat.quant-gas↗