SearcharxivSearch

arXiv subjects

Yoshiyuki Kabashima

Publications and source records attributed to Yoshiyuki Kabashima.

At least 19 recordsLinked to original sources

Dynamical Regimes of Discrete Diffusion Models

Diffusion models generate high-dimensional data such as images by learning a process that gradually removes noise from corrupted data. Recent studies have shown that the backward dynamics of diffusion models exhibit two characteristic transitions: the speciation transition, at which generated samples begin to capture the global structure of the training data, and the collapse transition, at which the generation dynamics starts committing to individual training samples. While these transitions have been theoretically analyzed for continuous data, the same theoretical criteria have not been applied for discrete diffusion models, which are diffusion models for discrete data. In this work, we propose a simple effective model for discrete diffusion models trained on two-class Ising variable data with a general mixture ratio and analyze its backward dynamics using methods from statistical mechanics. We show that, as in the previous study on continuous data, the speciation transition can be determined through a second-order phase transition analysis using high-temperature expansion, while the collapse transition corresponds to a condensation transition described by the Random Energy Model. An analytical expression for the speciation time is obtained, and we show that its scaling becomes consistent with that of the continuous case when the noise increases with time as in practical diffusion models. These theoretical predictions are confirmed by numerical simulations and experiments with trained discrete diffusion models on real datasets. These results suggest that the original theoretical framework for continuous data remain valid for discrete data, and may provide a useful starting point for the statistical-mechanics analysis of discrete generative diffusion in more realistic settings.

cond-mat.stat-mech

Analysis of the Hopfield Model Incorporating the Effects of Unlearning

We analyze a variant of the Hopfield model that incorporates an unlearning mechanism based on spin correlations in the high-temperature regime. In the large system limit where extensively many patterns are stored, we employ the replica method under the replica symmetric ansatz to characterize the model analytically. Our analysis provides a systematic and self-consistent framework that yields order-parameter equations and stability conditions at finite temperatures over a wide range of parameter settings. The resulting theory accurately captures the behavior of the signal-to-noise ratio, the memory capacity, and the criteria for selecting optimal hyperparameters, in agreement with the qualitative findings of Nokura (1996 \textit{J. Phys. A: Math. Gen.} \textbf{29} 3871). Moreover, the theoretical predictions show good agreement with numerical simulations, supporting the conclusion that unlearning enhances memory capacity by suppressing spurious memories.

cond-mat.dis-nn

DMFT analysis of Hopfield network with plasticity

We study a fully connected Hopfield-type associative memory network with online activity-dependent synaptic plasticity, where neural states and synaptic couplings coevolve during retrieval. Using the generating-functional formalism, we derive a dynamical mean-field theory (DMFT) in the large-system limit with extensively many stored random patterns, and show that the many-body dynamics reduces to an effective single-site stochastic process with colored Gaussian crosstalk noise and delayed feedback terms. Numerical solutions of the DMFT equations agree well with direct simulations. We find that moderate plasticity enlarges the basin of attraction and increases the maximum retrievable memory load by generating a positive delayed feedback that stabilizes retrieval against crosstalk noise. However, excessively strong plasticity causes the network to imprint the imperfect initial cue itself, leading to spurious attractors and degraded retrieval performance. Consequently, an optimal plasticity strength emerges from the trade-off between memory stabilization and premature cue imprinting. These results extend the DMFT description of associative memory to networks with coevolving neural and synaptic dynamics.

cond-mat.dis-nn

Learning Linear Regression with Low-Rank Tasks in-Context

In-context learning (ICL) is a key building block of modern large language models, yet its theoretical mechanisms remain poorly understood. It is particularly mysterious how ICL operates in real-world applications where tasks have a common structure. In this work, we address this problem by analyzing a linear attention model trained on low-rank regression tasks. Within this setting, we precisely characterize the distribution of predictions and the generalization error in the high-dimensional limit. Moreover, we find that statistical fluctuations in finite pre-training data induce an implicit regularization. Finally, we identify a sharp phase transition of the generalization error governed by task structure. These results provide a framework for understanding how transformers learn to learn the task structure.

cond-mat.dis-nn

Testing the Role of Diagonal Interactions in High-Order Hopfield Models via Dynamical Mean-Field Theory

High-order extensions of the Hopfield model are known to exhibit dramatically enhanced storage capacity at equilibrium, while their dynamical retrieval properties remain less well understood. In our previous work, we carried out a dynamical mean-field theory (DMFT) analysis of the Krotov--Hopfield-type dense associative memory and found that the transition between successful and failed retrieval is accompanied by pronounced slow dynamics. As a consequence, the effective basin of attraction observed in numerical simulations extends well beyond that predicted by equilibrium statistical mechanics. A natural hypothesis is that this discrepancy originates from diagonal (self-interaction) contributions in the Krotov--Hopfield model, which generate a large number of lower-order interaction terms and may induce glassy relaxation near the retrieval boundary. To test this hypothesis, we analyze an alternative high-order associative memory model, namely the Abbott--Arian-type $p$-body Hopfield model, in which such diagonal contributions are absent by construction. Using dynamical mean-field theory, we derive an effective single-site process together with closed macroscopic equations governing the retrieval dynamics. Our analysis reveals that both slow dynamics and a substantial enlargement of the apparent basin of attraction persist even in this model. These results indicate that the dynamical slowdown near the retrieval boundary cannot be attributed primarily to diagonal self-interaction effects, but instead originates from intrinsic properties of high-order interactions.

cond-mat.stat-mech

Dynamical mean field approach to associative memory model with non-monotonic transfer functions

The Hopfield associative memory model stores random patterns in synaptic couplings according to Hebb's rule and retrieves them through gradient descent on an energy function. This conventional setting, where neurons are assumed to have monotonic transfer functions, has been central to understanding associative memory. Morita (1993, Neural Netw. 6 115), however, showed that introducing non-monotonic transfer functions can dramatically enhance retrieval performance. While this phenomenon has been qualitatively examined, a full quantitative theory remains elusive due to the difficulty of analysis in the absence of an underlying energy function. In this work, we apply dynamical mean-field theory to the discrete-time synchronous retrieval dynamics of the non-monotonic model, which succeeds in accurately characterizing its macroscopic dynamical properties. We also derive conditions for retrieval states, and clarify their relation to previous studies. Our results provide new insights into the non-equilibrium retrieval dynamics of associative memory models.

cond-mat.dis-nn

Storage capacity of perceptron with variable selection

A central challenge in machine learning is to distinguish genuine structure from chance correlations in high-dimensional data. In this work, we address this issue for the perceptron, a foundational model of neural computation. Specifically, we investigate the relationship between the pattern load $α$ and the variable selection ratio $ρ$ for which a simple perceptron can perfectly classify $P = αN$ random patterns by optimally selecting $M = ρN$ variables out of $N$ variables. While the Cover--Gardner theory establishes that a random subset of $ρN$ dimensions can separate $αN$ random patterns if and only if $α< 2ρ$, we demonstrate that optimal variable selection can surpass this bound by developing a method, based on the replica method from statistical mechanics, for enumerating the combinations of variables that enable perfect pattern classification. This not only provides a quantitative criterion for distinguishing true structure in the data from spurious regularities, but also yields the storage capacity of associative memory models with sparse asymmetric couplings.

cs.IT

Dynamical Properties of Dense Associative Memory

Dense associative memory, a fundamental instance of modern Hopfield networks, can store a large number of memory patterns as equilibrium states of recurrent networks. While the stationary-state storage capacity has been investigated, its dynamical properties have not yet been discussed. In this paper, we analyze the dynamics using an exact approach based on generating functional analysis. We show results on convergence properties of memory retrieval, such as the convergence time and the size of the attraction basins. Our analysis enables a quantitative evaluation of the convergence time and the storage capacity of dense associative memory, which is useful for model design. Unlike the traditional Hopfield model, the retrieval of a pattern does not act as additional noise to itself, suggesting that the structure of modern networks makes recall more robust. Furthermore, the methodology addressed here can be applied to other energy-based models, and thus has the potential to contribute to the design of future architectures.

cond-mat.dis-nn

Quantum effects in rotationally invariant spin glass models

This study investigates the quantum effects in transverse-field Ising spin glass models with rotationally invariant random interactions. The primary aim is to evaluate the validity of a quasi-static approximation that captures the imaginary-time dependence of the order parameters beyond the conventional static approximation. Using the replica method combined with the Suzuki--Trotter decomposition, we established a stability condition for the replica symmetric solution, which is analogous to the de Almeida--Thouless criterion. Numerical analysis of the Sherrington--Kirkpatrick model estimates a value of the critical transverse field, $Γ_\mathrm{c}$, which agrees with previous Monte Carlo-based estimations. For the Hopfield model, it provides an estimate of $Γ_\mathrm{c}$, which has not been previously evaluated. For the random orthogonal model, our analysis suggests that quantum effects alter the random first-order transition scenario in the low-temperature limit. This study supports a quasi-static treatment for analyzing quantum spin glasses and may offer useful insights into the analysis of quantum optimization algorithms.

cond-mat.dis-nn

Improving Decoupled Posterior Sampling for Inverse Problems using Data Consistency Constraint

Diffusion models have shown strong performances in solving inverse problems through posterior sampling while they suffer from errors during earlier steps. To mitigate this issue, several Decoupled Posterior Sampling methods have been recently proposed. However, the reverse process in these methods ignores measurement information, leading to errors that impede effective optimization in subsequent steps. To solve this problem, we propose Guided Decoupled Posterior Sampling (GDPS) by integrating a data consistency constraint in the reverse process. The constraint performs a smoother transition within the optimization process, facilitating a more effective convergence toward the target distribution. Furthermore, we extend our method to latent diffusion models and Tweedie's formula, demonstrating its scalability. We evaluate GDPS on the FFHQ and ImageNet datasets across various linear and nonlinear tasks under both standard and challenging conditions. Experimental results demonstrate that GDPS achieves state-of-the-art performance, improving accuracy over existing methods.

cs.LG

Alpha helices are more evolutionarily robust to environmental perturbations than beta sheets: Bayesian learning and statistical mechanics for protein evolution

How typical elements that shape organisms, such as protein secondary structures, have evolved, or how evolutionarily susceptible/resistant they are to environmental changes, are significant issues in evolutionary biology, structural biology, and biophysics. According to Darwinian evolution, natural selection and genetic mutations are the primary drivers of biological evolution. However, the concept of ``robustness of the phenotype to environmental perturbations across successive generations," which seems crucial from the perspective of natural selection, has not been formalized or analyzed. In this study, through Bayesian learning and statistical mechanics we formalize the stability of the free energy in the space of amino acid sequences that can design particular protein structure against perturbations of the chemical potential of water surrounding a protein as such robustness. This evolutionary stability is defined as a decreasing function of a quantity analogous to the susceptibility in the statistical mechanics of magnetic bodies specific to the amino acid sequence of a protein. Consequently, in a two-dimensional square lattice protein model composed of 36 residues, we found that as we increase the stability of the free energy against perturbations in environmental conditions, the structural space shows a steep step-like reduction. Furthermore, lattice protein structures with higher stability against perturbations in environmental conditions tend to have a higher proportion of $α$-helices and a lower proportion of $β$-sheets. This result is qualitatively confirmed by comparing the histograms of the percentage of secondary structures of evolutionarily robust proteins and randomly selected proteins through an empirical validation using a protein database.

physics.bio-ph

Exact Replica Symmetric solution for transverse field Hopfield model under finite Trotter size

We analyze the quantum Hopfield model in which an extensive number of patterns are embedded in the presence of a uniform transverse field. This analysis employs the replica method under the replica symmetric ansatz on the Suzuki-Trotter representation of the model, while keeping the number of Trotter slices $M$ finite. The statistical properties of the quantum Hopfield model in imaginary time are reduced to an effective $M$-spin long-range classical Ising model, which can be extensively studied using a dedicated Monte Carlo algorithm. This approach contrasts with the commonly applied static approximation, which ignores the imaginary time dependency of the order parameters, but allows $M \to \infty$ to be taken analytically. During the analysis, we introduce an exact but fundamentally weaker static relation, referred to as the quasi-static relation. We present the phase diagram of the model with respect to the transverse field strength and the number of embedded patterns, indicating a small but quantitative difference from previous results obtained using the static approximation.

cond-mat.dis-nn

Forecasting long-time dynamics in quantum many-body systems by dynamic mode decomposition

Reliable numerical computation of quantum dynamics is a fundamental challenge when the long-ranged quantum entanglement plays essential roles as in the cases governed by quantum criticality in strongly correlated systems. Here we apply a method that utilizes reliable short-time data of physical quantities to accurately forecast long-time behavior of the strongly entangled systems. We straightforwardly employ the simple dynamic mode decomposition (DMD), which is commonly used in fluid dynamics. Despite the simplicity of the method, the effectiveness and applicability of the DMD in quantum many-body systems such as the Ising model in the transverse field at the critical point are demonstrated, even when the time evolution at long time exhibits complicated features such as a volume-law entanglement entropy and consequential power-law decays of correlations characteristic of systems with long-ranged quantum entanglements unlike fluid dynamics. The present method, though simple, enables accurate forecasts amazingly at time as long as nearly an order of magnitude longer than that of the short-time training data. Effects of noise on the accuracy of the forecast are also investigated, because they are important especially when dealing with the experimental data. We find that a few percentages of noise do not affect the prediction accuracy destructively.

quant-ph

Diffusion Model Based Posterior Sampling for Noisy Linear Inverse Problems

With the rapid development of diffusion models and flow-based generative models, there has been a surge of interests in solving noisy linear inverse problems, e.g., super-resolution, deblurring, denoising, colorization, etc, with generative models. However, while remarkable reconstruction performances have been achieved, their inference time is typically too slow since most of them rely on the seminal diffusion posterior sampling (DPS) framework and thus to approximate the intractable likelihood score, time-consuming gradient calculation through back-propagation is needed. To address this issue, this paper provides a fast and effective solution by proposing a simple closed-form approximation to the likelihood score. For both diffusion and flow-based models, extensive experiments are conducted on various noisy linear inverse problems such as noisy super-resolution, denoising, deblurring, and colorization. In all these tasks, our method (namely DMPS) demonstrates highly competitive or even better reconstruction performances while being significantly faster than all the baseline methods.

cs.LG

Detection of diffusion anisotropy from an individual short particle trajectory

In parallel with advances in microscale imaging techniques, the fields of biology and materials science have focused on precisely extracting particle properties based on their diffusion behavior. Although the majority of real-world particles exhibit anisotropy, their behavior has been studied less than that of isotropic particles. In this study, we introduce a new method for estimating the diffusion coefficients of individual anisotropic particles using short-trajectory data on the basis of a maximum likelihood framework. Traditional estimation techniques often use mean-squared displacement (MSD) values or other statistical measures that inherently remove angular information. Instead, we treated the angle as a latent variable and used belief propagation to estimate it while maximizing the likelihood using the expectation-maximization algorithm. Compared to conventional methods, this approach facilitates better estimation of shorter trajectories and faster rotations, as confirmed by numerical simulations and experimental data involving bacteria and quantum rods. Additionally, we performed an analytical investigation of the limits of detectability of anisotropy and provided guidelines for the experimental design. In addition to serving as a powerful tool for analyzing complex systems, the proposed method will pave the way for applying maximum likelihood methods to more complex diffusion phenomena.

cond-mat.mes-hall

QCM-SGM+: Improved Quantized Compressed Sensing With Score-Based Generative Models

In practical compressed sensing (CS), the obtained measurements typically necessitate quantization to a limited number of bits prior to transmission or storage. This nonlinear quantization process poses significant recovery challenges, particularly with extreme coarse quantization such as 1-bit. Recently, an efficient algorithm called QCS-SGM was proposed for quantized CS (QCS) which utilizes score-based generative models (SGM) as an implicit prior. Due to the adeptness of SGM in capturing the intricate structures of natural signals, QCS-SGM substantially outperforms previous QCS methods. However, QCS-SGM is constrained to (approximately) row-orthogonal sensing matrices as the computation of the likelihood score becomes intractable otherwise. To address this limitation, we introduce an advanced variant of QCS-SGM, termed QCS-SGM+, capable of handling general matrices effectively. The key idea is a Bayesian inference perspective on the likelihood score computation, wherein expectation propagation is employed for its approximate computation. Extensive experiments are conducted, demonstrating the substantial superiority of QCS-SGM+ over QCS-SGM for general sensing matrices beyond mere row-orthogonality.

eess.SP

Generating gradients in the energy landscape using rectified linear type cost functions for efficiently solving 0/1 matrix factorization in Simulated Annealing

The 0/1 matrix factorization defines matrix products using logical AND and OR as product-sum operators, revealing the factors influencing various decision processes. Instances and their characteristics are arranged in rows and columns. Formulating matrix factorization as an energy minimization problem and exploring it with Simulated Annealing (SA) theoretically enables finding a minimum solution in sufficient time. However, searching for the optimal solution in practical time becomes problematic when the energy landscape has many plateaus with flat slopes. In this work, we propose a method to facilitate the solution process by applying a gradient to the energy landscape, using a rectified linear type cost function readily available in modern annealing machines. We also propose a method to quickly obtain a solution by updating the cost function's gradient during the search process. Numerical experiments were conducted, confirming the method's effectiveness with both noise-free artificial and real data.

cs.LG

Compressed Sensing Radar Detectors based on Weighted LASSO

The compressed sensing (CS) model can represent the signal recovery process of a large number of radar systems. The detection problem of such radar systems has been studied in many pieces of literature through the technology of debiased least absolute shrinkage and selection operator (LASSO). While naive LASSO treats all the entries equally, there are many applications in which prior information varies depending on each entry. Weighted LASSO, in which the weights of the regularization terms are tuned depending on the entry-dependent prior, is proven to be more effective with the prior information by many researchers. In the present paper, existing results obtained by methods of statistical mechanics are utilized to derive the debiased weighted LASSO estimator for randomly constructed row-orthogonal measurement matrices. Based on this estimator, we construct a detector, termed the debiased weighted LASSO detector (DWLD), for CS radar systems and prove its advantages. The threshold of this detector can be calculated by false alarm rate, which yields better detection performance than the naive weighted LASSO detector (NWLD) under the Neyman-Pearson principle. The improvement of the detection performance brought by tuning weights is demonstrated by numerical experiments. With the same false alarm rate, the detection probability of DWLD is obviously higher than those of NWLD and the debiased (non-weighted) LASSO detector (DLD).

eess.SP