Searcharxiv⌕ Search

arXiv subjects

Koji Hukushima

Publications and source records attributed to Koji Hukushima.

At least 19 recordsLinked to original sources

Statistical-mechanics of a three-body Hopfield model with finite connectivity

A three-body Hopfield model defined on sparse random graphs with finite mean connectivity is studied using a replica-symmetric (RS) analysis. By extending the functional replica framework for conventional two-body finite-connectivity Hopfield models, self-consistent equations for the local-field distribution are derived and solved numerically by population dynamics. In contrast to two-body sparse Hopfield models, three-body interactions induce a discontinuous retrieval transition, coexistence of paramagnetic and retrieval solutions, and a distinct spinodal structure. Starting the RS population dynamics from an uninformative finite-amplitude field distribution leads to a spin-glass-like fixed point at low temperatures rather than to the retrieval state. This trapping already occurs for a single embedded pattern, where conventional cross-talk noise among stored patterns is absent, indicating that it originates from the combination of sparse connectivity and three-body interactions. Within the RS description, increasing the number of embedded patterns drives a crossing between the retrieval and glassy free-energy branches, making the glassy branch thermodynamically stable at high load. Finite-size Monte Carlo simulations using population annealing support the RS description of the retrieval branch and reproduce the trapping behavior under cooling, while deviations from the RS spin-glass branch point to replica-symmetry breaking in the low-temperature glassy regime. The characteristic storage load scales linearly with the mean connectivity, reflecting the sparse number of couplings. These results clarify how many-body discontinuity and sparse graph disorder simultaneously influence the accessibility and capacity of associative memory.

cond-mat.dis-nn↗

Phase transition in large language models and the criticality of natural languages

Generation of text and speech in natural languages can be modeled as a stochastic process. This idea dates back to the seminal work of Markov and, later, to that of Shannon and also underlies the recent development of large language models (LLMs). The stochastic processes corresponding to natural languages should be distinct from those that generate nonlinguistic sequences. One of the features that discriminate linguistic and nonlinguistic sequences is power-law behavior, which is universally observed across different languages. In statistical physics, such behavior suggests that natural languages are critical: They lie near a phase transition point in a parametrized space of stochastic processes. However, testing this conjecture is not straightforward. A phase transition, even if it exists, cannot be directly observed in real-world natural languages because they do not have any controllable parameters. Here, we use LLMs as controllable effective models of natural languages. Through statistical analyses of texts generated by LLMs, we find that, when a parameter analogous to physical temperature is varied, LLMs undergo a phase transition. The transition separates a low-temperature phase with complex repetitive structures in generated texts from a high-temperature phase in which LLMs generate incomprehensible texts. At the critical point between these phases, generated texts display the power-law behavior similar to that of natural languages and most closely resemble natural languages as measured by a standard metric in natural language processing. These findings strongly suggest that natural languages are indeed critical.

cond-mat.dis-nn↗

Tensor-Network Population Annealing

We propose a hybrid sampling method, tensor-network population annealing (TNPA), which combines tensor-network (TN) initialization with population annealing (PA). We apply this method to the two-dimensional Edwards-Anderson Ising spin glass. The approach is motivated by the limitations of existing methods: TN-based samplers can become numerically unstable in frustrated spin systems at low temperatures, whereas conventional PA requires a long annealing schedule when started from the high-temperature limit. In TNPA, TN contractions are used only within a reliable temperature range to generate initial configurations that are close to the equilibrium distribution. The subsequent low-temperature equilibration is then carried out by PA. To stabilize the initialization process, we introduce a diagnostic based on the effective sample size that adaptively selects the initialization temperature. The proposed framework provides a practical and physically motivated route to low-temperature sampling by combining the complementary strengths of TN and PA.

cond-mat.stat-mech↗

Phase Transitions in a Modified Ising Spin Glass Model: A Tensor-Network-based Sampling Approach

Phase transitions in a modified Nishimori model, including the model considered by Kitatani, on a two-dimensional square lattice are investigated using a tensor-network-based sampling scheme. In this model, generating bond configurations is computationally demanding because of the correlated random interactions. The employed sampling method enables hierarchical and independent sampling of both bonds and spins. This approach allows high-precision calculations for system sizes up to $L=256$. The results provide clear numerical evidence that the spin-glass and ferromagnetic transitions are separated on the Nishimori line, supporting the existence of an intermediate Mattis-like spin-glass phase. This finding is consistent with the reentrant transition numerically observed in the two-dimensional Edwards-Anderson (EA) model. Furthermore, critical exponents estimated via finite-size-scaling analysis indicate that the universality class of the transitions differs from that of the standard independent and identically distributed EA model.

cond-mat.dis-nn↗

Rethinking the Relationship between the Power Law and Hierarchical Structures

Statistical analysis of corpora provides an approach to quantitatively investigate natural languages. This approach has revealed that several power laws consistently emerge across different corpora and languages, suggesting universal mechanisms underlying languages. In particular, the power-law decay of correlations has been interpreted as evidence of underlying hierarchical structures in syntax, semantics, and discourse. This perspective has also been extended beyond corpora produced by human adults, including child speech, birdsong, and chimpanzee action sequences. However, the argument supporting this interpretation has not been empirically tested in natural languages. To address this gap, the present study examines the validity of the argument for syntactic structures. Specifically, we test whether the statistical properties of parse trees align with the assumptions in the argument. Using English and Japanese corpora, we analyze the mutual information, deviations from probabilistic context-free grammars (PCFGs), and other properties in natural language parse trees, as well as in the PCFG that approximates these parse trees. Our results indicate that the assumptions do not hold for syntactic structures and that it is difficult to apply the proposed argument not only to sentences by human adults but also to other domains, highlighting the need to reconsider the relationship between the power law and hierarchical structures.

cs.CL↗

Structure-resolved free energy estimation of the 38-atom Lennard Jones cluster via population annealing

We systematically investigate the thermodynamic landscape of the 38-atom Lennard--Jones cluster LJ$_{38}$ using Population Annealing (PA), a method suited for systems with challenging double-funnel energy landscapes. By employing an adaptive temperature schedule, we demonstrate that thermodynamic observables, such as internal energy and heat capacity, converge robustly when the population size is sufficiently large. To gain deeper insights into the competing basins, we introduce an integrated framework that combines PA reweighting factors with structure-resolved analysis. Using quenched configurations characterized by potential energy and Steinhardt's bond-orientational order parameters, we identify three structural basins, FCC-like, icosahedral, and liquid-like, via dimensionality reduction and clustering. This framework enables the direct computation of structure-resolved free energy differences from population fractions, providing a quantitative mapping of the thermodynamic competition between the funnels. The resulting structural crossovers are consistent with the heat-capacity peak, demonstrating PA as a promising and scalable framework for structure-resolved thermodynamics in complex molecular systems.

physics.chem-ph↗

Phase transitions in time complexity of Brownian circuits

Brownian circuits perform computations using stochastic transitions driven by thermal fluctuations. While the energetic costs of such fluctuation-driven computation have been extensively studied within stochastic thermodynamics, much less is known about its computational complexity, in particular, how computation time scales with circuit size. In this work, the computation time for explicitly designed Brownian circuits is numerically investigated via the first-passage time to a completed state. For arithmetic circuits such as adders, varying the forward transition rate induces a sharp change in the scaling behavior of the mean computation time with circuit size, from linear to exponential. This change can be interpreted as an easy-hard transition in computational time complexity. The transition suggests that, for meaningful computational tasks, achieving efficient polynomial-time computation generally requires a finite forward bias corresponding to a nonzero energy input. As a counterexample, we show that arbitrary logical operations can be reduced to an effective one-dimensional stochastic process in which the zero-bias limit lies within the computationally efficient (easy) regime. However, achieving such a one-dimensional normal form unavoidably leads to an exponential increase in circuit size. These results reveal a fundamental trade-off between computation time, circuit size, and energy input in Brownian circuits and demonstrate that phase transitions in time complexity provide a natural framework for characterizing the cost of fluctuation-driven computation.

cond-mat.stat-mech↗

High-dimensional Asymptotics of VAEs: Threshold of Posterior Collapse and Dataset-Size Dependence of Rate-Distortion Curve

In variational autoencoders (VAEs), the variational posterior often collapses to the prior, known as posterior collapse, which leads to poor representation learning quality. An adjustable hyperparameter beta has been introduced in VAEs to address this issue. This study sharply evaluates the conditions under which the posterior collapse occurs with respect to beta and dataset size by analyzing a minimal VAE in a high-dimensional limit. Additionally, this setting enables the evaluation of the rate-distortion curve of the VAE. Our results show that, unlike typical regularization parameters, VAEs face "inevitable posterior collapse" beyond a certain beta threshold, regardless of dataset size. Moreover, the dataset-size dependence of the derived rate-distortion curve suggests that relatively large datasets are required to achieve a rate-distortion curve with high rates. These findings robustly explain generalization behavior observed in various real datasets with highly non-linear VAEs.

stat.ML↗

Tracking Cluster Continuity and Dynamics in Time-Series Data: Application to Chromatin Polymer Simulations

This study presents an enhanced method for analyzing cluster dynamics, with a particular focus on tracking clusters' continuity over time using time-series data from molecular dynamics (MD) simulation. The proposed method was applied to spatio-temporal cluster data obtained from a non-equilibrium MD simulation of a chromatin polymer model. In this model, clusters are formed on the polymer by binding molecules that stochastically and temporarily bind to the polymer segments at finite rates. Our analysis successfully tracked the dynamics of clusters, including merging and splitting events, and revealed that clusters exhibit a percolation transition in both spatial and temporal domains. This suggests that clusters in the chromatin polymer model can persist even under finite rates of attractive interactions, demonstrating that the method can capture complex cluster dynamics over time.

cond-mat.soft↗

Damage spreading and coupling in spin glasses and hard spheres

We study the connection between damage spreading, a phenomenon long discussed in the physics literature, and the coupling of Markov chains, a technique used to bound the mixing time. We discuss in parallel the Edwards-Anderson spin glass model and the hard-disk system, focusing on how coupling provides insights into the region of fast coupling within the paramagnetic and liquid phases. We also work out the connection between path coupling and damage spreading. Numerically, the scaling analysis of the mean coupling time determines a critical point between fast and slow couplings. The exact relationship between fast coupling and disordered phases has not been established rigorously, but we suggest that it will ultimately enhance our understanding of phase behavior in disordered systems.

cond-mat.dis-nn↗

Ratio Divergence Learning Using Target Energy in Restricted Boltzmann Machines: Beyond Kullback--Leibler Divergence Learning

We propose ratio divergence (RD) learning for discrete energy-based models, a method that utilizes both training data and a tractable target energy function. We apply RD learning to restricted Boltzmann machines (RBMs), which are a minimal model that satisfies the universal approximation theorem for discrete distributions. RD learning combines the strength of both forward and reverse Kullback-Leibler divergence (KLD) learning, effectively addressing the "notorious" issues of underfitting with the forward KLD and mode-collapse with the reverse KLD. Since the summation of forward and reverse KLD seems to be sufficient to combine the strength of both approaches, we include this learning method as a direct baseline in numerical experiments to evaluate its effectiveness. Numerical experiments demonstrate that RD learning significantly outperforms other learning methods in terms of energy function fitting, mode-covering, and learning stability across various discrete energy-based models. Moreover, the performance gaps between RD learning and the other learning methods become more pronounced as the dimensions of target models increase.

stat.ML↗

Statistical Mechanics of Min-Max Problems

Min-max optimization problems, also known as saddle point problems, have attracted significant attention due to their applications in various fields, such as fair beamforming, generative adversarial networks (GANs), and adversarial learning. However, understanding the properties of these min-max problems has remained a substantial challenge. This study introduces a statistical mechanical formalism for analyzing the equilibrium values of min-max problems in the high-dimensional limit, while appropriately addressing the order of operations for min and max. As a first step, we apply this formalism to bilinear min-max games and simple GANs, deriving the relationship between the amount of training data and generalization error and indicating the optimal ratio of fake to real data for effective learning. This formalism provides a groundwork for a deeper theoretical analysis of the equilibrium properties in various machine learning methods based on min-max problems and encourages the development of new algorithms and architectures.

cs.LG↗

Statistical properties of probabilistic context-sensitive grammars

Probabilistic context-free grammars (PCFGs), which are commonly used to generate trees randomly, have been well analyzed theoretically, leading to applications in various domains. Despite their utility, the distributions that the grammar can express are limited to those in which the distribution of a subtree depends only on its root and not on its context. This limitation presents a challenge for modeling various real-world phenomena, such as natural languages. To overcome this limitation, a probabilistic context-sensitive grammar (PCSG) is introduced, where the distribution of a subtree depends on its context. Numerical analysis of a PCSG reveals that the distribution of a symbol does not constitute a qualitative difference from that in the context-free case, but mutual information does. Furthermore, a novel metric introduced to directly quantify the breaking of this limitation detects a distinct difference between PCFGs and PCSGs. This metric, applicable to an arbitrary distribution of a tree, allows for further investigation and characterization of various tree structures that PCFGs cannot express.

cond-mat.dis-nn↗

Effect of constraint relaxation on dynamic critical phenomena in minimum vertex cover problem

The effects of constraint relaxation on dynamic critical phenomena in the Minimum Vertex Cover (MVC) problem on Erdős-Rényi random graphs are investigated using Markov chain Monte Carlo simulations. Following our previous work that revealed the reduction of the critical temperature by constraint relaxation based on the penalty function method, this study focuses on investigating the critical properties of the relaxation time along its phase boundary. It is found that the dynamical correlation function of MVC with respect to the problem size and the constraint strength follows a universal scaling function. The analysis shows that the relaxation time decreases as the constraints are relaxed. This decrease is more pronounced for the critical amplitude than for the critical exponent, and this result is interpreted in terms of the system's microscopic energy barriers due to the constraint relaxation.

cond-mat.stat-mech↗

Effect of Constraint Relaxation on the Minimum Vertex Cover Problem in Random Graphs

A statistical-mechanical study of the effect of constraint relaxation on the minimum vertex cover problem in Erdős-Rényi random graphs is presented. Using a penalty-method formulation for constraint relaxation, typical properties of solutions, including infeasible solutions that violate the constraints, are analyzed by means of the replica method and cavity method. The problem involves a competition between reducing the number of vertices to be covered and satisfying the edge constraints. The analysis under the replica-symmetric (RS) ansatz clarifies that the competition leads to degeneracies in the vertex and edge states, which determine the quantitative properties of the system, such as the cover and penalty ratios. A precise analysis of these effects improves the accuracy of RS approximation for the minimum cover ratio in the replica symmetry breaking (RSB) region. Furthermore, the analysis based on the RS cavity method indicates that the RS/RSB boundary of the ground states with respect to the mean degree of the graphs is expanded, and the critical temperature is lowered by constraint relaxation.

cond-mat.stat-mech↗

Adaptive Flip Graph Algorithm for Matrix Multiplication

This study proposes the "adaptive flip graph algorithm", which combines adaptive searches with the flip graph algorithm for finding fast and efficient methods for matrix multiplication. The adaptive flip graph algorithm addresses the inherent limitations of exploration and inefficient search encountered in the original flip graph algorithm, particularly when dealing with large matrix multiplication. For the limitation of exploration, the proposed algorithm adaptively transitions over the flip graph, introducing a flexibility that does not strictly reduce the number of multiplications. Concerning the issue of inefficient search in large instances, the proposed algorithm adaptively constraints the search range instead of relying on a completely random search, facilitating more effective exploration. Numerical experimental results demonstrate the effectiveness of the adaptive flip graph algorithm, showing a reduction in the number of multiplications for a $4\times 5$ matrix multiplied by a $5\times 5$ matrix from $76$ to $73$, and that from $95$ to $94$ for a $5 \times 5$ matrix multiplied by another $5\times 5$ matrix. These results are obtained in characteristic two.

cs.SC↗

Proof of avoidability of the quantum first-order transition in transverse magnetization in quantum annealing of finite-dimensional spin glasses

It is rigorously shown that an appropriate quantum annealing for any finite-dimensional spin system has no quantum first-order transition in transverse magnetization. This result can be applied to finite-dimensional spin-glass systems, where the ground state search problem is known to be hard to solve. Consequently, it is strongly suggested that the quantum first-order transition in transverse magnetization is not fatal to the difficulty of combinatorial optimization problems in quantum annealing.

quant-ph↗

Learning Dynamics in Linear VAE: Posterior Collapse Threshold, Superfluous Latent Space Pitfalls, and Speedup with KL Annealing

Variational autoencoders (VAEs) face a notorious problem wherein the variational posterior often aligns closely with the prior, a phenomenon known as posterior collapse, which hinders the quality of representation learning. To mitigate this problem, an adjustable hyperparameter $β$ and a strategy for annealing this parameter, called KL annealing, are proposed. This study presents a theoretical analysis of the learning dynamics in a minimal VAE. It is rigorously proved that the dynamics converge to a deterministic process within the limit of large input dimensions, thereby enabling a detailed dynamical analysis of the generalization error. Furthermore, the analysis shows that the VAE initially learns entangled representations and gradually acquires disentangled representations. A fixed-point analysis of the deterministic process reveals that when $β$ exceeds a certain threshold, posterior collapse becomes inevitable regardless of the learning period. Additionally, the superfluous latent variables for the data-generative factors lead to overfitting of the background noise; this adversely affects both generalization and learning convergence. The analysis further unveiled that appropriately tuned KL annealing can accelerate convergence.

stat.ML↗