SearcharxivSearch

arXiv subjects

Kyunghoo Mun

Publications and source records attributed to Kyunghoo Mun.

5 recordsLinked to original sources

Phase transitions for the noisy transformer model in arbitrary dimension

We study the McKean--Vlasov free energy on the unit sphere associated with the unnormalized self-attention (USA) model for noisy transformer dynamics. We prove a sharp global-minimizer dichotomy in every dimension $d\ge2$. There is a unique $β_*^{(d)}>0$ such that \begin{equation*} \frac{I_{d/2+1}(β_*^{(d)})}{I_{d/2}(β_*^{(d)})}=\frac1d, \end{equation*} where $I_ν$ is the modified Bessel function of the first kind. For $0<β\le β_*^{(d)}$, the uniform density remains the unique global minimizer up to the linear-stability threshold \begin{equation*} K_\#^{(d)}(β)=\frac{β^{d/2}}{2^{d/2}Γ(d/2)I_{d/2}(β)}, \end{equation*} and the phase transition is continuous. For $β>β_*^{(d)}$, the uniform density is not globally minimizing at $K_\#^{(d)}(β)$, so the critical coupling satisfies $K_c<K_\#^{(d)}(β)$ and the transition is discontinuous. This result generalizes the authors' recent $d=2$ work arXiv:2604.16288 to arbitrary dimension. The proof uses the sharp Beckner--Onofri/logarithmic Hardy-Littlewood-Sobolev (HLS) inequality on the sphere, together with a Funk--Hecke/Bessel coefficient computation and a degree-two quartic obstruction.

math.AP

Phase transitions in Doi-Onsager, Noisy Transformer, and other multimodal models

We study phase transitions for repulsive-attractive mean-field free energies on the circle. For a $\frac{1}{n+1}$-periodic interaction whose Fourier coefficients satisfy a certain decay condition, we prove that the critical coupling strength $K_c$ coincides with the linear stability threshold $K_\#$ of the uniform distribution and that the phase transition is continuous, in the sense that the uniform distribution is the unique global minimizer at criticality. The proof is based on a sharp coercivity estimate for the free energy obtained from the constrained Lebedev--Milin inequality. We apply this result to three motivating models for which the exact value of the phase transition and its (dis)continuity in terms of the model parameters was not fully known. For the two-dimensional Doi--Onsager model $W(θ)=-|\sin(2πθ)|$, we prove that the phase transition is continuous at $K_c=K_\#=3π/4$. For the noisy transformer model $W_β(θ)=(e^{β\cos(2πθ)}-1)/β$, we identify the sharp threshold $β_*$ such that $K_c(β) = K_\#(β)$ and the phase transition is continuous for $β\leq β_*$, while $K_c(β) β_*$. We also obtain the corresponding sharp dichotomy for the noisy Hegselmann--Krause model $W_{R}(θ) = (R-2π|θ|)_{+}^2$ .

math.AP

Phase transitions and linear stability for the mean-field Kuramoto-Daido model

We consider the mean-field noisy Kuramoto-Daido model, which is a McKean-Vlasov equation on the circle with bimodal interaction $W(θ)=\cosθ+m\cos2θ$ for $m\ge 0$ and interaction strength $K$, generalizing the celebrated noisy Kuramoto model corresponding to $m=0$. Our first contribution is to characterize the phase transition threshold $K_{c}$ by comparing it to the linear stability threshold $K_\# = \min (1, m^{-1})$ of the uniform distribution. When $m \leq 1/2,$ $K_{c}=1$, coinciding with that of the Kuramoto model. On the other hand, for $m \geq 2$, we show $K_c= m^{-1}$. We also classify the regimes in which the phase transition is continuous or discontinuous. Our second contribution is to analyze the linear stability of a global minimizer $q$ (the ``ordered phase'') of the mean-field free energy in the supercritical regime $K>1$. This stationary solution of the Kuramoto-Daido equation is unique up to translation invariance and distinct from the uniform distribution (the ``disordered phase''). Our approach extends the Dirichlet form method of Bertini et al. from the unimodal to bimodal setting. In particular, for $m \leq 1.590 \times 10^{-4}$ and $K>1$, we show an explicit lower bound on the spectral gap of the linearized McKean-Vlasov operator at $q$. To our knowledge, this is the first rigorous stability analysis for this class of models with bimodal interactions.

math.AP

Dynamical phase transition for the homogeneous multi-component Curie-Weiss-Potts model

In this paper, we study the homogeneous multi-component Curie-Weiss-Potts model with $q \geq 3$ spins. The model is defined on the complete graph $K_{Nm}$, whose vertex set is equally partitioned into $m$ components of size $N$. For a configuration $σ: \{1, \cdots, Nm\} \to \{1, \cdots, q\},$ the Gibbs measure is defined by $$ μ_{N,β}(σ) =\frac{1}{Z_{N,β}} \exp\Big(\fracβ{N} \sum_{v,w=1}^{Nm}\mathcal{J}(v,w)\, \mathbb{1}_{\{σ(v)=σ(w)\}}\Big), $$ where $Z_{N, β}$ is a normalizing constant, and $β>0$ is the inverse temperature parameter. The interaction coefficients are $ \mathcal{J}(v, w) = \frac{J}{1 + (m-1) λ}$, for $v, w$ in the same component, and $\mathcal{J}(v, w) = \frac{J λ}{1 + (m-1)λ}$ for $v, w$ in the different components, where $λ\in (0, 1)$ is the relative strength of inter-component interaction to intra-component interaction, and $J>0$ is the effective interaction strength. We identify a dynamical phase transition at the critical inverse temperature $β_{\operatorname{cr}} = β_{s}(q)/J$, where $β_{s}(q)$ is maximal inverse temperature guaranteeing a unique critical point of the free energy in the Curie-Weiss-Potts model arXiv:1204.4503. By extending the aggregate path method arXiv:1312.6728 to our multi-component setting, we prove $O(N \log N)$ mixing time in the high-temperature regime $β<β_{s}(q)/J.$ In the low-temperature regime $β> β_{s}(q)/J,$ we further show exponential mixing time by a metastability. This is the first result for the dynamical phase transition in the multi-component Potts model.

math.PR

Mini-Batch Optimization of Contrastive Loss

Contrastive learning has gained significant attention as a method for self-supervised learning. The contrastive loss function ensures that embeddings of positive sample pairs (e.g., different samples from the same class or different views of the same object) are similar, while embeddings of negative pairs are dissimilar. Practical constraints such as large memory requirements make it challenging to consider all possible positive and negative pairs, leading to the use of mini-batch optimization. In this paper, we investigate the theoretical aspects of mini-batch optimization in contrastive learning. We show that mini-batch optimization is equivalent to full-batch optimization if and only if all $\binom{N}{B}$ mini-batches are selected, while sub-optimality may arise when examining only a subset. We then demonstrate that utilizing high-loss mini-batches can speed up SGD convergence and propose a spectral clustering-based approach for identifying these high-loss mini-batches. Our experimental results validate our theoretical findings and demonstrate that our proposed algorithm outperforms vanilla SGD in practically relevant settings, providing a better understanding of mini-batch optimization in contrastive learning.

cs.LG