SearcharxivSearch

arXiv subjects

Kaiyuan Cui

Publications and source records attributed to Kaiyuan Cui.

7 recordsLinked to original sources

Log-Sobolev Inequality for Wolff Dynamics and Application to the Condensation of Eigen Microstate in the 1D Ising Model

The Wolff dynamics is a non-local Markov chain widely used for simulating the Ising model due to its effectiveness in reducing critical slowing down compared to the Glauber dynamics. Despite extensive algorithmic and numerical studies, a rigorous probabilistic understanding remains limited. In this paper, we take a first step toward addressing this gap. For the one-dimensional (1D) Ising model, we first derive the transition probabilities of the Wolff dynamics and show that, at the critical point, it converges to the two fully aligned configurations and subsequently oscillates between them. This behavior is absent in the Glauber dynamics. Second, we establish a log-Sobolev inequality with an explicit constant for the Wolff dynamics in the entire subcritical regime and derive quantitative bounds on its ergodic averages. As a by-product, at infinite temperature, the obtained constant coincides with the classical log-Sobolev constant of the random walk on the hypercube. Finally, we apply these results to analyze the spectrum of the sample covariance matrix generated by the Wolff dynamics, which was used by Chen et al. to study condensation of eigen microstate. We prove that the spectral behavior agrees with their simulations in the 1D Ising model, thereby providing theoretical support for their findings.

math.PR

Toward Universal and Transferable Jailbreak Attacks on Vision-Language Models

Vision-language models (VLMs) extend large language models (LLMs) with vision encoders, enabling text generation conditioned on both images and text. However, this multimodal integration expands the attack surface by exposing the model to image-based jailbreaks crafted to induce harmful responses. Existing gradient-based jailbreak methods transfer poorly, as adversarial patterns overfit to a single white-box surrogate and fail to generalise to black-box models. In this work, we propose Universal and transferable jailbreak (UltraBreak), a framework that constrains adversarial patterns through transformations and regularisation in the vision space, while relaxing textual targets through semantic-based objectives. By defining its loss in the textual embedding space of the target LLM, UltraBreak discovers universal adversarial patterns that generalise across diverse jailbreak objectives. This combination of vision-level regularisation and semantically guided textual supervision mitigates surrogate overfitting and enables strong transferability across both models and attack targets. Extensive experiments show that UltraBreak consistently outperforms prior jailbreak methods. Further analysis reveals why earlier approaches fail to transfer, highlighting that smoothing the loss landscape via semantic objectives is crucial for enabling universal and transferable jailbreaks. The code is publicly available in our \href{https://github.com/kaiyuanCui/UltraBreak}{GitHub repository}.

cs.LG

The log-Sobolev inequality and correlation functions for the renormalization of 1D Ising model

The renormalization group (RG) method is an important tool for studying critical phenomena. In this paper, we employ stochastic analysis techniques to investigate the stochastic partial differential equation (SPDE) derived by regularizing and continuousizing the discrete stochastic equation, which is a variant of stochastic quantization equation of the one dimensional (1D) Ising model. Firstly, we give the regularity estimates for the solution to SPDE. Secondly, we prove the Clark-Ocone-Haussmann formula and derive the log-Sobolev inequality up to the terminal time $T$, as well as obtain a priori form of the renormalization relation. Finally, we verify the correctness of the renormalization procedure based on the partition function, and prove that the two point correlation functions of SPDE on lattices converge to the two point correlation functions of the 1D Ising model at the stable fixed point of the RG transformation as $T\rightarrow +\infty$.

math.PR

Importance Weighted Score Matching for Diffusion Samplers with Enhanced Mode Coverage

Training neural samplers directly from unnormalized densities without access to target distribution samples presents a significant challenge. A critical desideratum in these settings is achieving comprehensive mode coverage, ensuring the sampler captures the full diversity of the target distribution. However, prevailing methods often circumvent the lack of target data by optimizing reverse KL-based objectives. Such objectives inherently exhibit mode-seeking behavior, potentially leading to incomplete representation of the underlying distribution. While alternative approaches strive for better mode coverage, they typically rely on implicit mechanisms like heuristics or iterative refinement. In this work, we propose a principled approach for training diffusion-based samplers by directly targeting an objective analogous to the forward KL divergence, which is conceptually known to encourage mode coverage. We introduce \textit{Importance Weighted Score Matching}, a method that optimizes this desired mode-covering objective by re-weighting the score matching loss using tractable importance sampling estimates, thereby overcoming the absence of target distribution data. We also provide theoretical analysis of the bias and variance for our proposed Monte Carlo estimator and the practical loss function used in our method. Experiments on increasingly complex multi-modal distributions, including 2D Gaussian Mixture Models with up to 120 modes and challenging particle systems with inherent symmetries -- demonstrate that our approach consistently outperforms existing neural samplers across all distributional distance metrics, achieving state-of-the-art results on all benchmarks.

cs.LG

Sampling from Binary Quadratic Distributions via Stochastic Localization

Sampling from binary quadratic distributions (BQDs) is a fundamental but challenging problem in discrete optimization and probabilistic inference. Previous work established theoretical guarantees for stochastic localization (SL) in continuous domains, where MCMC methods efficiently estimate the required posterior expectations during SL iterations. However, achieving similar convergence guarantees for discrete MCMC samplers in posterior estimation presents unique theoretical challenges. In this work, we present the first application of SL to general BQDs, proving that after a certain number of iterations, the external field of posterior distributions constructed by SL tends to infinity almost everywhere, hence satisfy Poincaré inequalities with probability near to 1, leading to polynomial-time mixing. This theoretical breakthrough enables efficient sampling from general BQDs, even those that may not originally possess fast mixing properties. Furthermore, our analysis, covering enormous discrete MCMC samplers based on Glauber dynamics and Metropolis-Hastings algorithms, demonstrates the broad applicability of our theoretical framework. Experiments on instances with quadratic unconstrained binary objectives, including maximum independent set, maximum cut, and maximum clique problems, demonstrate consistent improvements in sampling efficiency across different discrete MCMC samplers.

math.ST

Graph Neural Aggregation-diffusion with Metastability

Continuous graph neural models based on differential equations have expanded the architecture of graph neural networks (GNNs). Due to the connection between graph diffusion and message passing, diffusion-based models have been widely studied. However, diffusion naturally drives the system towards an equilibrium state, leading to issues like over-smoothing. To this end, we propose GRADE inspired by graph aggregation-diffusion equations, which includes the delicate balance between nonlinear diffusion and aggregation induced by interaction potentials. The node representations obtained through aggregation-diffusion equations exhibit metastability, indicating that features can aggregate into multiple clusters. In addition, the dynamics within these clusters can persist for long time periods, offering the potential to alleviate over-smoothing effects. This nonlinear diffusion in our model generalizes existing diffusion-based models and establishes a connection with classical GNNs. We prove that GRADE achieves competitive performance across various benchmarks and alleviates the over-smoothing issue in GNNs evidenced by the enhanced Dirichlet energy.

cs.LG

The local Poincare inequality of stochastic dynamic and application to the Ising model

Inspired by the idea of stochastic quantization proposed by Parisi and Wu, we construct the transition probability matrix which plays a central role in the renormalization group through a stochastic differential equation. By establishing the discrete time stochastic dynamics, the renormalization procedure can be characterized from the perspective of probability. Hence, we will focus on the investigation of the infinite dimensional stochastic dynamic. From the stochastic point of view, the discrete time stochastic dynamic can induce a Markov chain. Via calculating the square field operator and the Bakry-Émery curvature for a class of two-points functions, the local Poincaré inequality is established, from which the estimate of correlation functions can also be obtained. Finally, under the condition of ergodicity, by choosing the couple relationship between the system parameter $K$ and the system time $T$ properly when $T\rightarrow +\infty$, the two-points correlation functions for limit system are also estimated.

math.PR