SearcharxivSearch

arXiv subjects

Marc Mezard

Publications and source records attributed to Marc Mezard.

At least 19 recordsLinked to original sources

Classifier-Free Guidance: From High-Dimensional Analysis to Generalized Guidance Forms

Classifier-Free Guidance (CFG) is a widely adopted technique in diffusion and flow-based generative models, enabling high-quality conditional generation. A key theoretical challenge is characterizing the distribution induced by CFG, particularly in high-dimensional settings relevant to real-world data. Previous works have shown that CFG modifies the target distribution, steering it towards a distribution sharper than the target one, more shifted towards the boundary of the class. In this work, we provide a high-dimensional analysis of CFG, showing that these distortions vanish as the data dimension grows. We present a blessing-of-dimensionality result demonstrating that in sufficiently high and infinite dimensions, CFG accurately reproduces the target distribution. Using our high-dimensional theory, we show that there is a large family of guidances enjoying this property, in particular non-linear CFG generalizations. We study a simple non-linear power-law version, for which we demonstrate improved robustness, sample fidelity and diversity. Our findings are validated with experiments on class-conditional and text-to-image generation using state-of-the-art diffusion and flow-matching models.

cs.LG

Optimizing Noise Schedules of Generative Models in High Dimensionss

Recent works have shown that diffusion models can undergo phase transitions, the resolution of which is needed for accurately generating samples. This has motivated the use of different noise schedules, the two most common choices being referred to as variance preserving (VP) and variance exploding (VE). Here we revisit these schedules within the framework of stochastic interpolants. Using the Gaussian Mixture (GM) and Curie-Weiss (CW) data distributions as test case models, we first investigate the effect of the variance of the initial noise distribution and show that VP recovers the low-level feature (the distribution of each mode) but misses the high-level feature (the asymmetry between modes), whereas VE performs oppositely. We also show that this dichotomy, which happens when denoising by a constant amount in each step, can be avoided by using noise schedules specific to VP and VE that allow for the recovery of both high- and low-level features. Finally we show that these schedules yield generative models for the GM and CW model whose probability flow ODE can be discretized using $\Theta_d(1)$ steps in dimension $d$ instead of the $\Theta_d(\sqrt{d})$ steps required by constant denoising.

cs.LG

Mean-field message-passing equations in the Hopfield model and its generalizations

Motivated by recent progress in using restricted Boltzmann machines as preprocessing algorithms for deep neural network, we revisit the mean-field equations (belief-propagation and TAP equations) in the best understood such machine, namely the Hopfield model of neural networks, and we explicit how they can be used as iterative message-passing algorithms, providing a fast method to compute the local polarizations of neurons. In the "retrieval phase" where neurons polarize in the direction of one memorized pattern, we point out a major difference between the belief propagation and TAP equations : the set of belief propagation equations depends on the pattern which is retrieved, while one can use a unique set of TAP equations. This makes the latter method much better suited for applications in the learning process of restricted Boltzmann machines. In the case where the patterns memorized in the Hopfield model are not independent, but are correlated through a combinatorial structure, we show that the TAP equations have to be modified. This modification can be seen either as an alteration of the reaction term in TAP equations, or, more interestingly, as the consequence of message passing on a graphical model with several hidden layers, where the number of hidden layers depends on the depth of the correlations in the memorized patterns. This layered structure is actually necessary when one deals with more general restricted Boltzmann machines.

cond-mat.dis-nn

Belief Propagation Reconstruction for Discrete Tomography

We consider the reconstruction of a two-dimensional discrete image from a set of tomographic measurements corresponding to the Radon projection. Assuming that the image has a structure where neighbouring pixels have a larger probability to take the same value, we follow a Bayesian approach and introduce a fast message-passing reconstruction algorithm based on belief propagation. For numerical results, we specialize to the case of binary tomography. We test the algorithm on binary synthetic images with different length scales and compare our results against a more usual convex optimization approach. We investigate the reconstruction error as a function of the number of tomographic measurements, corresponding to the number of projection angles. The belief propagation algorithm turns out to be more efficient than the convex-optimization algorithm, both in terms of recovery bounds for noise-free projections, and in terms of reconstruction quality when moderate Gaussian noise is added to the projections.

math.NA

Level statistics of disordered spin-1/2 systems and its implications for materials with localized Cooper pairs

The origin of continuous energy spectra in large disordered interacting quantum systems is one of the key unsolved problems in quantum physics. While small quantum systems with discrete energy levels are noiseless and stay coherent forever in the absence of any coupling to external world, most large-scale quantum systems are able to produce thermal bath and excitation decay. This intrinsic decoherence is manifested by a broadening of energy levels which aquire a finite width. The important question is what is the driving force and the mechanism of transition(s) between two different types of many-body systems - with and without intrinsic decoherence? Here we address this question via the numerical study of energy level statistics of a system of spins-1/2 with anisotropic exchange interactions and random transverse fields. Our results present the first evidence for a well-defined quantum phase transition between domains of discrete and continous many-body spectra in a class of random spin models. Because this model also describes the physics of the superconductor-insulator transition in disordered superconductors like InO and similar materials, our results imply the appearance of novel insulating phases in the vicinity of this transition.

cond-mat.mes-hall

Emergence of rigidity at the structural glass transition: a first principle computation

We compute the shear modulus of structural glasses from a first principle approach based on the cloned liquid theory. We find that the intra-state shear-modulus, which corresponds to the plateau modulus measured in linear visco-elastic measurements, strongly depends on temperature and vanishes continuously when the temperature is increased beyond the glass temperature.

cond-mat.soft

Effect of coupling asymmetry on mean-field solutions of direct and inverse Sherrington-Kirkpatrick model

We study how the degree of symmetry in the couplings influences the performance of three mean field methods used for solving the direct and inverse problems for generalized Sherrington-Kirkpatrick models. In this context, the direct problem is predicting the potentially time-varying magnetizations. The three theories include the first and second order Plefka expansions, referred to as naive mean field (nMF) and TAP, respectively, and a mean field theory which is exact for fully asymmetric couplings. We call the last of these simply MF theory. We show that for the direct problem, nMF performs worse than the other two approximations, TAP outperforms MF when the coupling matrix is nearly symmetric, while MF works better when it is strongly asymmetric. For the inverse problem, MF performs better than both TAP and nMF, although an ad hoc adjustment of TAP can make it comparable to MF. For high temperatures the performance of TAP and MF approach each other.

cond-mat.dis-nn

On the solution of a `solvable' model of an ideal glass of hard spheres displaying a jamming transition

We discuss the analytical solution through the cavity method of a mean field model that displays at the same time an ideal glass transition and a set of jamming points. We establish the equations describing this system, and we discuss some approximate analytical solutions and a numerical strategy to solve them exactly. We compare these methods and we get insight into the reliability of the theory for the description of finite dimensional hard spheres.

cond-mat.dis-nn

The cavity method for quantum disordered systems: from transverse random field ferromagnets to directed polymers in random media

After reviewing the basics of the cavity method in classical systems, we show how its quantum version, with some appropriate approximation scheme, can be used to study a system of spins with random ferromagnetic interactions and a random transverse field. The quantum cavity equations describing the ferromagnetic-paramagnetic phase transition can be transformed into the well-known problem of a classical directed polymer in a random medium. The glass transition of this polymer problem translates ino the existence of a `Griffith phase' close to the quantum phase transition of the quantum spin problem, where the physics is dominated by rare events.

cond-mat.dis-nn

The Hierarchical Random Energy Model

We introduce a Random Energy Model on a hierarchical lattice where the interaction strength between variables is a decreasing function of their mutual hierarchical distance, making it a non-mean field model. Through small coupling series expansion and a direct numerical solution of the model, we provide evidence for a spin glass condensation transition similar to the one occuring in the usual mean field Random Energy Model. At variance with mean field, the high temperature branch of the free-energy is non-analytic at the transition point.

cond-mat.stat-mech

Glasses and replicas

We review the approach to glasses based on the replica formalism. The replica approach presented here is a first principle's approach which aims at deriving the main glass properties from the microscopic Hamiltonian. In contrast to the old use of replicas in the theory of disordered systems, this replica approach applies also to systems without quenched disorder (in this sense, replicas have nothing to do with computing the average of a logarithm of the partition function). It has the advantage of describing in an unified setting both the behaviour near the dynamic transition (mode coupling transition) and the behaviour near the equilibrium `transition' (Kauzmann transition) that is present in fragile glasses. The replica method may be used to solve simple mean field models, providing explicit examples of systems that may be studied analytically in great details and behave similarly to the experiments. Finally, using the replica formalism and some well adapted approximation schemes, it is possible to do explicit analytic computations of the properties of realistic models of glasses. The results of these first-principle computations are in reasonable agreement with numerical simulations. Draft of a chapter prepared for the book "Structural Glasses and Supercooled Liquids: Theory, Experiment, and Applications."

cond-mat.dis-nn

Energy transport in strongly disordered superconductors and magnets

We develop an analytical theory for quantum phase transitions driven by disorder in magnets and superconductors. We study these transitions with a cavity approximation which becomes exact on a Bethe lattice with large branching number. We find two different disordered phases, characterized by very different relaxation rates, which both exhibit strong inhomogeneities typical of glassy physics.

cond-mat.supr-con

Constraint satisfaction problems and neural networks: a statistical physics perspective

A new field of research is rapidly expanding at the crossroad between statistical physics, information theory and combinatorial optimization. In particular, the use of cutting edge statistical physics concepts and methods allow one to solve very large constraint satisfaction problems like random satisfiability, coloring, or error correction. Several aspects of these developments should be relevant for the understanding of functional complexity in neural networks. On the one hand the message passing procedures which are used in these new algorithms are based on local exchange of information, and succeed in solving some of the hardest computational problems. On the other hand some crucial inference problems in neurobiology, like those generated in multi-electrode recordings, naturally translate into hard constraint satisfaction problems. This paper gives a non-technical introduction to this field, emphasizing the main ideas at work in message passing strategies and their possible relevance to neural networks modeling. It also introduces a new message passing algorithm for inferring interactions between variables from correlation data, which could be useful in the analysis of multi-electrode recording data.

q-bio.NC

Statistical Physics of Group Testing

This paper provides a short introduction to the group testing problem, and reviews various aspects of its statistical physics formulation. Two main issues are discussed: the optimal design of pools used in a two-stage testing experiment, like the one often used in medical or biological applications, and the inference problem of detecting defective items based on pool diagnosis. The paper is largely based on: M. Mézard and C. Toninelli, arXiv:0706.3104, and M. Mézard and M. Tarzia {\it Phys. Rev. E} {\bf 76}, 041124 (2007).

cond-mat.stat-mech

Pairs of SAT Assignment in Random Boolean Formulae

We investigate geometrical properties of the random K-satisfiability problem using the notion of x-satisfiability: a formula is x-satisfiable if there exist two SAT assignments differing in Nx variables. We show the existence of a sharp threshold for this property as a function of the clause density. For large enough K, we prove that there exists a region of clause density, below the satisfiability threshold, where the landscape of Hamming distances between SAT assignments experiences a gap: pairs of SAT-assignments exist at small x, and around x=1/2, but they donot exist at intermediate values of x. This result is consistent with the clustering scenario which is at the heart of the recent heuristic analysis of satisfiability using statistical physics analysis (the cavity method), and its algorithmic counterpart (the survey propagation algorithm). The method uses elementary probabilistic arguments (first and second moment methods), and might be useful in other problems of computational and physical interest where similar phenomena appear.

cond-mat.dis-nn

Group Testing with Random Pools: optimal two-stage algorithms

We study Probabilistic Group Testing of a set of N items each of which is defective with probability p. We focus on the double limit of small defect probability, p<<1, and large number of variables, N>>1, taking either p->0 after $N\to\infty$ or $p=1/N^β$ with $β\in(0,1/2)$. In both settings the optimal number of tests which are required to identify with certainty the defectives via a two-stage procedure, $\bar T(N,p)$, is known to scale as $Np|\log p|$. Here we determine the sharp asymptotic value of $\bar T(N,p)/(Np|\log p|)$ and construct a class of two-stage algorithms over which this optimal value is attained. This is done by choosing a proper bipartite regular graph (of tests and variable nodes) for the first stage of the detection. Furthermore we prove that this optimal value is also attained on average over a random bipartite graph where all variables have the same degree, while the tests have Poisson-distributed degrees. Finally, we improve the existing upper and lower bound for the optimal number of tests in the case $p=1/N^β$ with $β\in[1/2,1)$.

cs.DS

Risk Minimization through Portfolio Replication

We use a replica approach to deal with portfolio optimization problems. A given risk measure is minimized using empirical estimates of asset values correlations. We study the phase transition which happens when the time series is too short with respect to the size of the portfolio. We also study the noise sensitivity of portfolio allocation when this transition is approached. We consider explicitely the cases where the absolute deviation and the conditional value-at-risk are chosen as a risk measure. We show how the replica method can study a wide range of risk measures, and deal with various types of time series correlations, including realistic ones with volatility clustering.

physics.soc-ph

Asymmetric quantum error correcting codes

The noise in physical qubits is fundamentally asymmetric: in most devices, phase errors are much more probable than bit flips. We propose a quantum error correcting code which takes advantage of this asymmetry and shows good performance at a relatively small cost in redundancy, requiring less than a doubling of the number of physical qubits for error correction.

quant-ph