SearcharxivSearch

arXiv subjects

Tony Bonnaire

Publications and source records attributed to Tony Bonnaire.

12 recordsLinked to original sources

Learning cosmic web environments with diffusion models

The cosmic web, consisting of an intricate network of voids, walls, filaments, and nodes, encodes key information about structure formation and the cosmological parameters that govern it. In the era of high-precision cosmology, large ensembles of numerical simulations are required to analyse next-generation galaxy surveys, motivating the use of generative models to circumvent their high computational cost. While they have shown promise in emulating high-fidelity cosmic web simulations, their ability to capture distinct cosmic web environments remains largely unexplored. For this study, we trained a diffusion model on the Quijote N-body simulation suite to investigate the semantic information learnt by its self-attention maps. Using statistical estimators such as the Dice coefficient and cross-power spectra, we quantified the correspondence between attention maps and cosmic web environments defined by the T-Web classifier. We find that attention maps of varying spatial resolutions across different layers capture overdense and underdense structures in distinct ways, exhibiting strong positive correlations and anti-correlations with both the overall matter distribution and individual cosmic web environments. Moreover, the diffusion model predominantly encodes cosmological information at intermediate-to-large spatial scales, indicating that attention maps primarily capture globally coherent structures. Our results show that, beyond accurately reproducing two-point statistics, diffusion models learn a multi-scale representation of the cosmic web through self-attention, including non-Gaussian information.

astro-ph.CO

Topological Exploration of High-Dimensional Empirical Risk Landscapes: general approach, and applications to phase retrieval

We consider the landscape of empirical risk minimization for high-dimensional Gaussian single-index models (generalized linear models). The objective is to recover an unknown signal $\boldsymbol{\theta}^\star \in \mathbb{R}^d$ (where $d \gg 1$) from a loss function $\hat{R}(\boldsymbol{\theta})$ that depends on pairs of labels $(\mathbf{x}_i \cdot \boldsymbol{\theta}, \mathbf{x}_i \cdot \boldsymbol{\theta}^\star)_{i=1}^n$, with $\mathbf{x}_i \sim \mathcal{N}(0, I_d)$, in the proportional asymptotic regime $n \asymp d$. Using the Kac-Rice formula, we analyze different complexities of the landscape -- defined as the expected number of critical points -- corresponding to various types of critical points, including local minima. We first show that some variational formulas previously established in the literature for these complexities can be drastically simplified, reducing to explicit variational problems over a finite number of scalar parameters that we can efficiently solve numerically. Our framework also provides detailed predictions for properties of the critical points, including the spectral properties of the Hessian and the joint distribution of labels. We apply our analysis to the real phase retrieval problem for which we derive complete topological phase diagrams of the loss landscape, characterizing notably BBP-type transitions where the Hessian at local minima (as predicted by the Kac-Rice formula) becomes unstable in the direction of the signal. We test the predictive power of our analysis to characterize gradient flow dynamics, finding excellent agreement with finite-size simulations of local optimization algorithms, and capturing fine-grained details such as the empirical distribution of labels. Overall, our results open new avenues for the asymptotic study of loss landscapes and topological trivialization phenomena in high-dimensional statistical models.

stat.ML

X-ray emission in IllustrisTNG circum-cluster environments. II -- Possible origins of the soft X-ray excess emission

An excess of soft X-ray emission (0.2-1 keV) above the contribution from the hot intra-cluster medium (ICM) has been detected in a number of galaxy clusters, including the Coma cluster. The physical origin of this emitting medium above hot ICM has not yet been determined, especially whether it be thermal or non-thermal. We aim to investigate which gas phase and gas structure more accurately reproduce the soft excess radiation from the cluster core to the outskirts, using simulations. By using the simulation TNG300, we predict the radial profile of thermodynamic properties and the Soft-X-ray surface brightness of 138 clusters within 5 $R_{200}$. Their X-ray emission is simulated for the hot ICM gas phase, the entire Warm-Hot medium, the diffuse and low-density Warm-Hot Intergalactic Medium (WHIM). Inside clusters, the soft excess appears to be produced by substructures of the WARM gas phase which host dense warm clumps (i.e, the Warm Circum-Galactic Medium, WCGM), and in fact the inner soft excess is strongly correlated with substructure and WCGM mass fractions. Outside of the virial radius, the fraction of WHIM gas that is mostly inside filaments connected to clusters boosts the soft X-ray excess. The more diffuse the gas is, the higher the soft X-ray excess beyond the virial region. The thermal emission of WARM gas phase, in the form of WCGM clumps and WHIM diffuse filaments, reproduces well the soft excess emission that was observed up to the virial radius in Coma and in the inner regions of other massive clusters. Moreover, our analysis suggests that soft X-ray excess is a proxy of cluster dynamical state, with larger excess being observed in the most unrelaxed clusters.

astro-ph.CO

On the role of non-linear latent features in bipartite generative neural networks

We investigate the phase diagram and memory retrieval capabilities of bipartite energy-based neural networks, namely Restricted Boltzmann Machines (RBMs), as a function of the prior distribution imposed on their hidden units - including binary, multi-state, and ReLU-like activations. Drawing connections to the Hopfield model and employing analytical tools from statistical physics of disordered systems, we explore how the architectural choices and activation functions shape the thermodynamic properties of these models. Our analysis reveals that standard RBMs with binary hidden nodes and extensive connectivity suffer from reduced critical capacity, limiting their effectiveness as associative memories. To address this, we examine several modifications, such as introducing local biases and adopting richer hidden unit priors. These adjustments restore ordered retrieval phases and markedly improve recall performance, even at finite temperatures. Our theoretical findings, supported by finite-size Monte Carlo simulations, highlight the importance of hidden unit design in enhancing the expressive power of RBMs.

cond-mat.dis-nn

Why Diffusion Models Don't Memorize: The Role of Implicit Dynamical Regularization in Training

Diffusion models have achieved remarkable success across a wide range of generative tasks. A key challenge is understanding the mechanisms that prevent their memorization of training data and allow generalization. In this work, we investigate the role of the training dynamics in the transition from generalization to memorization. Through extensive experiments and theoretical analysis, we identify two distinct timescales: an early time $\tau_\mathrm{gen}$ at which models begin to generate high-quality samples, and a later time $\tau_\mathrm{mem}$ beyond which memorization emerges. Crucially, we find that $\tau_\mathrm{mem}$ increases linearly with the training set size $n$, while $\tau_\mathrm{gen}$ remains constant. This creates a growing window of training times with $n$ where models generalize effectively, despite showing strong memorization if training continues beyond it. It is only when $n$ becomes larger than a model-dependent threshold that overfitting disappears at infinite training times. These findings reveal a form of implicit dynamical regularization in the training dynamics, which allow to avoid memorization even in highly overparameterized settings. Our results are supported by numerical experiments with standard U-Net architectures on realistic and synthetic datasets, and by a theoretical analysis using a tractable random features model studied in the high-dimensional limit.

cs.LG

Tracing gaseous filaments connected to galaxy clusters: the case study of Abell 2744

Filaments connected to galaxy clusters are crucial environments to study the building up of cosmic structures as they funnel matter towards the clusters' deep gravitational potentials. Identifying gas in filaments is a challenge, due to their lower density contrast which produces faint signals. The best chance to detect these signals is therefore in the outskirts of galaxy clusters. We revisit the X-ray observation of the cluster Abell 2744 using statistical estimators of anisotropic matter distribution to identify filamentary patterns around it. We report for the first time the blind detection of filaments connected to a galaxy cluster from X-ray emission using a filament-finder technique and a multipole decomposition technique. We compare this result with filaments extracted from the distribution of spectroscopic galaxies, through which we demonstrate the robustness and reliability of our techniques in tracing a filamentary structure of 3 to 5 filaments connected to Abell 2744.

astro-ph.CO

The Role of the Time-Dependent Hessian in High-Dimensional Optimization

Gradient descent is commonly used to find minima in rough landscapes, particularly in recent machine learning applications. However, a theoretical understanding of why good solutions are found remains elusive, especially in strongly non-convex and high-dimensional settings. Here, we focus on the phase retrieval problem as a typical example, which has received a lot of attention recently in theoretical machine learning. We analyze the Hessian during gradient descent, identify a dynamical transition in its spectral properties, and relate it to the ability of escaping rough regions in the loss landscape. When the signal-to-noise ratio (SNR) is large enough, an informative negative direction exists in the Hessian at the beginning of the descent, i.e in the initial condition. While descending, a BBP transition in the spectrum takes place in finite time: the direction is lost, and the dynamics is trapped in a rugged region filled with marginally stable bad minima. Surprisingly, for finite system sizes, this window of negative curvature allows the system to recover the signal well before the theoretical SNR found for infinite sizes, emphasizing the central role of initialization and early-time dynamics for efficiently navigating rough landscapes.

cs.LG

Dynamical Regimes of Diffusion Models

Using statistical physics methods, we study generative diffusion models in the regime where the dimension of space and the number of data are large, and the score function has been trained optimally. Our analysis reveals three distinct dynamical regimes during the backward generative diffusion process. The generative dynamics, starting from pure noise, encounters first a 'speciation' transition where the gross structure of data is unraveled, through a mechanism similar to symmetry breaking in phase transitions. It is followed at later time by a 'collapse' transition where the trajectories of the dynamics become attracted to one of the memorized data points, through a mechanism which is similar to the condensation in a glass phase. For any dataset, the speciation time can be found from a spectral analysis of the correlation matrix, and the collapse time can be found from the estimation of an 'excess entropy' in the data. The dependence of the collapse time on the dimension and number of data provides a thorough characterization of the curse of dimensionality for diffusion models. Analytical solutions for simple models like high-dimensional Gaussian mixtures substantiate these findings and provide a theoretical framework, while extensions to more complex scenarios and numerical validations with real datasets confirm the theoretical predictions.

cs.LG

High-Dimensional Non-Convex Landscapes and Gradient Descent Dynamics

In these lecture notes we present different methods and concepts developed in statistical physics to analyze gradient descent dynamics in high-dimensional non-convex landscapes. Our aim is to show how approaches developed in physics, mainly statistical physics of disordered systems, can be used to tackle open questions on high-dimensional dynamics in Machine Learning.

cond-mat.dis-nn

Regularization of Mixture Models for Robust Principal Graph Learning

A regularized version of Mixture Models is proposed to learn a principal graph from a distribution of $D$-dimensional data points. In the particular case of manifold learning for ridge detection, we assume that the underlying manifold can be modeled as a graph structure acting like a topological prior for the Gaussian clusters turning the problem into a maximum a posteriori estimation. Parameters of the model are iteratively estimated through an Expectation-Maximization procedure making the learning of the structure computationally efficient with guaranteed convergence for any graph prior in a polynomial time. We also embed in the formalism a natural way to make the algorithm robust to outliers of the pattern and heteroscedasticity of the manifold sampling coherently with the graph structure. The method uses a graph prior given by the minimum spanning tree that we extend using random sub-samplings of the dataset to take into account cycles that can be observed in the spatial distribution.

cs.LG

Cosmology with cosmic web environments II. Redshift-space auto and cross power spectra

Degeneracies among parameters of the cosmological model are known to drastically limit the information contained in the matter distribution. In the first paper of this series, we shown that the cosmic web environments; namely the voids, walls, filaments and nodes; can be used as a leverage to improve the real-space constraints on a set of six cosmological parameters, including the summed neutrino mass. Following-upon these results, we propose to study the achievable constraints of environment-dependent power spectra in redshift space where the velocities add up information to the standard two-point statistics by breaking the isotropy of the matter density field. A Fisher analysis based on a set of thousands of Quijote simulations allows us to conclude that the combination of power spectra computed in the several cosmic web environments is able to break some degeneracies. Compared to the matter monopole and quadrupole information alone, the combination of environment-dependent spectra tightens down the constraints on key parameters like the matter density or the summed neutrino mass by up to a factor of $5.5$. Additionally, while the information contained in the matter statistic quickly saturates at mildly non-linear scales in redshift space, the combination of power spectra in the environments appears as a goldmine of information able to improve the constraints at all the studied scales from $0.1$ to $0.5$ $h$/Mpc and suggests that further improvements are reachable at even finer scales.

astro-ph.CO

Cosmology with cosmic web environments I. Real-space power spectra

We undertake the first comprehensive and quantitative real-space analysis of the cosmological information content in the environments of the cosmic web (voids, filaments, walls, and nodes) up to non-linear scales, $k = 0.5$ $h$/Mpc. Relying on the large set of $N$-body simulations from the Quijote suite, the environments are defined through the eigenvalues of the tidal tensor and the Fisher formalism is used to assess the constraining power of the power spectra derived in each of the four environments and their combination. Our results show that there is more information available in the environment-dependent power spectra, both individually and when combined all together, than in the matter power spectrum. By breaking some key degeneracies between parameters of the cosmological model such as $M_ν$--$σ_\mathrm{8}$ or $Ω_\mathrm{m}$--$σ_8$, the power spectra computed in identified environments improve the constraints on cosmological parameters by factors $\sim 15$ for the summed neutrino mass $M_ν$ and $\sim 8$ for the matter density $Ω_\mathrm{m}$ over those derived from the matter power spectrum. We show that these tighter constraints are obtained for a wide range of the maximum scale, from $k_\mathrm{max} = 0.1$ $h$/Mpc to highly non-linear regimes with $k_\mathrm{max} = 0.5$ $h$/Mpc. We also report an eight times higher value of the signal-to-noise ratio for the combination of spectra compared to the matter one. Importantly, we show that all the presented results are robust to variations of the parameters defining the environments hence suggesting a robustness to the definition we chose to define them.

astro-ph.CO