Searcharxiv⌕ Search

arXiv subjects

Yongseok Jo

Publications and source records attributed to Yongseok Jo.

16 recordsLinked to original sources

Constructing a Mock Galaxy Catalog for the All-sky SPECtroscopic Survey of Nearby Galaxies (A-SPEC) Using the Machine-assisted Semi-Simulation Model

We present a methodology for constructing a mock galaxy catalog for the All-sky SPECtroscopic survey of nearby galaxies (A-SPEC) using the Machine-assisted Semi-Simulation Model. The model is trained on the cosmological magnetohydrodynamical simulation IllustrisTNG to predict baryonic properties of subhalos from dark-matter-only features and is applied to our own N-body simulation tailored to satisfy the requirements of A-SPEC. We have improved the model's accuracy by introducing additional features such as subhalo anisotropy parameters and modified definitions of the subhalo environment, which result in the coefficient of determination R^2=0.96, 0.90, 0.70, 0.79 for stellar mass, gas mass, star formation rate, and gas metallicity, respectively. The resulting mock galaxies reproduce the luminosity-dependent clustering of the target galaxies when tuned to match the number density. We discuss avenues for further improvement, including the role of environment in the predictions. We release the mock galaxy catalog with the baryonic properties predicted from the model.

astro-ph.CO↗

From Dense Gas Clouds to Supermassive Black Hole Seeds: Hybrid Hydro/Direct $N$-body Simulations of Runaway Collision-driven Intermediate-mass Black Hole Formation

A population of dense stellar systems at high redshift has recently been uncovered by the JWST. To investigate the formation of supermassive black hole (SMBH) seeds in these dense environments without invoking any \textit{ad hoc} seeding mechanisms, we present star cluster-scale simulations performed with an updated version of the hydrodynamics code \texttt{Enzo-Abyss}, which self-consistently integrates the gravity using a direct $N$-body method coupled with stellar evolution. By modeling initially dense, metal-poor gas clouds with varying turbulence, we consistently find the formation of dense clusters resembling early-stage nuclear star clusters (NSCs), as well as the formation of very massive stars (VMSs) ranging from $343\;\mathrm{M_\odot}$ to $5108\;\mathrm{M_\odot}$ via runaway collisions, irrespective of stellar wind feedback strength. Following the direct collapse of these VMSs, the resulting intermediate-mass black holes (IMBHs) grow through Eddington-limited gas accretion and tidal disruption events (TDEs). In our most optimistic model, we find a mass accretion rate of $1.64\times10^{-4}\;\mathrm{M_\odot\;yr^{-1}}$, with TDEs contributing $23\%$ of the total accretion over $\sim10\;\mathrm{Myr}$. Assuming a steady gas supply into the NSC driven by rapid structural assembly in the high-redshift environment, together with a constant TDE rate, we project that an IMBH with an initial mass of $6747\;\mathrm{M_\odot}$ at the center of the NSC can grow to $\sim62000\;\mathrm{M_\odot}$ within $100\;\mathrm{Myr}$ of its formation. Our numerical study, conducted within a single self-consistent framework that incorporates the essential physical processes, suggests that VMSs can form in dense gas clouds, collapse into IMBHs, and subsequently provide viable seeds for the SMBHs observed at high redshift.

astro-ph.GA↗

Learning the Universe with the 2nd Generation of CAMELS: Varying 35 parameters of the IllustrisTNG model in (50Mpc/h)^3 boxes

We present a new set of 1,192 cosmological simulations as part of the CAMELS project, in which a space of 35 cosmological, astrophysical, and numerical parameters is explored around the fiducial IllustrisTNG model. The volume of each of these simulations is (50Mpc/h)^3, eight times larger than that of previous CAMELS simulations. This provides lower sample variance as well as access to more massive halos and more diverse environments. We focus this work on exploring the advantages these differences provide for parameter inference powered by neural networks. We generate training sets based on the matter power spectra, projected maps of the volumes, graphs representing galaxy spatial distributions, and thermodynamical properties of massive halos. We employ multilayer perceptrons, convolutional neural networks, graph neural networks, and Gaussian processes, respectively, to extract information on the simulation parameters from these inputs while comparing systematically to analogous results from our previous generation of (25Mpc/h)^3 simulations. We generally find that the new, larger volumes produce tighter marginal constraints on the parameters, to degrees that vary between the different inputs. The improvements, however, scale more weakly than with the square root of the increase in the amount of data (i.e., physical volume). We interpret this as originating either from information loss due to mode coupling or from complex degeneracies in parameter space. We also discuss the effects on statistics of the intergalactic medium temperature from four new parameters that are varied in these simulations, which control the amplitude and timing of the ionizing background radiation. We publicly release the simulation outputs and ancillary data at https://camels.readthedocs.io.

astro-ph.CO↗

Learning the Stellar Structure Equations via Self-supervised Physics-Informed Neural Networks

Stellar astrophysics relies critically on accurate descriptions of the physical conditions inside stars. Traditional solvers such as \texttt{MESA} (Modules for Experiments in Stellar Astrophysics), which employ adaptive finite-difference methods, can become computationally expensive and challenging to scale for large stellar population synthesis ($>10^9$ stars). In this work, we present an self-supervised physics-informed neural network (PINN) framework that provides a mesh-free and fully differentiable approach to solving the stellar structure equations under hydrostatic and thermal equilibrium. The model takes as input the stellar boundary conditions (at the center and surface) together with the chemical composition, and learns continuous radial profiles for mass $M_r(r)$, pressure $P(r)$, density $ρ(r)$, temperature $T(r)$, and luminosity $L_r(r)$ by enforcing the governing structure equations through physics-based loss terms. To incorporate realistic microphysics, we introduce auxiliary neural networks that approximate the equation of state and opacity tables as smooth, differentiable functions of the local thermodynamic state. These surrogates replace traditional tabulated inputs and enable end-to-end training. Once trained for a given star, the model produces continuous solutions across the entire radial domain without requiring discretization or interpolation. Validation against benchmark \texttt{MESA} models across a range of stellar masses yields a Mean Relative Absolute Error of $3.06\%$ and an average $R^2$ score of $99.98\%$. To our knowledge, this is the first demonstration that the stellar structure equations can be solved in a fully self-supervised and data-free fashion employing PINNs. This work establishes a foundation for scalable, physics-informed emulation of stellar interiors and opens the door to future extensions toward time-dependent stellar evolution.

astro-ph.SR↗

Evolution of Nuclear Star Cluster in Dwarf Galaxy through Mergers and In-Situ Star Formation

Nuclear Star Clusters (NSCs) are dense stellar systems located at the centers of galaxies. Employing Enzo-Abyss, which integrates hydrodynamics with a direct N-body solver, we introduce a simulation capable of resolving the evolution of NSCs within a live galaxy. This includes live dark matter, gaseous dynamics, star formation and feedback, collisional dynamics for star clusters. The evolution of NSCs is typically shaped by two main processes: mergers of star clusters and in-situ star formation. Our simulation enables investigation of the contributions of these mechanisms to the growth of NSCs. This work focuses on the impact of stellar physics and gas content on the growth of NSCs within a dwarf galaxy. To this end, we carry out four simulations, a fiducial simulation, one without supernova feedback, one with low star formation efficiency, and one with higher galactic gas content. This study shows a likelihood that both mergers and in-situ star formation contribute to NSC evolution comparably. In addition, mergers result in disruption of dense gas clumps within star clusters, indicating that in-situ star formation is suppressed when mergers occur. However, the limitations -- such as the lack of individual star physics and limited spatial/particle mass resolution -- hinder drawing a definite conclusion. Nevertheless, with further development, our simulations will serve as a cornerstone that untangles the complex interplay between mergers and in-situ star formation in shaping the structure and mass of NSCs, thereby providing insights into their formation and evolution.

astro-ph.GA↗

Towards Robustness Across Cosmological Simulation Models TNG, SIMBA, ASTRID, and EAGLE

The rapid advancement of large-scale cosmological simulations has opened new avenues for cosmological and astrophysical research. However, the increasing diversity among cosmological simulation models presents a challenge to the robustness. In this work, we develop the Model-Insensitive ESTimator (MIEST), a machine that can robustly estimate the cosmological parameters, $Ω_m$ and $σ_8$, from neural hydrogen maps of simulation models in the CAMELS project$-$TNG, SIMBA, ASTRID, and EAGLE. An estimator is considered robust if it possesses a consistent predictive power across all simulations, including those used during the training phase. We train our machine using multiple simulation models and ensure that it only extracts common features between the models while disregarding the model-specific features. This allows us to develop a novel model that is capable of accurately estimating parameters across a range of simulation models, without being biased towards any particular model. Upon the investigation of the latent space$-$a set of summary statistics, we find that the implementation of robustness leads to the blending of latent variables across different models, demonstrating the removal of model-specific features. In comparison to a standard machine lacking robustness, the average performance of MIEST on the unseen simulations during the training phase has been improved by $\sim17$% for $Ω_m$ and $\sim 38$% for $σ_8$. By using a machine learning approach that can extract robust, yet physical features, we hope to improve our understanding of galaxy formation and evolution in a (subgrid) model-insensitive manner, and ultimately, gain insight into the underlying physical processes responsible for robustness. This is a Learning the Universe publication.

astro-ph.CO↗

On the Significance of Covariance for Constraining Theoretical Models From Galaxy Observables

In this study, we investigate the impact of covariance within uncertainties on the inference of cosmological and astrophysical parameters, specifically focusing on galaxy stellar mass functions derived from the CAMELS simulation suite. Utilizing both Fisher analysis and Implicit Likelihood Inference (ILI), we explore how different covariance structures, including simple toy models and physics-motivated uncertainties, affect posterior distributions and parameter variances. Our methodology utilizes forward modeling via emulators that are trained on CAMELS simulations to produce stellar mass functions based on input parameters, subsequently incorporating Gaussian noise as defined by covariance matrices. We examine both toy model covariance matrices and physically motivated covariance matrices derived from observational factors like the stellar Initial Mass Function (IMF) and photometric aperture size. Our results demonstrate that covariance terms significantly influence parameter inference, often leading to tighter constraints or revealing complex, multimodal posterior distributions. These findings underscore the necessity of accounting for covariance when interpreting astrophysical observations, especially in fields where accurate parameter estimation is critical for model validation and hypothesis testing.

astro-ph.CO↗

Inferring Cosmological Parameters on SDSS via Domain-Generalized Neural Networks and Lightcone Simulations

We present a proof-of-concept simulation-based inference on $Ω_{\rm m}$ and $σ_{8}$ from the SDSS BOSS LOWZ NGC catalog using neural networks and domain generalization techniques without the need of summary statistics. Using rapid lightcone simulations, ${\rm L{\scriptsize -PICOLA}}$, mock galaxy catalogs are produced that fully incorporate the observational effects. The collection of galaxies is fed as input to a point cloud-based network, ${\texttt{Minkowski-PointNet}}$. We also add relatively more accurate ${\rm G{\scriptsize ADGET}}$ mocks to obtain robust and generalizable neural networks. By explicitly learning the representations which reduces the discrepancies between the two different datasets via the semantic alignment loss term, we show that the latent space configuration aligns into a single plane in which the two cosmological parameters form clear axes. Consequently, during inference, the SDSS BOSS LOWZ NGC catalog maps onto the plane, demonstrating effective generalization and improving prediction accuracy compared to non-generalized models. Results from the ensemble of 25 independently trained machines find $Ω_{\rm m}=0.339 \pm 0.056$ and $σ_{8}=0.801 \pm 0.061$, inferred only from the distribution of galaxies in the lightcone slices without relying on any indirect summary statistics. A single machine that best adapts to the ${\rm G{\scriptsize ADGET}}$ mocks yields a tighter prediction of $Ω_{\rm m}=0.282 \pm 0.014$ and $σ_{8}=0.786 \pm 0.036$. We emphasize that adaptation across multiple domains can enhance the robustness of the neural networks in observational data.

astro-ph.CO↗

Evolution of Star Cluster Within Galaxy using Self-consistent Hybrid Hydro/N-body Simulation

We introduce a GPU-accelerated hybrid hydro/N-body code (Enzo-N) designed to address the challenges of concurrently simulating star clusters and their parent galaxies. This task has been exceedingly challenging, primarily due to the considerable computational time required, which stems from the substantial scale difference between galaxies (~ 0.1 Mpc) and star clusters (~ pc). Yet, this significant scale separation means that particles within star clusters perceive those outside the star cluster in a semi-stationary state. By leveraging this aspect, we integrate the direct N-body code (Nbody6++GPU) into the cosmological (magneto-)hydrodynamic code (Enzo) through the utilization of the semi-stationary background acceleration approximation. We solve the dynamics of particles within star clusters using the direct N-body solver with regularization for few-body interactions, while evolving particles outside -- dark matter, gas, and stars -- using the particle-mesh gravity solver and hydrodynamic methods. We demonstrate that Enzo-N successfully simulates the co-evolution of star clusters and their parent galaxies, capturing phenomena such as core collapse of the star cluster and tidal stripping due to galactic tides. This comprehensive framework opens up new possibilities for studying the evolution of star clusters within galaxies, offering insights that were previously inaccessible.

astro-ph.GA↗

The CAMELS project: Expanding the galaxy formation model space with new ASTRID and 28-parameter TNG and SIMBA suites

We present CAMELS-ASTRID, the third suite of hydrodynamical simulations in the Cosmology and Astrophysics with MachinE Learning (CAMELS) project, along with new simulation sets that extend the model parameter space based on the previous frameworks of CAMELS-TNG and CAMELS-SIMBA, to provide broader training sets and testing grounds for machine-learning algorithms designed for cosmological studies. CAMELS-ASTRID employs the galaxy formation model following the ASTRID simulation and contains 2,124 hydrodynamic simulation runs that vary 3 cosmological parameters ($Ω_m$, $σ_8$, $Ω_b$) and 4 parameters controlling stellar and AGN feedback. Compared to the existing TNG and SIMBA simulation suites in CAMELS, the fiducial model of ASTRID features the mildest AGN feedback and predicts the least baryonic effect on the matter power spectrum. The training set of ASTRID covers a broader variation in the galaxy populations and the baryonic impact on the matter power spectrum compared to its TNG and SIMBA counterparts, which can make machine-learning models trained on the ASTRID suite exhibit better extrapolation performance when tested on other hydrodynamic simulation sets. We also introduce extension simulation sets in CAMELS that widely explore 28 parameters in the TNG and SIMBA models, demonstrating the enormity of the overall galaxy formation model parameter space and the complex non-linear interplay between cosmology and astrophysical processes. With the new simulation suites, we show that building robust machine-learning models favors training and testing on the largest possible diversity of galaxy formation models. We also demonstrate that it is possible to train accurate neural networks to infer cosmological parameters using the high-dimensional TNG-SB28 simulation set.

astro-ph.CO↗

Calibrating cosmological simulations with implicit likelihood inference using galaxy growth observables

In a novel approach employing implicit likelihood inference (ILI), also known as likelihood-free inference, we calibrate the parameters of cosmological hydrodynamic simulations against observations, which has previously been unfeasible due to the high computational cost of these simulations. For computational efficiency, we train neural networks as emulators on ~1000 cosmological simulations from the CAMELS project to estimate simulated observables, taking as input the cosmological and astrophysical parameters, and use these emulators as surrogates to the cosmological simulations. Using the cosmic star formation rate density (SFRD) and, separately, stellar mass functions (SMFs) at different redshifts, we perform ILI on selected cosmological and astrophysical parameters (Omega_m, sigma_8, stellar wind feedback, and kinetic black hole feedback) and obtain full 6-dimensional posterior distributions. In the performance test, the ILI from the emulated SFRD (SMFs) can recover the target observables with a relative error of 0.17% (0.4%). We find that degeneracies exist between the parameters inferred from the emulated SFRD, confirmed with new full cosmological simulations. We also find that the SMFs can break the degeneracy in the SFRD, which indicates that the SMFs provide complementary constraints for the parameters. Further, we find that the parameter combination inferred from an observationally-inferred SFRD reproduces the target observed SFRD very well, whereas, in the case of the SMFs, the inferred and observed SMFs show significant discrepancies that indicate potential limitations of the current galaxy formation modeling and calibration framework, and/or systematic differences and inconsistencies between observations of the stellar mass function.

astro-ph.CO↗

The CAMELS project: public data release

The Cosmology and Astrophysics with MachinE Learning Simulations (CAMELS) project was developed to combine cosmology with astrophysics through thousands of cosmological hydrodynamic simulations and machine learning. CAMELS contains 4,233 cosmological simulations, 2,049 N-body and 2,184 state-of-the-art hydrodynamic simulations that sample a vast volume in parameter space. In this paper we present the CAMELS public data release, describing the characteristics of the CAMELS simulations and a variety of data products generated from them, including halo, subhalo, galaxy, and void catalogues, power spectra, bispectra, Lyman-$α$ spectra, probability distribution functions, halo radial profiles, and X-rays photon lists. We also release over one thousand catalogues that contain billions of galaxies from CAMELS-SAM: a large collection of N-body simulations that have been combined with the Santa Cruz Semi-Analytic Model. We release all the data, comprising more than 350 terabytes and containing 143,922 snapshots, millions of halos, galaxies and summary statistics. We provide further technical details on how to access, download, read, and process the data at \url{https://camels.readthedocs.io}.

astro-ph.CO↗

The CAMELS Multifield Dataset: Learning the Universe's Fundamental Parameters with Artificial Intelligence

We present the Cosmology and Astrophysics with MachinE Learning Simulations (CAMELS) Multifield Dataset, CMD, a collection of hundreds of thousands of 2D maps and 3D grids containing many different properties of cosmic gas, dark matter, and stars from 2,000 distinct simulated universes at several cosmic times. The 2D maps and 3D grids represent cosmic regions that span $\sim$100 million light years and have been generated from thousands of state-of-the-art hydrodynamic and gravity-only N-body simulations from the CAMELS project. Designed to train machine learning models, CMD is the largest dataset of its kind containing more than 70 Terabytes of data. In this paper we describe CMD in detail and outline a few of its applications. We focus our attention on one such task, parameter inference, formulating the problems we face as a challenge to the community. We release all data and provide further technical details at https://camels-multifield-dataset.readthedocs.io.

cs.LG↗

Dark Matter Deficient Galaxies Produced Via High-velocity Galaxy Collisions In High-resolution Numerical Simulations

The recent discovery of diffuse dwarf galaxies that are deficient in dark matter appears to challenge the current paradigm of structure formation in our Universe. We describe the numerical experiments to determine if the so-called dark matter deficient galaxies (DMDGs) could be produced when two gas-rich, dwarf-sized galaxies collide with a high relative velocity of $\sim 300\,{\rm kms^{-1}}$. Using idealized high-resolution simulations with both mesh-based and particle-based gravito-hydrodynamics codes, we find that DMDGs can form as high-velocity galaxy collisions separate dark matter from the warm disk gas which subsequently is compressed by shock and tidal interaction to form stars. Then using a large simulated universe IllustrisTNG, we discover a number of high-velocity galaxy collision events in which DMDGs are expected to form. However, we did not find evidence that these types of collisions actually produced DMDGs in the TNG100-1 run. We argue that the resolution of the numerical experiment is critical to realize the "collision-induced" DMDG formation scenario. Our results demonstrate one of many routes in which galaxies could form with unconventional dark matter fractions.

astro-ph.GA↗

High-redshift Galaxy Formation with Self-consistently Modeled Stars and Massive Black Holes: Stellar Feedback and Quasar Growth

As computational resolution of modern cosmological simulations reach ever so close to resolving individual star-forming clumps in a galaxy, a need for "resolution-appropriate" physics for a galaxy-scale simulation has never been greater. To this end, we introduce a self-consistent numerical framework that includes explicit treatments of feedback from star-forming molecular clouds (SFMCs) and massive black holes (MBHs). In addition to the thermal supernovae feedback from SFMC particles, photoionizing radiation from both SFMCs and MBHs is tracked through full 3-dimensional ray tracing. A mechanical feedback channel from MBHs is also considered. Using our framework, we perform a state-of-the-art cosmological simulation of a quasar-host galaxy at z~7.5 for ~25 Myrs with all relevant galactic components such as dark matter, gas, SFMCs, and an embedded MBH seed of ~> 1e6 Ms. We find that feedback from SFMCs and an accreting MBH suppresses runaway star formation locally in the galactic core region. Newly included radiation feedback from SFMCs, combined with feedback from the MBH, helps the MBH grow faster by retaining gas that eventually accretes on to the MBH. Our experiment demonstrates that previously undiscussed types of interplay between gas, SFMCs, and a MBH may hold important clues about the growth and feedback of quasars and their host galaxies in the high-redshift Universe.

astro-ph.GA↗

Machine-assisted Semi-Simulation Model (MSSM): Estimating Galactic Baryonic Properties from their Dark Matter using a Machine Trained on Hydrodynamic Simulations

We present a pipeline to estimate baryonic properties of a galaxy inside a dark matter (DM) halo in DM-only simulations using a machine trained on high-resolution hydrodynamic simulations. As an example, we use the IllustrisTNG hydrodynamic simulation of a $(75 \,\,h^{-1}{\rm Mpc})^3$ volume to train our machine to predict e.g., stellar mass and star formation rate in a galaxy-sized halo based purely on its DM content. An extremely randomized tree (ERT) algorithm is used together with multiple novel improvements we introduce here such as a refined error function in machine training and two-stage learning. Aided by these improvements, our model demonstrates a significantly increased accuracy in predicting baryonic properties compared to prior attempts --- in other words, the machine better mimics IllustrisTNG's galaxy-halo correlation. By applying our machine to the MultiDark-Planck DM-only simulation of a large $(1 \,\,h^{-1}{\rm Gpc})^3$ volume, we then validate the pipeline that rapidly generates a galaxy catalogue from a DM halo catalogue using the correlations the machine found in IllustrisTNG. We also compare our galaxy catalogue with the ones produced by popular semi-analytic models (SAMs). Our so-called machine-assisted semi-simulation model (MSSM) is shown to be largely compatible with SAMs, and may become a promising method to transplant the baryon physics of galaxy-scale hydrodynamic calculations onto a larger-volume DM-only run. We discuss the benefits that machine-based approaches like this entail, as well as suggestions to raise the scientific potential of such approaches.

astro-ph.GA↗