SearcharxivSearch

arXiv subjects

Jaehong Park

Publications and source records attributed to Jaehong Park.

At least 19 recordsLinked to original sources

Semantic-Aware Reconstruction Error for Detecting AI-Generated Images

Recently, AI-generated image detection has gained increasing attention, as the rapid advancement of image generation technologies has raised serious concerns about their potential misuse. While existing detection methods have achieved promising results, their performance often degrades significantly when facing fake images from unseen, out-of-distribution (OOD) generative models, since they primarily rely on model-specific artifacts and thus overfit to the models used for training. To address this limitation, we propose a novel representation, namely Semantic-Aware Reconstruction Error (SARE), that measures the semantic difference between an image and its caption-guided reconstruction. The key hypothesis behind SARE is that real images, whose captions often fail to fully capture their complex visual content, may undergo noticeable semantic shifts during the caption-guided reconstruction process. In contrast, fake images, which closely align with their captions, show minimal semantic changes. By quantifying these semantic shifts, SARE provides a robust and discriminative feature for detecting fake images across diverse generative models. Additionally, we introduce a fusion module that integrates SARE into the backbone detector via a cross-attention mechanism. Image features attend to semantic representations extracted from SARE, enabling the model to adaptively leverage semantic information. Experimental results demonstrate that the proposed method achieves strong generalization, outperforming existing baselines on benchmarks including GenImage and ForenSynths. We further validate the effectiveness of caption guidance through a detailed analysis of semantic shifts, confirming its ability to enhance detection robustness.

cs.CV

Galaxy populations in protoclusters at cosmic noon

We investigate the physical properties and redshift evolution of simulated galaxies residing in protoclusters at cosmic noon, to understand the influence of the environment on galaxy formation. This work is to build clear expectations for the ongoing ODIN survey, devoted to mapping large-scale structures at z=2.4, 3.1, and 4.5 using Ly$α$-emitting galaxies (LAEs) as tracers. From the IllustrisTNG simulations, we define subregions centered on the most massive clusters ranked by total stellar mass at z=0 and study the properties of galaxies within, including LAEs. To model the LAE population, we take a semi-analytical approach that assigns Ly$α$ luminosity and equivalent width based on the UV luminosities to galaxies in a probabilistic manner. We investigate stellar mass, star formation rate, major mergers, and specific star formation rate of the population of star-forming galaxies and LAEs in the field and protocluster environment and trace their evolution. We find that the overall shape of the UV luminosity function (LF) in simulated protocluster environments is characterized by a shallower faint-end slope and an excess on the bright end, signaling different formation histories for galaxies therein. The difference is milder for the Ly$α$ LF. While protocluster galaxies follow the same SFR-$M_{\odot}$ scaling relation as average field galaxies, a larger fraction appears to have experienced major mergers in the last 200 Myr and as a result shows enhanced star formation at a ~60% level, leading to a flatter distribution in both SFR and $M_{\odot}$ relative to galaxies in the average field. We find that protocluster galaxies, including LAEs, begin to quench much earlier (z~0.8-1.6) than field galaxies (z~0.5-0.9); our result is in agreement with recent observational results and highlights the importance of large-scale environment on the overall formation history of galaxies.

astro-ph.GA

Testing Lyman Alpha Emitters and Lyman-Break Galaxies as Tracers of Large-Scale Structures at High Redshifts

We test whether Lyman alpha emitters (LAEs) and Lyman-break galaxies (LBGs) can be good tracers of high-z large-scale structures, using the Horizon Run 5 cosmological hydrodynamical simulation. We identify LAEs using the Lyα emission line luminosity and its equivalent width, and LBGs using the broad-band magnitudes at z~2.4, 3.1, and 4.5. We first compare the spatial distributions of LAEs, LBGs, all galaxies, and dark matter around the filamentary structures defined by dark matter. The comparison shows that both LAEs and LBGs are more concentrated toward the dark matter filaments than dark matter. We also find an empirical fitting formula for the vertical density profile of filaments as a binomial power-law relation of the distance to the filaments. We then compare the spatial distributions of the samples around the filaments defined by themselves. LAEs and LBGs are again more concentrated toward their filaments than dark matter. We also find the overall consistency between filamentary structures defined by LAEs, LBGs, and dark matter, with the median spatial offsets that are smaller than the mean separation of the sample. These results support the idea that the LAEs and LBGs could be good tracers of large-scale structures of dark matter at high redshifts.

astro-ph.GA

The One-hundred-deg^2 DECam Imaging in Narrowbands (ODIN): Survey Design and Science Goals

We describe the survey design and science goals for ODIN (One-hundred-deg^2 DECam Imaging in Narrowbands), a NOIRLab survey using the Dark Energy Camera (DECam) to obtain deep (AB~25.7) narrow-band images over an unprecedented area of sky. The three custom-built narrow-band filters, N419, N501, and N673, have central wavelengths of 419, 501, and 673 nm and respective full-widthat-half-maxima of 7.2, 7.4, and 9.8 nm, corresponding to Lya at z=2.4, 3.1, and 4.5 and cosmic times of 2.8, 2.1, and 1.4 Gyr, respectively. When combined with even deeper, public broad-band data from Hyper Suprime-Cam, DECam, and in the future, LSST, the ODIN narrow-band images will enable the selection of over 100,000 Lya-emitting (LAE) galaxies at these epochs. ODIN-selected LAEs will identify protoclusters as galaxy overdensities, and the deep narrow-band images enable detection of highly extended Lya blobs (LABs). Primary science goals include measuring the clustering strength and dark matter halo connection of LAEs, LABs, and protoclusters, and their respective relationship to filaments in the cosmic web. The three epochs allow the redshift evolution of these properties to be determined during the period known as Cosmic Noon, where star formation was at its peak. The two narrow-band filter wavelengths are designed to enable interloper rejection and further scientific studies by revealing [O II] and [O III] at z=0.34, Lya and He II 1640 at z=3.1, and Lyman continuum plus Lya at z=4.5. Ancillary science includes similar studies of the lower-redshift emission-line galaxy samples and investigations of nearby star-forming galaxies resolved into numerous [O III] and [S II] emitting regions.

astro-ph.GA

21cmfish: Fisher-matrix framework for fast parameter forecasts from the cosmic 21-cm signal

The 21-cm signal from neutral hydrogen in the early universe will provide unprecedented information about the first stars and galaxies. Extracting this information, however, requires accounting for many unknown astrophysical processes. Semi-numerical simulations are key for exploring the vast parameter space of said processes. These simulations use approximate techniques such as excursion-set and perturbation theory to model the 3D evolution of the intergalactic medium, at a fraction of the computational cost of hydrodynamic and/or radiative transfer simulations. However, exploring the enormous parameter space of the first galaxies can still be computationally expensive. Here we introduce 21cmfish, a Fisher-matrix wrapper for the semi-numerical simulation 21cmFAST. 21cmfish facilitates efficient parameter forecasts, scaling to significantly higher dimensionalities than MCMC approaches, assuming a multi-variate Gaussian posterior. Our method produces comparable parameter uncertainty forecasts to previous MCMC analyses but requires ~10$^4$x fewer simulations. This enables a rapid way to prototype analyses adding new physics and/or additional parameters. We carry out a forecast for HERA using the largest astrophysical parameter space to-date, with 10 free parameters, spanning both population II and III star formation. We find X-ray parameters for the first galaxies could be measured to sub-percent precision, and, though they are highly degenerate, the stellar-to-halo mass relation and ionizing photon escape fraction for population II and III galaxies can be constrained to ~10% precision (logarithmic quantities). Using a principal component analysis we find HERA is most sensitive to the product of the ionizing escape fraction and the stellar-to-halo mass fraction for population II galaxies.

astro-ph.CO

Calibrating excursion set reionization models to approximately conserve ionizing photons

The excursion set reionization framework is widely used, due to its speed and accuracy in reproducing the 3D topology of reionization. However, it is known that it does not conserve photon number. Here, we introduce an efficient, on-the-fly recipe to approximately account for photon conservation. Using a flexible galaxy model shown to reproduce current high-$z$ observables, we quantify the bias in the inferred reionization history and galaxy properties resulting from the non-conservation of ionizing photons. Using a mock 21-cm observation, we perform inference with and without correcting for ionizing photon conservation. We find that ignoring photon conservation results in very modest biases in the inferred galaxy properties, for our fiducial model. The notable exception is in the power-law scaling of the ionizing escape fraction with halo mass, which can be biased from the true value by $\sim2.4σ$ (corresponding to $\sim0.2$ in the power-law index). Our scheme is implemented in the public code ${\tt 21cmFAST}$.

astro-ph.CO

21cmFAST v3: A Python-integrated C code forgenerating 3D realizations of the cosmic 21cm signal

This brief code paper presents a new Python-wrapped version of the popular 21cm cosmology simulator, 21cmFAST. The new version, v3+, maintains the same core functionality of previous versions of 21cmFAST, but features a simple and intuitive interface, and a great deal more flexibility. This evolution represents the work of a formalized collaboration, and the new version, available publicly on GitHub, provides a single point-of-reference for all future upgrades and community-added features. In this paper, we describe simple usage of 21cmFAST, some of its new features, and provide a simple performance benchmark.

astro-ph.IM

Learning to Summarize Long Texts with Memory Compression and Transfer

We introduce Mem2Mem, a memory-to-memory mechanism for hierarchical recurrent neural network based encoder decoder architectures and we explore its use for abstractive document summarization. Mem2Mem transfers "memories" via readable/writable external memory modules that augment both the encoder and decoder. Our memory regularization compresses an encoded input article into a more compact set of sentence representations. Most importantly, the memory compression step performs implicit extraction without labels, sidestepping issues with suboptimal ground-truth data and exposure bias of hybrid extractive-abstractive summarization techniques. By allowing the decoder to read/write over the encoded input memory, the model learns to read salient information about the input article while keeping track of what has been generated. Our Mem2Mem approach yields results that are competitive with state of the art transformer based summarization methods, but with 16 times fewer parameters

cs.CL

Reionization inference from the CMB optical depth and E-mode polarization power spectra

The Epoch of Reionization (EoR) depends on the complex astrophysics governing the birth and evolution of the first galaxies and structures in the intergalactic medium. EoR models rely on cosmic microwave background (CMB) observations, and in particular the large-scale E-mode polarization power spectra (EE PS), to help constrain their highly uncertain parameters. However, rather than directly forward-modelling the EE PS, most EoR models are constrained using a summary statistic -- the Thompson scattering optical depth, $τ_e$. Compressing CMB observations to $τ_e$ requires adopting a basis set for the EoR history. The common choice is the unphysical, redshift-symmetric hyperbolic tangent (Tanh) function, which differs in shape from physical EoR models based on hierarchical structure formation. Combining public EoR and CMB codes, 21cmFAST and CLASS, here we quantify how inference using the $τ_e$ summary statistic impacts the resulting constraints on galaxy properties and EoR histories. Using the last Planck 2018 data release, we show that the marginalized constraints on the EoR history are more sensitive to the choice of the basis set (Tanh vs physical model) than to the CMB likelihood statistic ($τ_e$ vs PS). For example, EoR histories implied by the growth of structure show a small tail of partial reionization extending to higher redshifts. However, biases in inference using $τ_e$ are negligible for the Planck 2018 data. Using EoR constraints from high-redshift observations including the quasar dark fraction, galaxy UV luminosity functions and CMB EE PS, our physical model recovers $τ_e=0.0569^{+0.0081}_{-0.0066}$.

astro-ph.CO

A tale of two sites -- II: Inferring the properties of minihalo-hosted galaxies with upcoming 21-cm interferometers

The first generation of galaxies is expected to form in minihalos, accreting gas through ${\rm H}_2$ cooling, and possessing unique properties. Although unlikely to be directly detected in UV/infrared surveys, the radiation from these molecular-cooling galaxies (MCGs) could leave an imprint in the 21-cm signal from the Cosmic Dawn. Here we quantify their detectability with upcoming radio interferometers. We generate mock 21-cm power spectra using a model for both MCGs as well as more massive, atomic-cooling galaxies (AGCs), allowing both populations to have different properties and scaling relations. The galaxy parameters are chosen so as to be consistent with: (i) high-redshift UV luminosity functions; (ii) the upper limit on the neutral fraction from QSO spectra; (iii) the Thomson scattering optical depth to the CMB; and (iv) the timing of the recent putative EDGES detection. The latter implies a significant contribution of MCGs to the Cosmic Dawn, if confirmed to be cosmological. We then perform Bayesian inference on two models including and ignoring MCG contributions. Comparing their Bayesian evidences, we find a strong preference for the model including MCGs, despite the fact that it has more free parameters. This suggests that if MCGs indeed play a significant role in the Cosmic Dawn, it should be possible to infer their properties from upcoming 21-cm power spectra. Our study illustrates how these observations can discriminate among uncertain galaxy formation models with varying complexities, by maximizing the Bayesian evidence.

astro-ph.CO

A tale of two sites -- I: Inferring the properties of minihalo-hosted galaxies from current observations

The very first galaxies that started the cosmic dawn likely resided in so-called "minihaloes", with masses of $\sim10^5$-$10^8\mathrm{M}_\odot$, accreting their gas from the intergalactic medium through H$_2$ cooling. Such molecularly cooled galaxies (MCGs) mostly formed in pristine environments, hosted massive, metal-free stars, and were eventually sterilized by the build-up of a disassociating (Lyman-Werner; LW) background. Therefore, their properties might be very different from the galaxies we see in the later Universe. Although MCGs are probably too faint to be observed directly, we could nevertheless infer their properties from the imprint they leave in the cosmic 21-cm signal. Here we quantify this imprint by extending the public simulation code 21cmFAST to allow for a distinct population of MCGs. We allow MCGs to have different properties from other galaxies, including unique scaling relations for their stellar-to-halo mass ratios, ionizing escape fractions, and spectral energy distributions. We track inhomogeneous recombinations, disassociative LW feedback, and photoheating from reionization. After demonstrating how MCGs can shape the 21-cm signal, we explore to what extent current observations can already place constraints on their properties. The cosmic microwave background optical depth from Planck sets an upper limit on the product of the ionizing escape fraction and the stellar mass in MCGs. When including also the timing of the putative EDGES absorption signal, we find an additional strong degeneracy between the stellar mass and the X-ray luminosity of MCGs. If proven to be of cosmic origin, the timing of the EDGES signal would have been set by MCGs.

astro-ph.CO

On the impressive performance of randomly weighted encoders in summarization tasks

In this work, we investigate the performance of untrained randomly initialized encoders in a general class of sequence to sequence models and compare their performance with that of fully-trained encoders on the task of abstractive summarization. We hypothesize that random projections of an input text have enough representational power to encode the hierarchical structure of sentences and semantics of documents. Using a trained decoder to produce abstractive text summaries, we empirically demonstrate that architectures with untrained randomly initialized encoders perform competitively with respect to the equivalent architectures with fully-trained encoders. We further find that the capacity of the encoder not only improves overall model generalization but also closes the performance gap between untrained randomly initialized and full-trained encoders. To our knowledge, it is the first time that general sequence to sequence models with attention are assessed for trained and randomly projected representations on abstractive summarization.

cs.CL

Combining high-z galaxy luminosity functions with Bayesian evidence

Galaxy formation during the first billion years of our Universe remains a challenging problem at the forefront of astrophysical cosmology. Although these $z \geq 6$ galaxies are likely responsible for the last major phase change of our Universe, the epoch of reionization (EoR), detailed studies are possible only for relatively rare, bright objects. Characterizing the fainter galaxies which are more representative of the population as a whole is currently done mainly through their non-ionizing UV luminosity function (LF). Observing the faint end of the UV LFs is nevertheless challenging, and current estimates can differ by orders of magnitude. Here we propose a methodology to combine disparate high-$z$ UV LF data sets in a Bayesian framework: Bayesian Data Averaging (BDA). Using a flexible, physically-motivated galaxy model, we compute the relative evidence of various $z=6$ UV LFs within the magnitude range $-20 \leq M_{\rm UV} \leq -15$ which is common to the data sets. Our model, based primarily on power-law scalings of the halo mass function, naturally penalizes systematically jagged data points as well as mis-estimated errors. We then use the relative evidence to weigh the posteriors obtained from disparate LF observations during the EoR, $6 \leq z \leq 10$. The resulting LFs suggest that the star formation rate density (SFRD) integrated down to a UV magnitude of -17 represent $60.9^{+11.3}_{-9.6}\%$ / $28.2^{+9.3}_{-10.1}\%$ / $5.7^{+4.5}_{-4.7}\%$ of the total SFRD at redshifts 6 / 10 / 15. The BDA framework we introduce enables galaxy models to leverage multiple, analogous observational data sets.

astro-ph.GA

Properties of reionization-era galaxies from JWST luminosity functions and 21-cm interferometry

Next generation observatories will enable us to study the first billion years of our Universe in unprecedented detail. Foremost among these are 21-cm interferometry with the HERA and the SKA, and high-$z$ galaxy observations with the James Webb Space Telescope (JWST). Taking a basic galaxy model, in which we allow the star formation rates and ionizing escape fractions to have a power-law dependence on halo mass with an exponential turnover below some threshold, we quantify how observations from these instruments can be used to constrain the astrophysics of high-$z$ galaxies. For this purpose, we generate mock JWST LFs, based on two different hydrodynamical cosmological simulations; these have intrinsic luminosity functions (LFs) which turn over at different scales and yet are fully consistent with present-day observations. We also generate mock 21-cm power spectrum observations, using 1000h observations with SKA1 and a moderate foreground model. Using only JWST data, we predict up to a factor of 2-3 improvement (compared with HST) in the fractional uncertainty of the star formation rate to halo mass relation and the scales at which the LFs peak (i.e. turnover). Most parameters regulating the UV galaxy properties can be constrained at the level of $\sim 10$% or better, if either (i) we are able to better characterize systematic lensing uncertainties than currently possible; or (ii) the intrinsic LFs peak at magnitudes brighter than $M_{\rm UV} \lesssim -13$. Otherwise, improvement over HST-based inference is modest. When combining with upcoming 21-cm observations, we are able to significantly mitigate degeneracies, and constrain all of our astrophysical parameters, even for our most pessimistic assumptions about upcoming JWST LFs. The 21-cm observations also result in an order of magnitude improvement in constraints on the EoR history.

astro-ph.CO

Ion Acceleration in Laser Generated Mega Tesla Magnetic Vortex

Magnetic Vortex Acceleration (MVA) from near critical density targets is one of the promising schemes of laser-driven ion acceleration. 3D particle-in-cell simulations are used to explore a more extensive laser-target parameter space than previously reported on in the literature as well as to study the laser pulse coupling to the target, the structure of the fields, and the properties of the accelerated ion beam in the MVA scheme. The efficiency of acceleration depends on the coupling of the laser energy to the self-generated channel in the target. The accelerated proton beams demonstrate high level of collimation with achromatic angular divergence, and carry a significant amount of charge. For PW-class lasers, this acceleration regime provides favorable scaling of maximum ion energy with laser power for optimized interaction parameters. The mega Tesla-level magnetic fields generated by the laser-driven co-axial plasma structure in the target are prerequisite for accelerating protons to the energy of several hundred MeV.

physics.plasm-ph

Inferring the astrophysics of reionization and cosmic dawn from galaxy luminosity functions and the 21-cm signal

The properties of the first galaxies, expected to drive the Cosmic Dawn (CD) and the Epoch of Reionization (EoR), are encoded in the 3D structure of the cosmic 21-cm signal. Parameter inference from upcoming 21-cm observations promises to revolutionize our understanding of these unseen galaxies. However, prior inference was done using models with several simplifying assumptions. Here we introduce a flexible, physically-motivated parametrization for high-$z$ galaxy properties, implementing it in the public code 21cmFAST. In particular, we allow their star formation rates and ionizing escape fraction to scale with the masses of their host dark matter halos, and directly compute inhomogeneous, sub-grid recombinations in the intergalactic medium. Combining current Hubble observations of the rest-frame UV luminosity function (UV LFs) at high-$z$ with a mock 1000h 21-cm observation using the Hydrogen Epoch of Reionization Arrays (HERA), we constrain the parameters of our model using a Monte Carlo Markov Chain sampler of 3D simulations, 21CMMC. We show that the amplitude and scaling of the stellar mass with halo mass is strongly constrained by LF observations, while the remaining galaxy properties are constrained mainly by 21-cm observations. The two data sets compliment each other quite well, mitigating degeneracies intrinsic to each observation. All eight of our astrophysical parameters are able to be constrained at the level of $\sim 10\%$ or better. The updated versions of 21cmFAST and 21CMMC used in this work are publicly available.

astro-ph.GA

Building a Neural Machine Translation System Using Only Synthetic Parallel Data

Recent works have shown that synthetic parallel data automatically generated by translation models can be effective for various neural machine translation (NMT) issues. In this study, we build NMT systems using only synthetic parallel data. As an efficient alternative to real parallel data, we also present a new type of synthetic parallel corpus. The proposed pseudo parallel data are distinct from previous works in that ground truth and synthetic examples are mixed on both sides of sentence pairs. Experiments on Czech-German and French-German translations demonstrate the efficacy of the proposed pseudo parallel corpus, which shows not only enhanced results for bidirectional translation tasks but also substantial improvement with the aid of a ground truth real parallel corpus.

cs.CL

Dark-ages reionization and galaxy formation simulation XI: Clustering and halo masses of high redshift galaxies

We investigate the clustering properties of Lyman-break galaxies (LBGs) at $z\sim6$ - $8$. Using the semi-analytical model {\scshape Meraxes} constructed as part of the Dark-ages Reionization And Galaxy-formation Observables from Numerical Simulation (DRAGONS) project, we predict the angular correlation function (ACF) of LBGs at $z\sim6$ - $8$. Overall, we find that the predicted ACFs are in good agreement with recent measurements at $z\sim 6$ and $z\sim 7.2$ from observations consisting of the Hubble eXtreme Deep Field (XDF), the Hubble Ultra-Deep Field (HUDF) and Cosmic Assembly Near-infrared Deep Extragalactic Legacy Survey (CANDELS) field. We confirm the dependence of clustering on luminosity, with more massive dark matter haloes hosting brighter galaxies, remains valid at high redshift. The predicted galaxy bias at fixed luminosity is found to increase with redshift, in agreement with observations. We find that LBGs of magnitude $M_{\rm AB(1600)} < -19.4$ at $6\lesssim z \lesssim 8$ reside in dark matter haloes of mean mass $\sim 10^{11.0}$- $10^{11.5} M_{\rm \odot}$, and this dark matter halo mass does not evolve significantly during reionisation.

astro-ph.GA