SearcharxivSearch

arXiv subjects

Aaron Smith

Publications and source records attributed to Aaron Smith.

At least 19 recordsLinked to original sources

RIGEL: Ultra-faint dwarf galaxy diversity shaped by inhomogeneous cosmic reionization

Ultra-faint dwarf galaxies (UFDs) are among the smallest and oldest galaxies in the Universe and are widely regarded as relics of cosmic reionization. To investigate how reionization quenches star formation and shapes the diversity of UFDs, we present a suite of eight cosmological zoom-in simulations of isolated UFDs with present-day halo masses of $\sim10^9\,{\rm M}_\odot$. The simulations are performed with the radiation-magnetohydrodynamic galaxy formation framework Realistic ISM modeling in Galaxy Evolution and Lifecycles (RIGEL), coupled to realistic large-scale radiation fields extracted from the THESAN reionization simulation. Despite residing in similar $z=0$ halos, the simulated galaxies span nearly two orders of magnitude in stellar mass and broadly reproduce the observed luminosities, sizes, metallicities, and stellar kinematics of Local Group UFDs. We find that reionization quenches star formation through a two-stage process. The arrival of the ionization front rapidly photoionizes the diffuse circumgalactic and intergalactic gas, suppressing further gas accretion onto the galaxy. Star formation nevertheless continues for several hundred Myr using the surviving self-shielded gas reservoir and ceases only after this gas is consumed or dispersed. Within 500 Myr after reionization, less than 40% of the initial gas mass remains in the halo, with photoevaporation constituting the dominant gas-loss channel. We further show that the halo mass at the time of reionization is a key parameter governing the subsequent evolution of UFDs. Galaxies residing in more massive halos at reionization retain gas for longer periods and undergo more extended chemical enrichment. Consequently, the halo mass at reionization strongly correlates with the final stellar mass, stellar age spread, and chemical evolution of the galaxy.

astro-ph.GA

The THESAN-ZOOM project: clumpiness of high-redshift galaxies and its connection to bursty star formation

Recent JWST observations have revealed diverse high-redshift galaxy morphologies, including a population with irregular and clumpy structures. The physical origin of these structures, and the extent to which observational biases shape their appearance, remain uncertain. We present a power-spectrum-based method for quantifying galaxy clumpiness across spatial scales, using the radiation-hydrodynamic simulation suite THESAN-ZOOM, which employs a state-of-the-art galaxy formation model that resolves the multiphase interstellar medium (ISM). Although the total stellar mass distributions in THESAN-ZOOM galaxies are usually smooth, clumpy structures appear in the H$\alpha$, far-ultraviolet (FUV), and optical light distributions. Tracers sensitive to shorter-timescale star formation exhibit more pronounced small-scale structure ($\sim10^{2}$--$10^{3}{\rm pc}$). The corresponding projected light spectra follow $P(k)\propto k^{-1}$ to $k^{-2}$, with progressively shallower slopes for tracers sensitive to more recent star formation, reflecting enhanced small-scale power and greater spatial intermittency in young stellar populations. This behaviour is consistent with a highly compressible, shock-dominated ISM in which stellar feedback and outflows reorganise dense gas into filamentary and clumpy structures. We also find that galaxy clumpiness depends on the treatment of stellar feedback. Weaker early stellar feedback enhances small-scale power in both the mass and light distributions. Clumpiness also varies strongly over the bursty star formation cycle, implying that observed samples may be biased towards galaxies caught in phases of elevated star formation. Galaxy clumpiness, therefore, could provide a complementary probe of the bursty star formation in the early Universe.

astro-ph.GA

On the Pseudo-Mixing of Kac's Walk

Motivated by a conjecture of Vaikuntanathan and Zamir, we study the pseudo-mixing of Kac's walk on $\mathrm{SO}(n)$: whether short trajectories are indistinguishable from Haar measure by low-complexity tests. We prove that the first $k$ columns mix in Wasserstein distance in $O(n(k+\log n)\log n)$ steps for fixed accuracy, resolving a conjecture of Oliveira. Combining this with a representation-theoretic variance bound, we show that if $T=\omega(nk(k+\log n)\log n)$, then every degree-$k$ polynomial normalized to have unit Haar variance has expectation under the $T$-step law within $o(1)$ of its Haar expectation. As an application, we show that this pseudo-mixing estimate can be used to prove the effectiveness of a fast Johnson--Lindenstrauss transform with the usual target dimension.

math.PR

Investigating Click Behaviors On Google Search Result Pages That Produce an AI Overview

In 2024, Google introduced "AI Overviews," a feature that displays an AI-generated result summary at the top of many Google search pages. This study investigates the role of AI in Google search using one month of web browsing data from a representative panel of 900 U.S. adults. Our analysis of the panelists' Google searches sheds light on AI Overviews, when they appear in Google search results, and what user behaviors are associated with AI Overviews. We identify several attributes that make a search query more likely to generate an AI Overview, including the length of a query, whether the query begins with a question word, and whether the query contains both a noun and verb. When it comes to user behavior, we find that clicks to sources cited in AI Overviews are very rare, occurring in only about 1% of visits to AI Overviews. We also find that AI Overviews are associated with fewer clicks and higher rates of ending browsing sessions. Importantly, results from a mixed-effects logistic regression model indicate that these associations hold when controlling for random effects by panelist and query attributes that make AI Overviews more likely to appear.

cs.HC

A large-scale corpus of religious radio broadcast transcripts from webstream recordings in the United States

Religious radio is a widespread but understudied form of mass communication in the United States, and content-level analysis of it has been constrained by the absence of large-scale transcript data. This Data Descriptor presents a corpus of transcribed English-language religious radio broadcasts captured from live webstreams over a one-month period in July 2025. Fifteen-minute segments were recorded on a rolling schedule from 785 distinct streams, which together rebroadcast the signals of more than two thousand AM and FM stations, yielding over 700,000 recordings and more than 60 million diarized lines of speech. Each recording was transcribed and speaker-diarized with an automated pipeline, and segmented and labeled by programming format and topic using a large language model. The corpus is organized as linked tables of stream metadata, recording metadata, and transcript lines. It supports descriptive study of religious broadcasting across regions and traditions, analysis of how social and political issues are discussed in religious media, and speech-processing research in an underrepresented domain.

cs.CL

Predicting ionised gas emission in 3D with SKIRT. I. Framework and validation

Emission lines from ionised gas are key diagnostics of star formation, metallicity, and ionisation conditions in galaxies. Interpreting spatially resolved observations from integral-field surveys (e.g. MaNGA, MUSE, JWST/NIRSpec) and comparing them with hydrodynamical simulations requires 3D photoionisation models that handle realistic geometries, dust attenuation, and synthetic instrument output. We present a new photoionisation module for the Monte Carlo radiative transfer code SKIRT that predicts emission-line luminosities of ionised gas in 3D, combining pre-computed Cloudy tables for gas temperature and opacity with a direct calculation of ion fractions and line emissivities. The local ionising radiation field (1-6 Ryd) is characterised by log U and four spectral-shape ratios; Cloudy tables map these to temperature and opacity, converging through SKIRT's existing iteration cycle. An inline solver then determines ion fractions from the converged field and temperature and evaluates line emissivities. We validate against Cloudy on 60 spherical shell models and against COLT on a Milky Way-analogue galaxy. On the 1D grid, hydrogen recombination lines agree with Cloudy to within a few per cent (Halpha median ratio 0.97) and the forbidden lines to within ~5%, except [S II] 6717 (1.23), whose offset traces a temperature overestimate near the ionisation front. In 3D, integrated luminosities agree with COLT to within 18% for the hydrogen lines and 2% for [N II], while [O III] and [S II] are elevated by ~70 and ~80%. Pixel-by-pixel correlation coefficients reach r >= 0.92, with luminosity-weighted scatter of 0.14-0.31 dex and broadly consistent BPT ratios. The module enables self-consistent synthetic observations in which ionised-gas emission lines, dust attenuation, and dust re-emission are computed in a single MCRT run, applicable to any hydrodynamical simulation.

astro-ph.GA

Force convergence in Monte Carlo Lyman-alpha radiative transfer

Monte Carlo radiative transfer (MCRT) is widely used to model Lyman-alpha (Lya) resonant-line transport, but convergence is difficult to assess in optically thick media where photons undergo many scatterings before escape. This is especially important for internal quantities such as radiative acceleration and the force multiplier, which depend on momentum deposition throughout the gas rather than only on emergent spectra. We study the convergence of Lya MCRT momentum-transfer estimators in static spherical clouds. We first establish diffusion-limit benchmarks for radial acceleration profiles and integrated force multipliers, then develop a moment-based framework for diagnosing convergence from the photon-packet contribution distribution. This framework separates three distinct questions: whether the estimator converges to the correct mean, how large its finite-sampling uncertainty is, and whether the estimated uncertainty is itself stable. We apply this hierarchy to the direct event-based scattering estimator, a gradient-of-energy-density estimator, and a divergence-of-radiation-pressure estimator. Zeroth-order convergence is assessed with profile comparisons, integrated force-multiplier bias, and finite-group relative error. First-order convergence is quantified with fractional error, the photon number required to reach a target precision, and the corresponding runtime requirement. Second-order convergence is tested with the coefficient of variation of variance, which measures the reliability of the variance estimate used in the first-order diagnostics. Core-skipping prescriptions, source geometry, estimator construction, and spatial resolution enter this hierarchy in different ways. Our results provide a practical convergence framework for internal Lya MCRT force calculations and show why statistical precision, computational cost, and physical accuracy must be evaluated separately.

astro-ph.GA

The Thesan-Zoom project: bursty star formation is incompatible with prolonged dust survival

Cosmic dust is a key regulator of galaxy evolution, but its build-up and survival in the first billion years remain poorly constrained. We present a systematic analysis of dust in the thesan-zoom suite of radiation-hydrodynamical zoom-in simulations, which self-consistently model dust formation, growth, destruction, and its coupling to radiative transfer in galaxies at $z \geq 3$, a multi-phase ISM and bursty star formation histories. The simulated galaxies reproduce the observed trends of dust-to-gas and dust-to-metal ratios with gas metallicity, while showing a dust deficit at high specific star-formation rates. They also broadly match observed dust temperatures and UV-IR spatial offsets. We find that dust and its properties are strongly time-variable and tightly linked to bursty star formation, with short-lived IR-bright phases (median duration of $20.3^{+2.3}_{-2.4}$ Myr) and longer dust-poor phases, naturally producing a correlation between dust temperature and distance from the star-forming main sequence. The predicted attenuation at $1500$ \r{A} is low compared to observations, even when including unresolved dust through post processing, indicating that a mechanism able to shield dust from strong feedback events is necessary to reconcile our galaxy formation model with observations. In our model, bursty star formation prevents the survival of large dust reservoirs ($M_{dust} / M_{star} \geq 10^{-3}$) over a significant fraction of cosmic time. This implies that bursty star formation can produce the observed overabundance of UV-bright galaxies at $z \geq 10$ only if it rapidly settles down by $z \sim 8$ (where large dust reservoirs are detected). It is also possible that our models lack physical ingredients or emergent phenomena that aid the survival of dust. Future observations of high-redshift dust will be key to diagnose the physical mechanism at play in the first galaxies.

astro-ph.GA

Globular cluster abundance patterns inherited from giant molecular clouds

Globular clusters exhibit large star-to-star variations and anticorrelations in their light element abundances that are commonly interpreted in terms of in-cluster self-enrichment, in which ejecta from early-forming cluster stars pollute the gas from which later stars form over millions of years. Yet proposed self-enrichment scenarios suffer from a severe mass-budget problem or invoke exotic stellar populations. Using cosmological radiation-hydrodynamic simulations with a standard chemical enrichment model, we identify a population of giant molecular clouds whose internal abundance patterns reproduce several key globular cluster signatures: large light-element abundance spreads and nitrogen-oxygen anticorrelations at nearly constant iron abundance. These clouds form at the restart of star-formation activity after an earlier starburst, where previously ejected oxygen-rich gas collides with nitrogen-rich galactic gas, and are sites of dense star-cluster formation. In this picture, the chemical abundance patterns of globular clusters need not require extended in-cluster star formation, but can be inherited at birth from chemically structured interstellar gas shaped by the baryon cycle. Globular clusters therefore provide a fossil record of chemical enrichment and gas flows in high-redshift galaxies.

astro-ph.GA

The Lumina Project: Intergalactic Clumping and Recombination Sinks

Recombinations during the Epoch of Reionization are intrinsically inhomogeneous, with different regions of the intergalactic medium contributing unevenly depending on their density, temperature, ionization state, and spatial patchiness. We combine the high- and medium-resolution 95.5 cMpc Thesan-1 andh Thesan-2 runs with the significantly larger 500 cMpc Lumina simulation to measure clumping factors and recombination rates consistently across different resolutions and box sizes. We consider the standard ionized hydrogen clumping factor, $C_{\rm HII} \equiv \langle n_{\rm HII}^2\rangle/\langle n_{\rm HII}\rangle^2$, and a recombination-weighted clumping factor, $C_{\rm rec}$. Despite differences in resolution, volume, and reionization history, the simulations show an approximately universal clumping evolution at the 10-20% level when parametrized by the global ionized fraction $x_{\rm HII}$ rather than by redshift. Across all simulations, $C_{\rm rec}$ remains systematically below $C_{\rm HII}$, with the discrepancy increasing toward lower redshift as photoheating suppresses recombinations. In \lumina, the density-only prescription overpredicts the instantaneous recombination rate by factors of 1.29 at $z\approx8$ and 1.84 at $z\approx5$, and the cumulative recombination count by a factor of 1.45 by $z\approx5$. Mapping the recombination budget in the joint overdensity-temperature plane reveals that the dominant recombination ridges closely follow simple analytic thermal equilibrium bands. Finally, we introduce a phase-space recombination integral and define a phase-space clumping factor, $C_{\rm ps}(\Delta,T)$, which isolates the intrinsic recombination enhancement associated with ionization structure and thermal state at fixed overdensity and temperature.

astro-ph.GA

From THESAN-ZOOM to JWST: Predicting ionizing photon escape and the rise of UV-bright reionization sources

Understanding the sources and evolution of cosmic reionization remains a central challenge in astrophysics, with the escape of ionizing Lyman-continuum (LyC) photons from early galaxies representing a major uncertainty. In this work, we use more than 35,000 galaxy realisations from the THESAN-ZOOM cosmological radiation-hydrodynamic simulations to identify indirect diagnostics of the LyC photon escape fraction ($f_\mathrm{esc}$) and the LyC photon escape rate ($\dot{N}_\mathrm{ion,esc}$) across the redshift range $z=3-16$. We train random forest regression models using these diagnostics to predict both quantities. We present four models: two trained with the full set of simulation-derived indicators to predict $f_\mathrm{esc}$ and $\dot{N}_\mathrm{ion,esc}$, and two restricted to observables accessible to JWST photometric surveys. We find the 10-to-100$\,$Myr star-formation rate ratio ($\mathrm{SFR}_{10} / \mathrm{SFR}_{100}$) and the gas-to-stellar mass ratio ($M_\mathrm{gas} / M_*$) to be the strongest diagnostics of $f_\mathrm{esc}$, suggesting a strong relationship between ionizing photon escape and gas clearing through bursty star formation. In contrast, rest-frame UV ($1500 \, \r{A}$) absolute magnitude ($M_\mathrm{UV}$) dominates $\dot{N}_\mathrm{ion,esc}$ prediction. Motivated by the strong predictive power of $M_\mathrm{UV}$, we combine observed UV luminosity functions with derived $\dot{N}_\mathrm{ion,esc} - M_\mathrm{UV}$ relations to construct histories of reionization. These are consistent with observational constraints, avoiding the recently reported crisis in the ionizing photon budget. Our analysis suggests that the bulk of reionization occurred rapidly after $z \approx 8$, driven by UV-bright galaxies, with the $M_\mathrm{UV} < -17$ populations providing the dominant contribution.

astro-ph.GA

Lyman-alpha Pressure Strongly Enhances Pre-Supernova Feedback at Cosmic Dawn: The First Multi-Dimensional Lyman-alpha Radiation Hydrodynamics Simulations

The dynamical role of Lyman-$\alpha$ (Ly$\alpha$) radiation pressure feedback has been debated for nearly a century, with recent analytical and 1D numerical studies highlighting its potential dominance over other stellar feedback processes at Cosmic Dawn. Despite this, no multi-dimensional Ly$\alpha$ radiation hydrodynamics (RHD) simulations have been performed to date. In this paper, we present the first 2D Ly$\alpha$ RHD simulations using Lydion, an RHD code with a novel M1 moment method for Ly$\alpha$ transfer, and self-consistent dust dynamics. Lydion yields a $\sim \mathcal{O}(100) \,\times$ speed-up compared to Monte Carlo radiative transfer in simple benchmarks, making 2D Ly$\alpha$ RHD feasible. We perform simulations of star clusters and isolated stars embedded in dense, metal-poor ($Z/Z_\odot \leq 0.01$) clouds, and find that Ly$\alpha$ feedback dramatically boosts outflows and dominates over feedback from direct and infrared radiation pressure. Ly$\alpha$ leakage through lower-column density channels, Doppler shifts, and Ly$\alpha$ photon destruction, while important, cannot prevent the build-up of strong Ly$\alpha$ radiation pressure in H II regions, leading to radiative forces $\sim (2 - 16) \times L_{\rm bol}/c$, and Ly$\alpha$ force multipliers $M_{\rm F} \sim 10-60$. Ly$\alpha$ feedback may not preclude efficient star formation, but raises the threshold gas surface density for this to occur. We conclude that nearly all galaxy and star formation simulations are currently missing the strongest source of radiation pressure feedback in dense and metal-poor environments.

astro-ph.GA

Fast approximate Bayesian multidimensional scaling with consistency guarantees

Bayesian multidimensional scaling (BMDS) embeds $n$ objects in a low-dimensional space to approximately preserve an observed dissimilarity matrix. Compared to classic MDS, BMDS is more robust to model misspecification and supports posterior uncertainty quantification and joint estimation within hierarchical models. However, standard BMDS inference is computationally prohibitive, requiring $O(n^2)$ operations per MCMC iteration to evaluate the likelihood. We propose Barnes--Hut BMDS (BH-BMDS), which uses a tree-based approximation to the likelihood and a Gibbs sampler that leverages this structure, remaining compatible with hierarchical extensions. BH-BMDS reduces computational complexity to $O(n \log n)$ while preserving the geometric fidelity of the embedding. We further establish consistency for the stationary measure of BH-BMDS, proving that it concentrates around the true latent configuration even as the total error of the surrogate likelihood diverges. Notably, this consistency holds in the infinite-dimensional limit. We evaluate the approximation on datasets with diverse structure, including air traffic networks, arXiv abstracts, MNIST images and neural activity recordings from mouse models of tau pathology. Across all settings, BH-BMDS closely matches BMDS while achieving substantial computational gains, with approximately 10-fold speedups at $n=1{,}000$ and 70-fold speedups at $n=10{,}000$. These gains increase with $n$, demonstrating strong empirical scalability.

stat.CO

Accurate and Efficient MCMC for Latent Position Models

Latent position models (LPMs) are a large and popular class of models for random graphs. However, fitting Bayesian LPMs is computationally challenging - computing the likelihood even once takes time that is quadratic in the number of vertices $|V|$ of the observed graph $G = (V,E)$. Many previous papers have introduced approximate MCMC algorithms to speed this up, with the most similar to ours, Rastelli et al (2024), presenting an algorithm that has amortized running time that can be reduced almost to $O(|E|)$ and good empirical performance on reasonable inference problems. The present paper offers two algorithms for solving the same problem: a ``fast" algorithm with running time of the same almost-$O(|E|)$ order as astelli et al and much stronger accuracy guarantees, and a ``faster" algorithm with an improved running time of almost $O(|V|)$, and accuracy guarantees that are slightly improved compared to Rastelli et al (but not sufficient for all tasks). The main improvements come from the introduction of a simple auxiliary data structure that can be cheaply updated during an MCMC run; we suspect that the same ``cheap sketch" may be useful for other MCMC algorithms.

stat.CO

Convergence Rates of Ordering, Testing and Estimation Procedures for Graphons With Fast Boundary Decay Rates

In latent-position random graph models (LPMs), latent vertex positions $U_{1},\ldots,U_{n}$ are sampled from some distribution on a latent space $\Omega$, then edges of an observed graph $G = ([n],E)$ are sampled with some probability $\mathbb{P}[(i,j) \in E ]=w(U_i,U_j)$ that depends on the unobserved latent positions. LPMs are ubiquitous in the statistical analysis of networks, offering models that have good empirical performance, strong theoretical guarantees, and tractable algorithms. The special case $\Omega = [0,1]$ is important, as it corresponds to graphs with temporal or preference-based structure. In this paper, we study three problems related to LPMs with latent space $[0,1]$: \textit{ordering} the vertices according to the latent positions, \textit{estimating} the generating graphon $w$, and \textit{testing} whether an observed graph $G$ could have come from an LPM with state space $[0,1]$. Our results on the ordering problem greatly generalize two observations of Janssen/Smith (2022): (i) for \textit{some} families of graphons, the best estimate of the ordering converges much faster than the usual statistical rate of $\frac{1}{\sqrt{n}}$, and (ii) this occurs even though, for the same families of graphons, the best estimate of the latent positions still occurs at the usual $\frac{1}{\sqrt{n}}$ rate. As a main consequence, we develop a computationally-efficient graphon-estimation algorithm and show that it has the same convergence rate as the non-explicit optimal algorithm of Gao et al (2015). We also derive and analyze a testing procedure.

math.ST

Mixing on $k$ Columns of the Transvection Walk

In Diaconis and Saloff-Coste (1996), the authors introduced the simple ``transvection" walk on $\mathrm{GL}_n(\mathbb F_2)$: at each step, choose two distinct rows and add one to the other. In Ben-Hamou (2025), the author recently proved that this walk has mixing time $O(n^2\log n)$. Inspired by applications in cryptography (see Sotiraki (2016)), Ben-Hamou and Peres (2018) conjectured that the first $k$ columns of this walk mixed in $O(nk \log(n))$ steps. Our main result is a proof of this conjecture uniformly in $n$ and $k.$ Our proof is based on a local-to-global entropy estimate, in the spirit of block factorization results such as Caputo et al (2015), Caputo et al (2021). In our setting, the kernels that correspond roughly to the block kernels of Caputo et al (2021) do not have uniformly large log-Sobolev constants, and so naively applying these techniques does not improve over Ben-Hamou (2025). We avoid these bad blocks by combining our entropy estimates with a burn-in argument similar to classical drift-and-minorization arguments of Rosenthal (1995). This method may be of broader interest, and so we illustrate it by proving an analogous result for a family of product-replacement algorithms on the Heisenberg group.

math.PR

The Lumina Project: The Demographics of Active Galactic Nuclei from Quasars to Little Red Dots at $z\geq 3$

High-redshift active galactic nuclei (AGN) serve as powerful probes of early black-hole growth, galaxy formation, and the evolving intergalactic medium (IGM). In this work, we use Lumina, a cosmological radiation-hydrodynamic simulation spanning the epochs of hydrogen and helium reionization, which combines a large $(500\,{\rm cMpc})^3$ volume with $2\times 6000^3$ resolution elements, to explore high-redshift AGN. The simulation self-consistently follows hundreds of millions of galaxies and supermassive black holes (SMBHs), together with their impact on the ionization and thermal state of the IGM. We exploit this uniquely large dynamic range to predict multi-band AGN luminosity functions (LFs) at $z \geq 3$, from hard X-rays to the mid-infrared. These predictions encompass both moderately luminous quasars and the faint ``Little Red Dots'' (LRDs) uncovered by JWST. We develop an empirical model that maps simulated SMBHs onto observed AGN using bolometric and extinction/absorption corrections for canonical AGN and LRDs, and in which SMBHs with $M_{\rm BH}\leq 10\,M_{\rm seed} \sim 10^{7}\,{\rm M}_{\odot}$ stay in the LRD phase with a duty cycle of $30\%$. This simple framework reproduces the observed LFs and clustering of LRDs. Meanwhile, the pre-JWST quasar LF constraints are recovered, although we find that a $\sim 0.3$ dex log-normal scatter in bolometric luminosity is required to reproduce the bright end. We place the simulated AGN population in the cosmological context by quantifying the redshift evolution of AGN and LRD number densities, and their contributions to the integrated BH mass densities. The same AGN population is the dominant driver for the HeII reionization modelled self-consistently in Lumina. This empirical AGN model paves the way for general population-synthesis models of high-redshift AGN, including LRDs, in a unified cosmological framework.

astro-ph.GA

The Lumina Project: CMB Optical Depth Fluctuations from Patchy Reionization

Patchy reionization couples the ionized-bubble morphology to the underlying density field, making the CMB Thomson optical depth sensitive to both the global ionization history and anisotropic fluctuations on the sky. Using the large-volume radiation-hydrodynamical Lumina simulation, we compute $\tau_{\rm CMB}$ in two ways: (i) from global volume- and mass-weighted ionization histories, and (ii) from explicit line-of-sight integrations through on-the-fly light cones. We find that the sightline-averaged optical depth in the light cone, $\langle \tau_{\rm LOS} \rangle = 0.0550$, exceeds the value inferred from a global volume-weighted history, $\tau_{{\rm CMB},V} = 0.0515$, by $\approx 7\%$. This enhancement is largely captured by the global mass-weighted prediction, $\tau_{{\rm CMB},m} = 0.0544$, indicating that precision comparisons to CMB optical-depth constraints should use mass-weighted electron fractions or explicit light-cone integration rather than volume-weighted ionized fractions alone. The excess optical depth accumulates primarily near $z_{\rm LOS} = 8.0^{+1.9}_{-1.3}$, where the combination of high physical density and strong ionization-field patchiness is greatest. The resulting $\tau_{\rm LOS}$ field is non-Gaussian and exhibits $\gtrsim 5\%$ sightline-to-sightline scatter, with fluctuations tracing rare early-ionized overdensities and large-scale structure. Coarse-graining experiments show that smoothing the ionization field on $\gtrsim 3 {\rm cMpc}$ scales suppresses the density-ionization correlation and biases $\tau_{\rm CMB}$ low relative to the resolved calculation. Finally, angular power spectra and real-space correlation functions decomposed into HII, HeII, and HeIII auto- and cross-contributions reveal scale-dependent departures from simple hydrogen-helium co-tracing and evolving characteristic scales with redshift.

astro-ph.CO