SearcharxivSearch

arXiv subjects

Zechang Sun

Publications and source records attributed to Zechang Sun.

At least 19 recordsLinked to original sources

A Unified Generative Framework for Scalable Chemical Reaction Network Exploration

Chemical reaction networks (CRNs) are crucial for understanding reaction mechanisms and guiding chemical synthesis, yet the computational exploration remains limited by the combinatorial growth of chemical space, the reliability of reaction path screening, and the cost of evaluating thermodynamic and kinetic properties. Here, we present ByteCRN, an end-to-end framework for computational CRN exploration that combines chemically informed reaction enumeration with generative transition state modeling. A key component of our framework is a generative rectified flow architecture for both transition state generation and reaction validation, where it maps reactant-product pairs to candidate transition state structures and verifies connectivity by mapping back to reactants and products. This unified generative strategy replaces the most expensive steps of conventional computational workflows, namely iterative transition state search and intrinsic reaction coordinate validation, within a complete CRN construction pipeline. ByteCRN delivers a 10--100-fold acceleration over traditional workflows while maintaining high predictive fidelity for individual reactions. At the network scale, it effectively prunes $\sim$70-90% of the enumerated reactions, streamlining the exploration of complex reaction space. Its utility is illustrated through the discovery of novel pathways involving cyanoacetaldehyde and the successful modeling of the challenging $\gamma$-ketohydroperoxide network, demonstrating a practical, scalable approach to autonomous chemical exploration.

physics.chem-ph

(LRDs)$^2$: The Low-ReDshift Little Red Dots Survey. II. DESI DR1 Sample

JWST has revealed a substantial population of "Little Red Dots" (LRDs) at $z>4$, challenging conventional AGN frameworks. However, the low-redshift regime remains largely unexplored. In the second paper of the (LRDs)$^2$ series, we present a systematic selection from DESI DR1 and identify 27 LRDs at $z=0.2-0.9$, yielding a number density lower limit of $7.5 \times 10^{-10}$ cMpc$^{-3}$. We conducted near-IR spectroscopic follow-up observations for 18 of them, revealing their full SED shapes and emission lines. These low-$z$ LRDs share the hallmark properties of their high-$z$ counterparts: compact morphology, V-shaped UV-optical continua, broad Balmer emission with extreme decrements (median H$\alpha$/H$\beta \sim 16$), frequent Balmer absorption (67%), and blackbody-like optical-to-near-IR continua. All have low metallicity, occupy the same regions in the BPT diagram as high-$z$ LRDs, and have softer ionizing spectra than typical AGNs. The consistency between low-$z$ and high-$z$ LRD properties indicates the same physical processes at work. The correlation between broad-line Balmer luminosity and $L_{5100}$ deviates from that of local type-1 AGNs, limiting the direct application of local BH mass calibrations. Ionized [O III] outflows are ubiquitous (78%). One LRD at $z=0.196$, J1717+3807, shows robust long-term variability in $i$ and WISE bands. The optical-to-NIR continua of LRDs reveal a wide range of temperatures $\sim 2000-4700$ K (peak $0.6-1.5$ $\mu$m), with a subset showing cooler and larger envelopes than those at high $z$. Low-$z$ LRDs serve not only as proximate laboratories for probing the nature of LRDs, but also trace the cosmic evolution of this population from the cosmic dawn to the present day.

astro-ph.GA

A Post-starburst Galaxy Undergoing Ram-pressure Stripping at Redshift 3.06

Understanding how galaxies ignite and extinguish their star formation remains a cornerstone question in modern astrophysics. Recent JWST surveys have revealed an overabundance of massive quiescent galaxies in the first billion years of the Universe, challenging current models of galaxy evolution. In the nearby Universe, ram pressure stripping (RPS) is a major environmental mechanism capable of rapidly shutting down star formation, yet direct observation remains scarce at redshift $z\gtrsim1$, and its role at $z>2$ is even poorly constrained by simulations. Here, we utilize JWST and ALMA observations to present direct evidence of RPS in the post-starburst galaxy A2744-JF-z3, residing in a galaxy group at redshift 3.06, the earliest such detection to date. Spectroscopic diagnostics and spectral energy distribution modeling reveal the ongoing removal of cold gas and dust, coincident with the abrupt cessation of star formation. Contrary to hydrodynamical simulations that predict a reduced incidence of RPS at high redshift, our results instead imply that RPS can operate at $z>3$, suggesting a highly stochastic and impulsive stripping within a clumpy, filamentary intra-group and circumgalactic medium. These observations extend environmental quenching well into the epoch of galaxy assembly, highlighting RPS as a previously overlooked decisive pathway to rapid quenching in nascent groups and protoclusters in the early Universe.

astro-ph.GA

Cosmological Constraints from Full-Scale Clustering and Galaxy-Galaxy Lensing with DESI DR1

We present constraints on cosmic structure growth from the analysis of galaxy clustering and galaxy-galaxy lensing with galaxies from the Dark Energy Spectroscopic Instrument (DESI) Data Release 1. We analyze four samples drawn from the Bright Galaxy Survey (BGS) and the Luminous Red Galaxy (LRG) target classes. Projected galaxy clustering measurements from DESI are supplemented with lensing measurements from the Dark Energy Survey (DES), the Kilo-Degree Survey (KiDS), and the Hyper Suprime-Cam (HSC) survey around the same targets. Our method relies on a simulation-based modeling framework using the AbacusSummit simulations and a complex halo occupation distribution model that incorporates assembly bias. We analyze scales down to $0.4 \, h^{-1} \, \mathrm{Mpc}$ for clustering and $2.5 \, h^{-1} \, \mathrm{Mpc}$ for lensing, leading to stringent constraints on $S_8 = \sigma_8 \sqrt{\Omega_\mathrm{m} / 0.3}$ and $\Omega_\mathrm{m}$ when fixing other cosmological parameters to those preferred by the CMB. We find $S_8 = 0.794 \pm 0.023$ and $\Omega_\mathrm{m} = 0.295 \pm 0.012$ when using lensing measurements from DES and KiDS. Similarly, for HSC, we find $S_8 = 0.793 \pm 0.017$ and $\Omega_\mathrm{m} = 0.303 \pm 0.010$ when assuming the best-fit photometric redshift offset suggested by the HSC collaboration. Overall, our results are in good agreement with other results in the literature while continuing to highlight the constraining power of non-linear scales.

astro-ph.CO

AstroMLab 5: Structured Summaries and Concept Extraction for 400,000 Astrophysics Papers

We present a dataset of 408,590 astrophysics papers from arXiv (astro-ph), spanning 1992 through July 2025. Each paper has been processed through a multi-stage pipeline to produce: (1) structured summaries organized into six semantic sections (Background, Motivation, Methodology, Results, Interpretation, Implication), and (2) concept extraction yielding 9,999 unique concepts with detailed descriptions. The dataset contains 3.8 million paper-concept associations and includes semantic embeddings for all concepts. Comparison with traditional ADS keywords reveals that the concepts provide denser coverage and more uniform distribution, while analysis of embedding space structure demonstrates that concepts are semantically dispersed within papers-enabling discovery through multiple diverse entry points. Concept vocabulary and embeddings are publicly released at https://github.com/tingyuansen/astro-ph_knowledge_graph.

astro-ph.IM

A Unified Photometric Redshift Calibration for Weak Lensing Surveys using the Dark Energy Spectroscopic Instrument

The effective redshift distribution $n(z)$ of galaxies is a critical component in the study of weak gravitational lensing. Here, we introduce a new method for determining $n(z)$ for weak lensing surveys based on high-quality redshifts and neural network-based importance weights. Additionally, we present the first unified photometric redshift calibration of the three leading stage-III weak lensing surveys, the Dark Energy Survey (DES), the Hyper Suprime-Cam (HSC) survey and the Kilo-Degree Survey (KiDS), with state-of-the-art spectroscopic data from the Dark Energy Spectroscopic Instrument (DESI). We verify our method using a new, data-driven approach and obtain $n(z)$ constraints with statistical uncertainties of order $\sigma_{\bar z} \sim 0.01$ and smaller. Our analysis is largely independent of previous photometric redshift calibrations and, thus, provides an important cross-check in light of recent cosmological tensions. Overall, we find excellent agreement with previously published results on the DES Y3 and HSC Y1 data sets while there are some differences on the mean redshift with respect to the previously published KiDS-1000 results. We attribute the latter to mismatches in photometric noise properties in the COSMOS field compared to the wider KiDS SOM-gold catalog. At the same time, the new $n(z)$ estimates for KiDS do not significantly change estimates of cosmic structure growth from cosmic shear. Finally, we discuss how our method can be applied to future weak lensing calibrations with DESI data.

astro-ph.CO

Mephisto: Self-Improving Large Language Model-Based Agents for Automated Interpretation of Multi-band Galaxy Observations

Astronomical research has long relied on human expertise to interpret complex data and formulate scientific hypotheses. In this study, we introduce Mephisto -- a multi-agent collaboration framework powered by large language models (LLMs) that emulates human-like reasoning for analyzing multi-band galaxy observations. Mephisto interfaces with the CIGALE codebase (a library of spectral energy distribution, SED, models) to iteratively refine physical models against observational data. It conducts deliberate reasoning via tree search, accumulates knowledge through self-play, and dynamically updates its knowledge base. Validated across diverse galaxy populations -- including the James Webb Space Telescope's recently discovered "Little Red Dot" galaxies -- we show that Mephisto demonstrates proficiency in inferring the physical properties of galaxies from multi-band photometry, positioning it as a promising research copilot for astronomers. Unlike prior black-box machine learning approaches in astronomy, Mephisto offers a transparent, human-aligned reasoning process that integrates seamlessly with existing research practices. This work underscores the possibility of LLM-driven agent-based research for astronomy, establishes a foundation for fully automated, end-to-end artificial intelligence (AI)-powered scientific workflows, and unlocks new avenues for AI-augmented discoveries in astronomy.

astro-ph.IM

Teaching LLMs to Speak Spectroscopy

Pre-trained Large Language Models (LLMs) have revolutionized text processing, yet adapting Transformer-based neural networks to non-textual scientific modalities typically requires specialized architectures and extensive computational resources. We demonstrate that LLaMA-3.1-8B can be efficiently repurposed to predict galaxy redshifts from spectroscopic data through Low-Rank Adaptation (LoRA), achieving competitive performance while preserving its linguistic capabilities. Using only 16 GPU-hours and adapting 0.04% of model parameters, our approach achieves a mean absolute error of 0.04 in redshift prediction while retaining over 85% of performance on AstroBench and 89% on general QA tasks from eval-harness. This minimal-effort adaptation--requiring only simple standard fine-tuning APIs--lowers barriers to entry for domain scientists and enables integrated agentic workflows where a single model handles both spectroscopic data for quantitative analysis and natural language for reasoning.

astro-ph.IM

Can AI Dream of Unseen Galaxies? Conditional Diffusion Model for Galaxy Morphology Augmentation

Observational astronomy relies on visual feature identification to detect critical astrophysical phenomena. While machine learning (ML) increasingly automates this process, models often struggle with generalization in large-scale surveys due to the limited representativeness of labeled datasets, whether from simulations or human annotation, a challenge pronounced for rare yet scientifically valuable objects. To address this, we propose a conditional diffusion model to synthesize realistic galaxy images for augmenting ML training data (hereafter GalaxySD). Leveraging the Galaxy Zoo 2 dataset which contains visual feature, galaxy image pairs from volunteer annotation, we demonstrate that GalaxySD generates diverse, high-fidelity galaxy images that closely adhere to the specified morphological feature conditions. Moreover, this model enables generative extrapolation to project well-annotated data into unseen domains and advancing rare object detection. Integrating synthesized images into ML pipelines improves performance in standard morphology classification, boosting completeness and purity by up to 30% across key metrics. For rare object detection, using early-type galaxies with prominent dust lane features (~0.1% in GZ2 dataset) as a test case, our approach doubled the number of detected instances, from 352 to 872, compared to previous studies based on visual inspection. This study highlights the power of generative models to bridge gaps between scarce labeled data and the vast, uncharted parameter space of observational astronomy and sheds insight for future astrophysical foundation model developments. Our project homepage is available at https://galaxysd-webpage.streamlit.app/.

astro-ph.GA

Zangetsu: A Candidate of Isolated, Quiescent, and Backsplash Ultra-Diffuse Galaxy in the COSMOS Field

Deep imaging surveys have changed our view of the low surface brightness (LSB) Universe. The ``renaissance'' of the LSB galaxy population, as a prime example of this recent development, continues to challenge our understanding of galaxy formation. Here, we report the serendipitous discovery of Zangetsu, an isolated, quiescent, and distorted ultra-diffuse galaxy (UDG) candidate in the COSMOS field, using images from the Hyper Suprime-Cam Subaru Strategic Program (HSC-SSP). Zangetsu exhibits an extremely low central surface brightness ($\mathrm{\mu_{0,g}}=26.60\pm0.01$ mag arcsec$^{-2}$), a very shallow inner surface brightness profile ($\mathrm{n}_{\rm Sersic}=0.40\pm0.01$), and a large angular size ($\mathrm{R_e}\approx 10.44$ arcsec). Surprisingly, Zangetsu also has a quiescent stellar population ($\mathrm{g-i}=0.96$), an unusually elongated shape ($\mathrm{b/a}\sim 0.25$), and mild morphological asymmetry, making it a rare case among known UDGs. Surface brightness fluctuation analysis of HSC and Hubble Space Telescope (HST) images only provides a distance lower limit of $D>25.4$ Mpc (thus $\mathrm{R_e}>1.38$ kpc). However, Zangetsu remains an extreme outlier in the luminosity-size relation of known LSB galaxies, suggesting that it could be an exceptionally large and/or diffuse system. Classic internal or external UDG formation mechanisms alone struggle to explain such a system. A backsplash origin may account for its isolation and quiescent nature. This finding also raises the possibility that current works may overlook similarly extreme, elongated systems that could further our understanding of the LSB Universe.

astro-ph.GA

Luminous Mid-IR Selected Type-2 Quasars at Cosmic Noon in SDSS Stripe82 I: Selection, Composite Photometry, and Spectral Energy Distributions

We analyze 23 spectroscopically confirmed Type-2 quasars (QSOs) selected from the WISE 22$\rm \mu$m band in the SDSS Stripe 82 region, focusing on their multi-band photometry and spectral energy distributions (SEDs). These objects were selected to be IR-luminous ($\rm flux_{W4} > 5mJy$, i.e., $12.62 < W4 < 14.62 \rm\ AB \, magnitude$), optically faint ($r > 23$) or with red color ($r - W4 >8.38$). Gemini/GNIRS observations were conducted for all 24 candidates, and 18/24 were also observed with Keck/LRIS. The observations confirm 23 to be real Type-2 QSOs in the redshift range $0.88 - 2.99$ (12 are at $z>2$). We collect multi-band photometry and conduct SED fitting. The composite photometry probes the wavelength from 0.1$\rm \mu$m to 10$\rm \mu$m at the rest frame. The IR emission is dominated by dust torus implying an average torus luminosity for the sample of $L_{\rm torus} 10^{46.84} \rm erg/s$. The origin of the rest-UV/optical light is not definitive, but we present three possible scenarios: scattered light, stellar emission, and the reddened accretion disk. Assuming an obscured:unobscured ratio of approximately 1:1, our targets have $L_{\rm bol} = 10^{46.28} \rm erg \,s^{-1} - 10^{47.49} \rm erg \,s^{-1}$ and around SMBH masses $\rm 10^{8.18} M_{\odot} - 10^{9.39} M_{\odot}$, assuming they accreate at the Eddington limit. Compared to previous Type-2 AGN SEDs, our targets have a brighter dust torus and redder optical-IR color. By comparing the SED to the results from JWST `little red dots' (LRDs), we find that these IR-selected Type-2 QSOs have similar SED shapes to the LRDs. This pilot Type-2 QSO survey demonstrates that mid-IR selection is an efficient way to find luminous Type-2 QSOs at $z>2$. Finally, the composite photometry and Type-2 QSOs SED model generated by this sample provide a guide for finding more Type-2 QSOs at higher redshift.

astro-ph.GA

Measuring the Mean Free Path of HI Ionizing Photons at $3.2\leq z\leq4.6$ with DESI Y1 Quasars

The mean free path of ionizing photons ($\lambda_\mathrm{mfp}^{912}$) in the intergalactic medium (IGM) is a crucial quantity in modelling the ionization state of IGM and the extragalactic ultraviolet background (EUVB), and is widely used in hydrodynamical simulations of galaxies and reionization. We construct the largest quasar spectrum dataset to date -- 12,595 $\mathrm{S/N}>3$ spectra -- using the Y1 observation of Dark Energy Spectroscopic Instrument (DESI) to make the most precise model-independent measurement of the mean free path at $3.2\leq z\leq 4.6$. By stacking the spectra in 17 redshift bins and modelling the Lyman continuum profile, we get a redshift evolution $\lambda_\mathrm{mfp}^{912}\propto(1+z)^{-4.27}$ at $2\leq z\leq 5$, which is much shallower than previous estimate $\lambda_\mathrm{mfp}^{912}\propto(1+z)^{-5.4}$. We then explore the sources of systematic bias, including the choice of intrinsic quasar continuum, the consideration of Lyman series opacity and Lyman limit opacity evolution and the definition of $\lambda_\mathrm{mfp}^{912}$. Combining our results with estimates of $\lambda_\mathrm{mfp}^{912}$ at higher redshifts, we conclude at high confidence that the evolution in $\lambda_\mathrm{mfp}^{912}$ steepens at $z \approx 5$. We interpret this inflection as the transition from the end of HI reionization to a fully ionized plasma which characterizes the intergalactic medium of the past $\sim10$ billion years.

astro-ph.CO

AstroMLab 3: Achieving GPT-4o Level Performance in Astronomy with a Specialized 8B-Parameter Large Language Model

AstroSage-Llama-3.1-8B is a domain-specialized natural-language AI assistant tailored for research in astronomy, astrophysics, cosmology, and astronomical instrumentation. Trained on the complete collection of astronomy-related arXiv papers from 2007 to 2024 along with millions of synthetically-generated question-answer pairs and other astronomical literature, AstroSage-Llama-3.1-8B demonstrates remarkable proficiency on a wide range of questions. AstroSage-Llama-3.1-8B scores 80.9% on the AstroMLab-1 benchmark, greatly outperforming all models -- proprietary and open-weight -- in the 8-billion parameter class, and performing on par with GPT-4o. This achievement demonstrates the potential of domain specialization in AI, suggesting that focused training can yield capabilities exceeding those of much larger, general-purpose models. AstroSage-Llama-3.1-8B is freely available, enabling widespread access to advanced AI capabilities for astronomical education and research.

astro-ph.IM

Interpreting Multi-band Galaxy Observations with Large Language Model-Based Agents

Astronomical research traditionally relies on extensive domain knowledge to interpret observations and narrow down hypotheses. We demonstrate that this process can be emulated using large language model-based agents to accelerate research workflows. We propose mephisto, a multi-agent collaboration framework that mimics human reasoning to interpret multi-band galaxy observations. mephisto interacts with the CIGALE codebase, which includes spectral energy distribution (SED) models to explain observations. In this open-world setting, mephisto learns from its self-play experience, performs tree search, and accumulates knowledge in a dynamically updated base. As a proof of concept, we apply mephisto to the latest data from the James Webb Space Telescope. mephisto attains near-human proficiency in reasoning about galaxies' physical scenarios, even when dealing with a recently discovered population of "Little Red Dot" galaxies. This represents the first demonstration of agentic research in astronomy, advancing towards end-to-end research via LLM agents and potentially expediting astronomical discoveries.

astro-ph.IM

AstroMLab 1: Who Wins Astronomy Jeopardy!?

We present a comprehensive evaluation of proprietary and open-weights large language models using the first astronomy-specific benchmarking dataset. This dataset comprises 4,425 multiple-choice questions curated from the Annual Review of Astronomy and Astrophysics, covering a broad range of astrophysical topics. Our analysis examines model performance across various astronomical subfields and assesses response calibration, crucial for potential deployment in research environments. Claude-3.5-Sonnet outperforms competitors by up to 4.6 percentage points, achieving 85.0% accuracy. For proprietary models, we observed a universal reduction in cost every 3-to-12 months to achieve similar score in this particular astronomy benchmark. open-weights models have rapidly improved, with LLaMA-3-70b (80.6%) and Qwen-2-72b (77.7%) now competing with some of the best proprietary models. We identify performance variations across topics, with non-English-focused models generally struggling more in exoplanet-related fields, stellar astrophysics, and instrumentation related questions. These challenges likely stem from less abundant training data, limited historical context, and rapid recent developments in these areas. This pattern is observed across both open-weights and proprietary models, with regional dependencies evident, highlighting the impact of training data diversity on model performance in specialized scientific domains. Top-performing models demonstrate well-calibrated confidence, with correlations above 0.9 between confidence and correctness, though they tend to be slightly underconfident. The development for fast, low-cost inference of open-weights models presents new opportunities for affordable deployment in astronomy. The rapid progress observed suggests that LLM-driven research in astronomy may become feasible in the near future.

astro-ph.IM

Probing the cosmic web in Ly$\alpha$ emission over large scales: an Intensity Mapping forecast for DECaLS/BASS and DESI

Being the most prominent HI line, Ly$\alpha$ permeates the cosmic web in emission. Despite its potential as a cosmological probe, its detection on large scales remains elusive. We present a new methodology to perform Ly$\alpha$ intensity mapping with broad-band optical images, by cross-correlating them with Ly$\alpha$ forest data using a custom one-parameter estimator. We also develop an analytical large-scale Ly$\alpha$ emission model with two parameters (average luminosity $\langle L_{\rm Ly\alpha} \rangle$ and bias $b_{\rm e}$) that respects observational constraints from QSO luminosity functions. We compute a forecast for DECaLS/BASS $g$-band images cross-correlated with DESI Ly$\alpha$ forest data, setting guidelines for reducing images into Ly$\alpha$ intensity maps. Given the transversal scales of our cross-correlation (26.4 arcmin, $\sim$33 cMpc/h), our study effectively integrates Ly$\alpha$ emission over all the cosmic volume inside the DESI footprint at $2.2 < z < 3.4$ (the $g$-band Ly$\alpha$ redshift range). Over the parameter space ($\langle L_{\rm Ly\alpha} \rangle$, $b_{\rm e}$) sampled by our forecast, we find a 3$\sigma$ of large-scale structure in Ly$\alpha$ likely, with a probability of detection of 23.95\% for DESI-DECaLS/BASS, and 54.93\% for a hypothetical DESI phase II with twice as much Ly$\alpha$ QSOs. Without a detection, we derive upper bounds on $\langle L_{\rm Ly\alpha} \rangle$ competitive with optimistic literature estimates ($2.3 \pm 1 \cdot 10^{\rm 41}$ erg/s/cMpc$^3$ for DESI, and $\sim$35\% lower for its hypothetical phase II). Extrapolation to the DESI-Rubin overlap shows that a detection of large-scale structure with Ly$\alpha$ intensity mapping using next-generation imaging surveys is certain. [abridged]

astro-ph.CO

Knowledge Graph in Astronomical Research with Large Language Models: Quantifying Driving Forces in Interdisciplinary Scientific Discovery

Identifying and predicting the factors that contribute to the success of interdisciplinary research is crucial for advancing scientific discovery. However, there is a lack of methods to quantify the integration of new ideas and technological advancements in astronomical research and how these new technologies drive further scientific breakthroughs. Large language models, with their ability to extract key concepts from vast literature beyond keyword searches, provide a new tool to quantify such processes. In this study, we extracted concepts in astronomical research from 297,807 publications between 1993 and 2024 using large language models, resulting in a set of 24,939 concepts. These concepts were then used to form a knowledge graph, where the link strength between any two concepts was determined by their relevance through the citation-reference relationships. By calculating this relevance across different time periods, we quantified the impact of numerical simulations and machine learning on astronomical research. The knowledge graph demonstrates two phases of development: a phase where the technology was integrated and another where the technology was explored in scientific discovery. The knowledge graph reveals that despite machine learning has made much inroad in astronomy, there is currently a lack of new concept development at the intersection of AI and Astronomy, which may be the current bottleneck preventing machine learning from further transforming the field of astronomy.

astro-ph.IM

Blind QSO reconstruction challenge: Exploring methods to reconstruct the Ly$\alpha$ emission line of QSOs

Reconstructing the intrinsic Ly$\alpha$ line flux from high-$z$ QSOs can place constraints on the neutral hydrogen content of the intergalactic medium during reionisation. There are now $\gtrsim10$ different Ly$\alpha$ reconstruction pipelines using different methodologies to predict the Ly$\alpha$ line flux from correlations with the spectral information redward of Ly$\alpha$. However, there have been few attempts to directly compare the performance of these pipelines. Therefore, we devised a blind QSO challenge to compare these reconstruction pipelines on a uniform set of objects. Each author was provided de-identified, observed rest-frame QSO spectra with spectral information only redward of 1260\AA\ rest-frame to ensure unbiased reconstruction. We constructed two samples of 30 QSOs, from X-Shooter and SDSS both spanning $3.5<z<4.5$. Importantly, the purpose of this comparison study was not to champion a single, best performing reconstruction pipeline but rather to explore the relative performance of these pipelines over a range of QSOs with broad observational characteristics to infer general trends. In summary, we find machine learning approaches in general provide the strongest ``best guesses" but underestimate the accompanying statistical uncertainty, although these can be recalibrated, whilst pipelines that decompose the spectral information, for example principal component or factor analysis generally perform better at predicting the Ly$\alpha$ profile. Further, we found that reconstruction pipelines trained on SDSS QSOs performed similarly on average for both the X-Shooter and SDSS samples indicating no discernible biases owing to differences in the observational characteristics of the training set or QSO being reconstructed, although the recovered distributions of reconstructions for X-Shooter were broader likely due to an increased fraction of outliers.

astro-ph.CO