SearcharxivSearch

arXiv subjects

Xinyi Song

Publications and source records attributed to Xinyi Song.

9 recordsLinked to original sources

Life 2.0: A Scalable Distributed Space-Telescope Array for Biosignature Spectroscopy

Answering the question "Are we alone?" requires atmospheric spectroscopy of nearby terrestrial planets. For an Earth--Sun analog, even the strongest transmission signals are expected to be of order 1 part per million (ppm). Unlike short-period planets, Earth 2.0 planets transit only about once per year, so single-transit sensitivity, rather than stacking repeated observations, is the fundamental design driver. Life 2.0 is a scalable space-mission concept linking Earth 2.0 candidates discovered by PLATO and the Earth 2.0 (ET) mission with atmospheric characterization and biosignature assessment. The baseline architecture comprises 900 one-meter space telescopes, each equipped with a high-throughput Waveguide Integrated Miniature Spectrograph and an ultra-low-read-noise CMOS detector. After independent calibration, spectra acquired simultaneously during a transit are combined, providing the photon-collecting capability of an approximately 30-m aperture at the selected spectral resolution while retaining a modular architecture. The baseline 0.2--1.05 $\mu$m range covers O$_3$, O$_2$, H$_2$O, Rayleigh scattering, and other diagnostics, with extension into the infrared as detector technologies mature. Prototype Waveguide Spectral Lens devices have demonstrated 40--66\% throughput at resolving powers from $R \sim 200$ to $R \sim 20{,}000$. Lightweight silicon-carbide mirrors and sub-electron-noise CMOS detectors support replicated production. Life 2.0 must address detector systematics, instrument stability, and stellar variability; rather than assuming these limitations disappear, it builds on calibration, detector-characterization, and data-analysis techniques advanced during the JWST era. The concept offers a scalable alternative to a monolithic 30-m-class space telescope and a staged pathway toward biosignature spectroscopy of nearby Earth-like planets.

astro-ph.IM

The influence of hypothetical exomoons on planetary thermal phase curves

More than 200 moons exist in our Solar System, yet no exomoon has been confirmed to date. While the innermost two planets of the Solar System lack natural satellites and most studies favour the existence of exomoons around long-period planets, some theoretical studies that take tidal dissipation, orbital decay, and migration processes into account suggest that exomoons may survive around short-period exoplanets. We investigated the impact of exomoons on planetary thermal phase curves and assessed their detectability within a theoretical framework. We simulated the thermal phase curves of exomoon-exoplanet systems, including mutual transits and occultations, and explored their dependence on planetary orbital periods across a wide range of systems. Close-in airless exomoons maintain large day-night temperature contrasts, amplifying the thermal phase-curve signal of the system. When the exomoon transits or is occulted by the exoplanet, the transit depth varies with the planetary phase, and the occultation depth varies with the exomoon's phase. The maximum occultation depth can reach $\sim$ 20 ppm for long-period systems. For short-period planets, the signal can reach up to $\sim$100 ppm, although such configurations may not be dynamically stable over long timescales. If exomoons are not accounted for, the planetary temperature distribution retrieved from observed thermal phase curves may overestimate the planetary day-night temperature contrast and underestimate the planetary horizontal heat transport. In principle, the periodic exomoon-exoplanet mutual occultation signal could be extracted using methods such as box-fitting least squares, providing a framework for future observational studies and instrument planning.

astro-ph.EP

StatLLM: A Dataset for Evaluating the Performance of Large Language Models in Statistical Analysis

The coding capabilities of large language models (LLMs) have opened up new opportunities for automatic statistical analysis in machine learning and data science. However, before their widespread adoption, it is crucial to assess the accuracy of code generated by LLMs. A major challenge in this evaluation lies in the absence of a benchmark dataset for statistical code (e.g., SAS and R). To fill in this gap, this paper introduces StatLLM, an open-source dataset for evaluating the performance of LLMs in statistical analysis. The StatLLM dataset comprises three key components: statistical analysis tasks, LLM-generated SAS code, and human evaluation scores. The first component includes statistical analysis tasks spanning a variety of analyses and datasets, providing problem descriptions, dataset details, and human-verified SAS code. The second component features SAS code generated by ChatGPT 3.5, ChatGPT 4.0, and Llama 3.1 for those tasks. The third component contains evaluation scores from human experts in assessing the correctness, effectiveness, readability, executability, and output accuracy of the LLM-generated code. We also illustrate the unique potential of the established benchmark dataset for (1) evaluating and enhancing natural language processing metrics, (2) assessing and improving LLM performance in statistical coding, and (3) developing and testing of next-generation statistical software - advancements that are crucial for data science and machine learning research.

stat.AP

Performance Evaluation of Large Language Models in Statistical Programming

The programming capabilities of large language models (LLMs) have revolutionized automatic code generation and opened new avenues for automatic statistical analysis. However, the validity and quality of these generated codes need to be systematically evaluated before they can be widely adopted. Despite their growing prominence, a comprehensive evaluation of statistical code generated by LLMs remains scarce in the literature. In this paper, we assess the performance of LLMs, including two versions of ChatGPT and one version of Llama, in the domain of SAS programming for statistical analysis. Our study utilizes a set of statistical analysis tasks encompassing diverse statistical topics and datasets. Each task includes a problem description, dataset information, and human-verified SAS code. We conduct a comprehensive assessment of the quality of SAS code generated by LLMs through human expert evaluation based on correctness, effectiveness, readability, executability, and the accuracy of output results. The analysis of rating scores reveals that while LLMs demonstrate usefulness in generating syntactically correct code, they struggle with tasks requiring deep domain understanding and may produce redundant or incorrect results. This study offers valuable insights into the capabilities and limitations of LLMs in statistical programming, providing guidance for future advancements in AI-assisted coding systems for statistical analysis.

stat.AP

Applied Statistics in the Era of Artificial Intelligence: A Review and Vision

The advent of artificial intelligence (AI) technologies has significantly changed many domains, including applied statistics. This review and vision paper explores the evolving role of applied statistics in the AI era, drawing from our experiences in engineering statistics. We begin by outlining the fundamental concepts and historical developments in applied statistics and tracing the rise of AI technologies. Subsequently, we review traditional areas of applied statistics, using examples from engineering statistics to illustrate key points. We then explore emerging areas in applied statistics, driven by recent technological advancements, highlighting examples from our recent projects. The paper discusses the symbiotic relationship between AI and applied statistics, focusing on how statistical principles can be employed to study the properties of AI models and enhance AI systems. We also examine how AI can advance applied statistics in terms of modeling and analysis. In conclusion, we reflect on the future role of statisticians. Our paper aims to shed light on the transformative impact of AI on applied statistics and inspire further exploration in this dynamic field.

stat.AP

A Comprehensive Case Study on the Performance of Machine Learning Methods on the Classification of Solar Panel Electroluminescence Images

Photovoltaics (PV) are widely used to harvest solar energy, an important form of renewable energy. Photovoltaic arrays consist of multiple solar panels constructed from solar cells. Solar cells in the field are vulnerable to various defects, and electroluminescence (EL) imaging provides effective and non-destructive diagnostics to detect those defects. We use multiple traditional machine learning and modern deep learning models to classify EL solar cell images into different functional/defective categories. Because of the asymmetry in the number of functional vs. defective cells, an imbalanced label problem arises in the EL image data. The current literature lacks insights on which methods and metrics to use for model training and prediction. In this paper, we comprehensively compare different machine learning and deep learning methods under different performance metrics on the classification of solar cell EL images from monocrystalline and polycrystalline modules. We provide a comprehensive discussion on different metrics. Our results provide insights and guidelines for practitioners in selecting prediction methods and performance metrics.

stat.AP

Critical role of vertical radiative cooling contrast in triggering episodic deluges in small-domain hothouse climates

Seeley and Wordsworth (2021) showed that in small-domain cloud-resolving simulations the pattern of precipitation transforms in extremely hot climates ($\ge$ 320 K) from quasi-steady to organized episodic deluges, with outbursts of heavy rain alternating with several dry days. They proposed a mechanism for this transition involving increased water vapor absorption of solar radiation leading to net lower-tropospheric radiative heating. This heating inhibits lower-tropospheric convection and decouples the boundary layer from the upper troposphere during the dry phase, allowing lower-tropospheric moist static energy to build until it discharges, resulting in a deluge. We perform cloud-resolving simulations in polar night and show that the same transition occurs, implying that some revision of their mechanism is necessary. We show that episodic deluges can occur even if the lower-tropospheric radiative heating rate is negative, as long as the magnitude of the upper-tropospheric radiative cooling is about twice as large. We find that in the episodic deluge regime the mean precipitation can be inferred from the atmospheric column energy budget and the period can be predicted from the time for radiation and reevaporation to cool the lower atmosphere.

physics.ao-ph

Cloud Behaviour on Tidally Locked Rocky Planets from Global High-resolution Modeling

Determining the behaviour of convection and clouds is one of the biggest challenges in our understanding of exoplanetary climates. Given the lack of in situ observations, one of the most preferable approaches is to use cloud-resolving or cloud-permitting models (CPM). Here we present CPM simulations in a quasi-global domain with high spatial resolution (4$\times$4 km grid) and explicit convection to study the cloud regime of 1 to 1 tidally locked rocky planets orbiting around low-mass stars. We show that the substellar region is covered by deep convective clouds and cloud albedo increases with increasing stellar flux. The CPM produces relatively less cloud liquid water concentration, smaller cloud coverage, lower cloud albedo, and deeper H2O spectral features than previous general circulation model (GCM) simulations employing empirical convection and cloud parameterizations. Furthermore, cloud streets--long bands of low-level clouds oriented nearly parallel to the direction of the mean boundary-layer winds--appear in the CPM and substantially affect energy balance and surface precipitation at a local level.

astro-ph.EP

Asymmetry and Variability in the Transmission Spectra of Tidally Locked Habitable Planets

Spatial heterogeneity and temporal variability are general features in planetary weather and climate, due to the effects of planetary rotation, uneven stellar flux distribution, fluid motion instability, etc. In this study, we investigate the asymmetry and variability in the transmission spectra of 1:1 spin--orbit tidally locked (or called synchronously rotating) planets around low-mass stars. We find that for rapidly rotating planets, the transit atmospheric thickness on the evening terminator (east of the substellar region) is significantly larger than that of the morning terminator (west of the substellar region). The asymmetry is mainly related to the spatial heterogeneity in ice clouds, as the contributions of liquid clouds and water vapor are smaller. The underlying mechanism is that there are always more ice clouds on the evening terminator, due to the combined effect of coupled Rossby--Kelvin waves and equatorial superrotation that advect vapor and clouds to the east, especially at high levels of the atmosphere. For slowly rotating planets, the asymmetry reverses (the morning terminator has a larger transmission depth than the evening terminator) but the magnitude is small or even negligible. For both rapidly and slowly rotating planets, there is strong variability in the transmission spectra. The asymmetry signal is nearly impossible to be observed by the James Webb Space Telescope (JWST), because the magnitude of the asymmetry (about 10 ppm) is smaller than the instrumental noise and the high variability further increases the challenge.

astro-ph.EP