SearcharxivSearch

arXiv subjects

Shichao Han

Publications and source records attributed to Shichao Han.

6 recordsLinked to original sources

Galactic HII regions in LAMOST Medium-Resolution Spectroscopic Survey of Nebulae

Based on LAMOST Medium-Resolution Spectroscopic Survey of Nebulae (MRS-N) data and WISE Galactic HII region catalog, we construct a sample of 280 Galactic HII regions and candidates in the Outer Galaxy (80$^{\circ}$ $\lesssim$ l $\lesssim$ 220$^{\circ}$). Using MRS-N optical spectra, we measure four emission lines (H$\alpha$, [NII]$\lambda$6584, [SII]$\lambda\lambda$6717,6731) and use line-ratios to spectroscopically confirm 255 HII regions, including 90 previously "Known" HII regions and 165 newly classified ones. We measure their $T_{\rm e}$, $n_{\rm e}$ and oxygen abundance, and determine distances via associated OB stars and the kinematic method. The sample spans $R_{\rm gal}$ from 8.16 to 15.36 kpc, enabling investigation of radial gradients in physical properties. We find [NII]/H$\alpha$ and [SII]/H$\alpha$ decrease with increasing $R_{\rm gal}$, while [SII]/[NII] remains nearly flat; these trends are quite different from diffuse ionized gas (DIG). We derive the $T_{\rm e}$ gradient of 344.530 $\pm$ 78.083 K kpc$^{-1}$, and the $\log n_{\rm e}$ gradient of -0.143 $\pm$ 0.041 cm$^{-3}$ kpc$^{-1}$. Oxygen abundance shows a steep slope of -0.044 $\pm$ 0.010 dex kpc$^{-1}$ in the inner disk and a shallow slope of -0.016 $\pm$ 0.005 dex kpc$^{-1}$ in the outer disk, with a global slope of -0.014 $\pm$ 0.005 dex kpc$^{-1}$. We also examine the two-dimensional distributions of $T_{\rm e}$, $n_{\rm e}$, and oxygen abundance, and find the gradients vary with azimuth. There is no obvious difference between spiral arm and interarm regions, and no trend appears along individual arms. From [NII]/H$\alpha$-[SII]$\lambda$6717/H$\alpha$ diagram, HII regions have a S$^+$/S ratio (0.32), lower than DIG (0.43); however, heavy overlap prevents clear separation from this diagram alone.

astro-ph.GA

The size-velocity dispersion relationship of Galactic HII regions

The size-velocity dispersion ($\sigma$) relation, while well established for giant HII regions, remains uncertain for their smaller counterparts (physical radii R < 20 pc). Thanks to the LAMOST MRS-N dataset's large sky coverage and high spatial/spectral resolution, we examined this relationship using 10 isolated Galactic HII regions with R < 20 pc. Our results reveal two key findings: (1) these small-size HII regions remarkably follow the same size-$\sigma$ relation as giant HII regions, suggesting this correlation could serve as a novel distance indicator for Galactic HII regions; and (2) we find distinct dynamical behaviors between younger and older HII regions. Specifically, in younger (< 0.5 Myr), ionization-bounded HII regions, the velocity dispersion shows no correlation with expansion velocity, indicating that turbulence is driven primarily by stellar winds and ionization processes. In contrast, in older (> 0.5 Myr), matter-bounded HII regions, a clear correlation emerges, implying that expansion-driven processes begin to play a significant role in generating turbulence. We therefore propose an evolutionary transition in the primary turbulence mechanisms, from being dominated by stellar winds and radiation to being increasingly influenced by expansion-driven dynamics, during the evolution of HII regions. Considering the small sample size used in this work, particularly the inclusion of only two young HII regions, which also have large uncertainties in their expansion velocities, further confirmation of this interpretation will require higher-resolution 2D spectroscopy to resolve blended kinematic components along the line of sight for more accurate estimation of expansion velocities, along with an expanded sample that specifically includes more young HII regions.

astro-ph.GA

Enhancing External Validity of Experiments with Ongoing Sampling

Participants in online experiments often enroll over time, which can compromise sample representativeness due to temporal shifts in covariates. This issue is particularly critical in A/B tests, online controlled experiments extensively used to evaluate product updates, since these tests are cost-sensitive and typically short in duration. We propose a novel framework that dynamically assesses sample representativeness by dividing the ongoing sampling process into three stages. We then develop stage-specific estimators for Population Average Treatment Effects (PATE), ensuring that experimental results remain generalizable across varying experiment durations. Leveraging survival analysis, we develop a heuristic function that identifies these stages without requiring prior knowledge of population or sample characteristics, thereby keeping implementation costs low. Our approach bridges the gap between experimental findings and real-world applicability, enabling product decisions to be based on evidence that accurately represents the broader target population. We validate the effectiveness of our framework on three levels: (1) through a real-world online experiment conducted on WeChat; (2) via a synthetic experiment; and (3) by applying it to 600 A/B tests on WeChat in a platform-wide application. Additionally, we provide practical guidelines for practitioners to implement our method in real-world settings.

econ.GN

MultiObjMatch: Matching with Optimal Tradeoffs between Multiple Objectives in R

In an observational study, matching aims to create many small sets of similar treated and control units from initial samples that may differ substantially in order to permit more credible causal inferences. The problem of constructing matched sets may be formulated as an optimization problem, but it can be challenging to specify a single objective function that adequately captures all the design considerations at work. One solution, proposed by \citet{pimentel2019optimal} is to explore a family of matched designs that are Pareto optimal for multiple objective functions. We present an R package, \href{https://github.com/ShichaoHan/MultiObjMatch}{\texttt{MultiObjMatch}}, that implements this multi-objective matching strategy using a network flow algorithm for several common design goals: marginal balance on important covariates, size of the matched sample, and average within-pair multivariate distances. We demonstrate the package's flexibility in exploring user-defined tradeoffs of interest via two case studies, a reanalysis of the canonical National Supported Work dataset and a novel analysis of a clinical dataset to estimate the impact of diabetic kidney disease on hospitalization costs.

stat.ME

Estimating Treatment Effects under Algorithmic Interference: A Structured Neural Networks Approach

Online user-generated content platforms allocate billions of dollars of promotional traffic through algorithms in two-sided marketplaces. To evaluate updates to these algorithms, platforms frequently rely on creator-side randomized experiments. However, because treated and control creators compete for exposure, such experiments suffer from algorithmic interference: exposure outcomes depend on competitors' treatment status. We show that commonly used difference-in-means estimators can therefore be severely biased and may even recommend deploying inferior algorithms. To address this challenge, we develop a structured semiparametric framework that explicitly models the competitive allocation mechanism underlying exposure. Our approach combines an algorithm choice model that characterizes how exposure is allocated across competing content with a viewer response model that captures engagement conditional on exposure. We construct a debiased estimator grounded in the double machine learning framework to recover the global treatment effect of platform-wide rollout. Methodologically, we extend DML asymptotic theory to accommodate correlated samples arising from overlapping consideration sets. Using Monte Carlo simulations and a large-scale field experiment on a major short-video platform, we show that our estimator closely matches an interference-free benchmark obtained from a costly double-sided experimental design. In contrast, standard estimators exhibit substantial bias and, in some cases, even reverse the sign of the effect.

econ.EM

Treatment Effect Detection with Controlled FDR under Dependence for Large-Scale Experiments

Online controlled experiments (also known as A/B Testing) have been viewed as a golden standard for large data-driven companies since the last few decades. The most common A/B testing framework adopted by many companies use "average treatment effect" (ATE) as statistics. However, it remains a difficult problem for companies to improve the power of detecting ATE while controlling "false discovery rate" (FDR) at a predetermined level. One of the most popular FDR-control algorithms is BH method, but BH method is only known to control FDR under restrictive positive dependence assumptions with a conservative bound. In this paper, we propose statistical methods that can systematically and accurately identify ATE, and demonstrate how they can work robustly with controlled low FDR but a higher power using both simulation and real-world experimentation data. Moreover, we discuss the scalability problem in detail and offer comparison of our paradigm to other more recent FDR control methods, e.g., knockoff, AdaPT procedure, etc.

stat.ME