SearcharxivSearch

arXiv subjects

Shu Wang

Publications and source records attributed to Shu Wang.

At least 19 recordsLinked to original sources

A Uniform Census of Variable Stars in M101 Based on HST/ACS Photometry

We present a homogeneous HST/ACS F555W/F814W time-series census of variable stars in two M101 disk fields, using a uniform multi-band reduction and reproducible variable-selection workflow for nearby SN-host galaxies. Supplementary archival F555W epochs are used only to refine selected long-period solutions. Candidates are selected with Lomb-Scargle period searches, Fourier light-curve modeling, two-band consistency statistics, artificial-star-test photometric corrections, and mode-aware Cepheid refinement. The final catalog contains 1417 variable-star candidates, including 1102 secure fundamental-mode and 78 secure first-overtone classical Cepheids. Of these secure Cepheids, 420 are newly identified relative to the Shappee & Stanek catalog. The catalog also contains luminous candidates outside the classical Cepheid locus. We release the catalog and source-level two-band light curves for studies of Cepheid physics, stellar variability, and distance-scale applications.

astro-ph.SR

The Intermediate-Mass Black Hole Reverberation Mapping Project: Scientific Overview and Sample Characteristics

Recent discoveries with the James Webb Space Telescope of massive black holes at high redshift have highlighted fundamental questions about black hole seed formation and the coevolution of black holes with their host galaxies. Because the initial seed population cannot yet be observed directly, nearby intermediate-mass black holes provide a complementary fossil record of black hole formation and early growth. Motivated by this opportunity, we present the Intermediate-Mass Black Hole Reverberation Mapping (IMBH-RM) project and construct a homogeneous Sloan Digital Sky Survey sample of active broad-line IMBHs by uniformly reanalyzing literature candidates with consistent spectral decomposition and black hole mass estimation. Our sample contains 192 reliable IMBH candidates at $z\lesssim0.3$ with $\log(M_{\rm BH}/M_\odot)<6$, including four particularly compelling sources with $\log(M_{\rm BH}/M_\odot)<5$. The primary goal of IMBH-RM is to obtain reliable black hole masses from direct measurements and characteristic sizes of the broad-line region and accretion disk for a carefully selected subsample. These measurements will provide robust low-mass anchors for calibrating single-epoch black hole mass estimates and extending black hole--galaxy scaling relations into the IMBH regime. By building a statistically meaningful reverberation-mapped sample spanning $10^4-10^6\,M_\odot$, we aim to constrain the local IMBH mass distribution and place observational constraints on competing black hole seed formation scenarios. The future Multi-Channel Imager aboard the Chinese Space-station Survey Telescope provides a particularly promising platform for achieving these goals.

astro-ph.GA

Differential Reddening and Extinction Law Analyses of Galactic Open Clusters

Extinction significantly affects open cluster parameters and their use in studies of Galactic structure, yet homogeneous large sample measurements of open cluster extinction properties remain limited. Using Gaia-era open cluster member samples combined with multi-band photometry and stellar parameters, we derive color excesses of member stars and provide the homogeneous characterization of the mean reddening, differential reddening, and color excess ratio (CER) at the cluster scale. Differential reddening increases systematically with mean reddening, with highly reddened clusters near the Galactic plane showing stronger extinction variations. Star-by-star reddening corrections narrow color--magnitude diagram (CMD) sequences in 369 of 435 clusters (85%) with reliable CMD-width measurements, and cluster color excess maps reveal small-scale extinction structures. The median CER is compatible with the standard diffuse interstellar medium extinction curve, while the broad CER distribution and its large-scale variations across the Galactic disk likely reflect differences in the dominant dust environments sampled along different Galactic sight lines.

astro-ph.GA

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization

Direct Preference Optimization (DPO) aggregates token-level log-probability ratios via uniform summation, implicitly treating all tokens as contributing equally to the preference signal. However, the contribution of individual tokens to the preference signal varies. We introduce token credit, which modulates each token's KL regularization based on its contribution to the preference outcome. We derive that effective token credit is proportional to the magnitude of each token's implicit reward, and observe that this quantity evolves substantially during training. This implies that static token credit becomes increasingly misaligned as training progresses. In this work, we propose Se-DPO (Self-Evolving Token Credit for DPO), a live mechanism that derives token credit from the model's own evolving internal signals during DPO training. Since the reward signal varies in reliability across positions, Se-DPO calibrates token credit based on both the strength and the confidence of each token's contribution. Se-DPO requires no external models, adding only a lightweight calibration network with minimal computational overhead. Experiments show that Se-DPO improves over DPO by up to 9.8 points on AlpacaEval~2 and 12.2 points on Arena-Hard.

cs.CL

JWST Reveals a Galaxy-Wide Association of Red Supergiants with OB Stars in NGC 5584

The spatial relationship between young and evolved massive-star tracers offers a means of probing recent star formation across galactic disks. We combine JWST/NIRCam, $Swift$/UVOT UVW2, and GALEX FUV/NUV images to investigate the spatial association between red-supergiant (RSG) candidates and UV-selected star-forming complexes (SFCs) in the spiral galaxy NGC 5584. We identify 5,310 RSG candidates using colour--magnitude criteria and define 106 UV-selected SFCs from the UVOT/UVW2 morphology, with their FUV and NUV luminosities measured independently using the GALEX images. We find that 82% of the UV-selected SFCs overlap significant RSG overdensities, indicating a galaxy-wide statistical association between UV-bright SFCs and the evolved massive-star population. Among the 93 SFCs with reliable JWST coverage, the UV-$\beta$-corrected FUV and NUV luminosities correlate tightly with the JWST/NIRCam luminosities, with Pearson correlation coefficients of $r=0.962$--$0.971$. MIST isochrone comparisons for 43 RSG-rich SFCs indicate a characteristic age of $\log(\mathrm{age/yr})\simeq 7.16$. These results show that RSG overdensities and UV-selected SFCs trace related but distinct phases of recent massive-star formation. Their spatial correspondence is statistical rather than one-to-one and likely reflects differences in stellar age, dust extinction, spatial resolution, aperture coverage, and local star-formation history.

astro-ph.SR

A-SR: Self-Evolving Agentic LLMs for Symbolic Regression via Hierarchical Coordination

Symbolic regression aims to discover closed-form equations from data, but existing LLM-guided methods often rely on a unified proposal loop that compresses heterogeneous search failures into a scalar score and a single prompt. We propose A-SR, a self-evolving agentic framework that shifts the control unit from expression edits to role-conditioned evidence views. A-SR coordinates formula discovery through routing among coordination protocols, an online evaluator-reward role policy, and state-routed process memory. During search, evaluator feedback characterizes reliability and productivity, updates role-level utilities, and routes elite motifs, failure traces, and validity diagnostics to different agents. The framework self-evolves at two timescales: within a run, it adapts the search process without updating LLM parameters; across runs, recorded trajectories can be distilled into open-source LLMs as role-conditioned proposal priors. Averaged over the four LSR-Synth scientific domains in LLM-SRBench, A-SR improves Acc@0.01 over baselines from 25.79% to 48.30% with Llama3.1-8B, while A-SR-LoRA improves the corresponding Qwen3-4B result from 24.58% to 38.29%. On four real-world scientific discovery tasks, A-SR obtains the best in-distribution or out-of-distribution normalized mean squared error on 7 of 8 reported metrics.

cs.CL

A Stable Mineral Fingerprint in the Fading Warm Debris Disk around HD 15407A

Extreme debris disks provide time-domain probes of rocky planet assembly after giant impacts. We compare Spitzer and JWST spectra of the warm debris disk around HD 15407A, obtained 14.95 yr apart. Across 6-25$\mu$m, the stellar-subtracted mid-infrared excess declined by $14.0\pm1.2$%, while the mineral fingerprint is nearly unchanged. Within the adopted common reduced mineral basis, Mg-rich pyroxene dominates the normalized fitted mineral weights at about 65%, while silica and forsterite contribute about 22-23% and 12-13%, respectively. In both epochs, about 93-94% of the fitted mineral weight is carried by 5.0$\mu$m grains. The fitted optically thin surface-component amplitude drops by 32%, consistent with reduced mineral-emitting flux from the disk surface rather than the appearance of a new dust composition. A simple collisional-cascade guide fitted to the IRAS-to-JWST multi-epoch ratios gives $t_c=87^{+17}_{-12}$ yr, for which the Spitzer-normalized dust-excess ratio reaches 0.5 after one $t_c$. HD 15407A is consistent with an evolved post-impact reservoir with limited recent supply of very small grains: its rocky mineral fingerprint has remained stable since at least 2008, while its warm mineral-emitting flux is gradually fading over several hundred years.

astro-ph.EP

RUBRIC: Realism--Utility Balanced Ranking for Imbalanced Classification

Class imbalance poses a fundamental challenge in risk-sensitive applications such as fraud detection and medical diagnosis, where minority-class samples are scarce yet critical for accurate classification. Existing oversampling methods generate synthetic samples to rebalance class distributions; however, they often produce large numbers of low-quality candidates that distort decision boundaries or introduce artifacts, leading to overfitting and degraded generalization. In this work, we introduce RUBRIC, a generator-agnostic filtering framework that formulates synthetic sample selection as a quality-over-quantity optimization problem. RUBRIC ranks candidates using a realism-utility trade-off: realism is quantified by a learned discriminator that distinguishes real samples from synthetic samples, while utility captures proximity to the decision boundary through a concave margin-based scoring function. We show that, under mild regularity conditions, the proposed filtering strategy monotonically tightens the generalization bound for margin-based classifiers by jointly reducing distribution shift and suppressing near-negative tail contributions. Through extensive experiments on credit-card fraud detection and other imbalanced benchmarks, we demonstrate that RUBRIC improves F1-macro and recall while maintaining comparable ROC-AUC across several generators. We also provide explicit lambda-sensitivity analysis to show how users can recover AUPRC when ranking quality is prioritized.

cs.LG

Divergent Evolution of Radial Metallicity Gradients in the Thin and Thick Disks of the Milky Way

Using 200,388 red clump stars from LAMOST and APOGEE, we investigate the radial metallicity gradients of the Galactic disk as a function of vertical height and stellar age. The thin disk displays a pronounced negative radial metallicity gradient near the Galactic mid-plane that progressively flattens with increasing $|Z|$, following $\Delta \mathrm{[Fe/H]}/\Delta R$ = $-$0.0784 $+$ 0.0776 (1 $-$ exp ($-$ $|Z|$/1.42)). The thin disk also exhibits a clear age dependence in radial metallicity gradients, evolving smoothly from a strong gradient regime for young stars to a weak gradient regime for old stars, following $\Delta \mathrm{[Fe/H]}/\Delta R$ = $-$0.0438 $+$ 0.0233 tanh (($\tau$ $-$ 11.29)/4.21). The thick disk shows weakly positive radial metallicity gradients that remain statistically invariant with respect to both vertical height and stellar age, following respectively, $\Delta \mathrm{[Fe/H]}/\Delta R$ = 0.0038 $+$ 0.0009 $|Z|$ and $\Delta \mathrm{[Fe/H]}/\Delta R$ = 0.0146 $-$ 0.0007 $\tau$. These results indicate that the thin disk retains radial metallicity gradients shaped by relatively ordered inside-out growth and long-term secular evolution processes. The thick disk exhibits spatially and temporally homogeneous radial metallicity gradients, which are consistent with a formation environment characterized by mergers of gas-rich systems and/or the turbulent ISM.

astro-ph.GA

Euclid Q1 reveals spatial variations of the extinction law in the dense cloud LDN 1641

Dust extinction laws are essential for precision photometry and provide a direct probe of grain properties, but their behaviour in dense molecular clouds remains poorly constrained at high extinction. Using Euclid Quick Data Release 1 (Q1) imaging of the Orion A dark cloud Lynds Dark Nebula 1641 (LDN 1641), we measured the extinction law from the broad Visible Instrument (VIS) band and the Near-Infrared Spectrometer and Photometer (NISP) $Y$, $J$, and $H$ bands along sightlines reaching $A_V\sim 30$ mag towards the cloud core. We derived colour-excess ratios $E(\lambda-H)/E(Y-H)$ from linear fits to colour--colour diagrams of $(\lambda-H)$ versus $(Y-H)$ and converted them into relative extinctions, $A_\lambda/A_H$. The near-infrared extinction in LDN 1641 is well described by a power law, $A_\lambda\propto \lambda^{-\alpha}$, with $\alpha=1.57 \pm 0.06$, corresponding to $A_{\rm VIS}/A_H=4.23 \pm 0.24$, $A_Y/A_H=2.13 \pm 0.18$, and $A_J/A_H=1.47 \pm 0.11$. We further find significant spatial variations: $\alpha$ changes by up to $27\%$, with systematically smaller values and therefore flatter extinction curves in higher-extinction regions. This flattening is consistent with an enhanced large-grain population and supports substantial grain growth from the diffuse outskirts to the dense core of a single molecular cloud.

astro-ph.GA

Distance-Ladder Measurements of the Hubble Constant: Recent Progress, Systematics, and Prospects

The Hubble constant, \(H_0\), links the nearby distance scale to the present cosmic expansion rate. Local distance-ladder measurements now reach percent-level precision and remain more than \(5\sigma\) higher than the value inferred from cosmic microwave background (CMB) observations in base-\(\Lambda\)CDM, making the reliability of the local ladder a central issue in the Hubble tension. We describe the ladder as a covariance network connecting level-0 geometric anchors, level-1 stellar distance indicators, and level-2 Hubble-flow probes. The Cepheid--Type Ia supernova (SN Ia) route remains the most precise single local ladder, but independent indicators including the tip of the red giant branch (TRGB), J-region asymptotic giant branch (JAGB) stars, Mira variables, surface-brightness fluctuations (SBF), the Tully--Fisher relation, and Type II supernovae (SNe II) now test shared and method-specific systematics. In a compact seven-route covariance summary, combining the Cepheid--SN Ia route with three level-1 alternatives (TRGB, JAGB, and Mira) and three level-2 alternatives (SBF, Tully--Fisher, and SNe II) gives \(H_0=73.30\pm0.92~{\rm km~s^{-1}~Mpc^{-1}}\), still \(5.6\sigma\) above Planck base-\(\Lambda\)CDM. JWST has already tested Cepheid crowding and is making independent TRGB-based \(H_0\) measurements increasingly feasible. Over the next five years, a reliable one-percent local \(H_0\) requires larger calibrator samples, cross-validated level-1 zero points, explicit covariance propagation, and AI-assisted, reproducible, pre-specified selection criteria for distance-indicator measurements.

astro-ph.CO

A Planetary Nebula from a 5.7 $M_{\odot}$ Progenitor in a 90 Myr M31 Star Cluster

Planetary nebulae (PNe) trace the late evolution of low-to-intermediate-mass stars, yet the masses of their progenitors are rarely measured directly. Here we present a PN physically associated with a young star cluster in M31, providing an unprecedented extragalactic empirical anchor in the poorly constrained high-mass regime of PN progenitors. High-resolution Hubble Space Telescope imaging shows that the nebula lies near the cluster center, and spectral decomposition of the blended cluster-plus-nebula spectrum yields consistent stellar and nebular radial velocities, strongly supporting a physical association. Isochrone fitting to the color-magnitude diagram indicates a cluster age of ~90 Myr and a near-solar metallicity, implying a progenitor initial mass of $5.66^{+0.42}_{-0.37}\,M_{\odot}$. This value is among the highest empirical progenitor-mass constraints yet reported for any PN and approaches the lower boundary of the super-asymptotic giant branch (super-AGB) regime. We further find that the nebula is strongly nitrogen-enhanced, with an N/O ratio ~7 times the solar value, broadly consistent with hot bottom burning in a relatively massive AGB progenitor. This system therefore provides a rare opportunity to test PN formation and nucleosynthesis at the high-mass end of the PN progenitor distribution.

astro-ph.SR

MT-EditFlow: Reinforcement Learning for Multi-Turn Image Editing with Flow Matching

Recent breakthroughs in instruction-based image editing have captured significant attention, as models are now capable of handling real-world editing demands with the practicality required by everyday users. However, editing models trained primarily for single-turn edits often break down in multi-turn editing--the natural interactive setting where a user iteratively refines an image based on the model's own previous outputs. This failure stems from the all-or-nothing requirement, where a single failed turn compromises the entire sequence, and error propagation, where exposure bias leads to compounding editing errors. To address these challenges, we introduce MT-EditFlow, a flow-matching reinforcement learning framework designed to optimize reward signals for sequential image editing. MT-EditFlow integrates a multi-turn perspective with a multi-reward formulation to provide a unified structure applicable to both GRPO and NFT-based reinforcement learning methods. We systematically analyze and optimize the reward signal by investigating effective scoring strategies for turn-level aggregation, VLM reasoning modes to trade off reward bias and variance, and advantage fusion levels to prevent reward hacking. Our findings reveal that broadcasting the aggregated advantage across the entire editing trajectory effectively bridges the gap between local planning and global multi-turn task success. Extensive experiments demonstrate that MT-EditFlow significantly improves performance across diverse base models. Notably, it boosts FLUX.1-Kontext-dev by 6.85 points in turn-3 overall performance, surpassing state-of-the-art open-source models such as Qwen-Image-Edit. By maintaining high marginal success rates and reducing exposure bias, MT-EditFlow provides a foundation for more reliable and natural human-AI collaboration in visual content creation.

cs.CV

Mean Field Competition of Optimal Switching: The Vanishing Entropy Regularization Approach

This paper studies a type of rank-based mean field game in which competing agents strategically switch among multiple effort regimes. We propose an entropy regularized auxiliary problem where the switching decisions are randomized to the control of transition probability for a continuous-time finite-state Markov chain. We first establish the existence of regularized equilibrium in this auxiliary problem. Assuming the convexity of reward scheme, we then prove that the equilibrium is unique and can be approximated by a fictitious play iteration scheme. Furthermore, as the entropy regularization vanishes, we establish the convergence analysis of the regularized equilibrium towards the relaxed equilibrium in the original MFG of optimal switching. The uniqueness of the population ranking distribution under the relaxed equilibrium is also obtained given a strictly convex reward scheme.

math.OC

Implementing the principal stratum strategy for intercurrent events with survival outcomes: a tutorial

The International Council for Harmonization (ICH) E9 (R1) addendum provides the estimand framework to formulate treatment effects in a clinical trial. One of the attributes of an estimand the framework describes is intercurrent events. Among the five strategies to intercurrent events the guidance lists, the principal stratum strategy is the most conceptually and technically challenging because it defines treatment effects on unobserved strata. Its application to survival outcomes is particularly inaccessible to practitioners. This tutorial reviews the methodology and implementation of the estimand framework with the principal stratum strategy to address intercurrent events with survival outcomes. We illustrate using a clinical trial in oncology and focus on a simple case with binary treatment and a single binary intercurrent event of discontinuation of the assigned treatment. We define the causal effects and review two main methods for estimating the effects: the mixture model method and the weighting method. For each method, we elaborate the associated assumptions, models, sensitivity analysis, software and provide example R code. We conduct simulation studies that mimic the real study to study the operation characteristics of these methods.

stat.ME

DBES: A Systematic Benchmark and Metric Suite for Evaluating Expert Specialization in Large-Scale MoEs

Expert specialization in Mixture-of-Experts (MoE) models remains poorly understood, with traditional evaluations conflating architectural load-balancing with functional specialization. We introduce DBES, a comprehensive diagnostic framework combining a multi-domain benchmark with five theoretically grounded metrics: Routing Specialization, Normalized Effective Rank, Domain Isolation, Routing Stiffness Score, and N-gram Expertise measures. Critical findings demonstrate distinct specialization paradigms across models: Qwen-series exhibit modular specialization with high domain isolation, while DeepSeek and GLM employ distributed collaboration. However, we emphasize that specialization is a diagnostic dimension, necessary but not sufficient for downstream performance. Most crucially, interventional evidence validates the actionability of these metrics: by using DBES to identify high-specialization expert paths during domain-specific post-training, we achieved 66% to 94.48% improvement in specialized domains with only 15% of original training resources, demonstrating that these diagnostic tools can be converted into concrete optimization operators. This work provides the first systematic methodology for evaluating expert specialization independently of accuracy metrics, offering crucial insights for the design and post-training optimization of next-generation MoE systems.

cs.LG

SkillRAE: Agent Skill-Based Context Compilation for Retrieval-Augmented Execution

Large Language Model (LLM)-based agents (e.g., OpenClaw) increasingly rely on reusable skill libraries to solve artifact-rich tasks such as document-centric workflows and data-intensive analysis. As these libraries grow, a few works have attempted to study the Retrieval-Augmented Execution (RAE), which often first retrieves some external skills and other knowledge, then compiles the context using retrieved skills, and finally executes the task. Existing works mainly focus on optimizing skill retrieval and task execution, and they pay little attention to how to effectively organize the selected skill evidence in a form that is compact, grounded, and immediately usable for the downstream executors to complete tasks. To fill this gap, we propose SkillRAE, a two-stage RAE approach focusing on skill-based context compilation, which consists of the offline and online stages. Specifically, in the offline indexing stage, it builds a multi-level skill graph over skill communities, skills, and reusable subunits, for capturing their relationships. In the online retrieval stage, it first performs skill-ranked retrieval with selected-subunit evidence export in the graph, and then applies rescue-aware compact compilation to recover the key evidence. Together, these components compile a coarse-ranked skill set into a task-specific context that is compact, grounded, and immediately usable. Experiments on two public benchmarks show that SkillRAE achieves a significant improvement over baselines for RAE. For example, on SkillsBench, it achieves an improvement of 11.7% over the SOTA method. Ablation studies further show that our context compilation is crucial, instead of a mere prompt addition.

cs.CL

ASTRA-QA: A Benchmark for Abstract Question Answering over Documents

Document-based question answering (QA) increasingly includes abstract questions that require synthesizing scattered information from long documents or across multiple documents into coherent answers. However, this setting is still poorly supported by existing benchmarks and evaluation methods, which often lack stable abstract references or rely on coarse similarity metrics and unstable head-to-head comparisons. To alleviate this issue, we introduce ASTRA-QA, a benchmark for AbSTRAct Question Answering over documents. ASTRA-QA contains 869 QA instances over academic papers and news documents, covering five abstract question types and three controlled retrieval scopes. Each instance is equipped with explicit evaluation annotations, including answer topic sets, curated unsupported topics, and aligned evidence. Building on these annotations, ASTRA-QA assesses whether answers cover required key points and avoid unsupported content by directly scoring topic coverage and curated unsupported content, enabling scalable evaluation without exhaustive head-to-head comparisons. Experiments with representative Retrieval-Augmented Generation (RAG) methods spanning vanilla, graph-based, and hierarchical retrieval settings show that ASTRA-QA provides reference-grounded diagnostics for coverage, hallucination, and retrieval-scope robustness. Our dataset and code are available at https://xinyangsally.github.io/astra-benchmark.

cs.CL