SearcharxivSearch

arXiv subjects

Yu Rong

Publications and source records attributed to Yu Rong.

At least 19 recordsLinked to original sources

A Direct DESI--SDSS Two-Fibre Test of Aperture Bias in Optical Emission-Line Diagnostics

Fixed angular apertures sample different physical regions of nearby galaxies and can therefore bias optical emission-line diagnostics. We use 20,545 high-confidence emission-line galaxies observed by both the Dark Energy Spectroscopic Instrument (DESI) and the Sloan Digital Sky Survey (SDSS) as a two-fibre experiment, in which the DESI 1.5 arcsec fibre is nested within the SDSS 3 arcsec fibre. To isolate aperture effects from line-measurement systematics, we refit every spectrum using the same eMILES stellar-continuum and emission-line model while preserving the wavelength-dependent line-spread function of each spectrum. This common refit yields 20,200 quality-controlled matched galaxies. Relative to SDSS, the smaller DESI aperture produces small but highly significant offsets: $\Delta\log({\rm N2})=+0.0067\pm0.0002$, $\Delta\log({\rm S2})=-0.0148\pm0.0003$, $\Delta\log({\rm O3})=-0.0246\pm0.0006$, $\Delta{\rm O3N2}=-0.0327\pm0.0007$, and $\Delta\log({\rm H}\alpha/{\rm H}\beta)=-0.0195\pm0.0003$, where $\Delta$ denotes DESI minus SDSS. These offsets persist in Baldwin--Phillips--Terlevich-selected star-forming galaxies and in a near-concentric subset. Their amplitudes vary across the redshift distribution and retain a measurable dependence on apparent fibre coverage after redshift is controlled, showing that redshift and relative aperture coverage jointly modulate the aperture terms. We conclude that, for high-S/N emission-line galaxies, DESI--SDSS comparisons carry measurable, diagnostic-dependent fibre-aperture terms even when continuum modelling, emission-line fitting, and spectral-resolution treatment are held fixed.

astro-ph.GA

On the Design Fundamentals of Pixel Text Representation Learning

Text-rich visual inputs require models that can read, retrieve, and compress language directly in pixel space, yet existing pixel-text encoders struggle with fixed resolution pretraining, visual shortcut learning, weak visual grounding, and multilingual visual text understanding. In this work, we investigate the fundamental design principles required for robust visual text representation learning. Through systematic controlled ablations, we identify four critical components: variable image resolutions and rendered font sizes provide spatial proxies for high-resolution document generalization; natural image-text pairs are indispensable for grounding and prevent text-only collapse; layout-aware rendering helps prevent pixel-level shortcuts; and a two-stage multilingual curriculum enables effective cross-lingual alignment. By integrating these principles into a scalable training recipe, we train Pixel Linguist II, a native-resolution vision encoder trained with on-the-fly rendering, unified contrastive grounding, and a multilingual curriculum over 280M training examples. Pixel Linguist II sets new state-of-the-art results on English, cross-lingual, and multilingual Visual STS and ViDoRe, while also enabling better MLLM downstream evaluation. Notably, Pixel Linguist II remains robust under 80\% visual token compression, showing great promise for optical context compression. Our code and resources are available at https://github.com/Pixel-Linguist/Pixel-Linguist-II.

cs.CV

GUI-CC: Benchmarking Contextual Consistency of GUI World Models as Agent Environments

GUI world models are increasingly evaluated as one-step next-screen predictors, yet their intended use is often as multi-step environments for GUI agents. This mismatch leaves a key requirement under-tested: generated states must remain contextually consistent when they are repeatedly reused for future interaction. We introduce GUI-CC, a benchmark that evaluates contextual consistency of GUI world models as agent environments rather than isolated next-screen predictors. GUI-CC contains two complementary tracks: an offline reference-action track that rolls models along real mobile GUI trajectories, and an online agent-loop track that lets fixed probing agents interact with model-generated UIs. We construct 500 offline trajectory tasks from GUIOdyssey and 200 emulator-verified online tasks across 30 mobile apps. GUI-CC evaluates transition fidelity, transition plausibility, contextual consistency, and task progress. Experiments show that plausible single-step generation does not guarantee reliable environment simulation: current models often produce usable-looking screens while failing to preserve task-relevant context or support executable multi-step rollouts.

cs.CL

Evidence for the transformation from lenticular to spiral galaxies

It is widely accepted that late-type galaxies, such as spirals, evolve into early-type systems, including elliptical and lenticular galaxies, through galaxy mergers and violent disk instability processes. Throughout this morphological transformation, star formation is typically suppressed by quenching mechanisms whose detailed nature remains the subject of active investigation. Here, we present compelling evidence for an evolutionary pathway that proceeds in the reverse direction. Using the integral field unit observations, we identify a population of spiral galaxies hosting quenched central cores (QCCs). These galaxies exhibit bimodal distributions in both their stellar population properties and their dynamical properties, along with sharp changes in radial gradients near the QCC boundary. These results indicate that the QCCs and the surrounding outer disks formed at distinct cosmic epochs and through different physical processes. Remarkably, QCCs closely resemble quiescent early-type galaxies, particularly lenticular galaxies, in their mass-size and mass-velocity dispersion scaling relations, as well as in their stellar population demographics and internal kinematics. These findings provide strong support for a rejuvenation scenario in which spiral disks are reassembled around pre-existing quiescent lenticular or early-type systems. Moreover, we show that such rejuvenation, accompanied by a reverse morphological transformation from early- to late-type appearance, is quite common. This indicates that quenching in galaxies is not invariably a terminal state and can be reversed under appropriate conditions.

astro-ph.GA

Electron Densities of Typical Low-Mass Galaxies at z~2-7 from Stacked JWST/NIRSpec Spectra

Direct electron-density measurements at high redshift are usually limited to galaxies with individually strong density-sensitive doublets, and therefore may not trace the average interstellar medium of ordinary low-mass galaxies. We stack public JWST/NIRSpec medium-resolution spectra from the DAWN JWST Archive to measure [SII]-based electron densities $n_e$ for low-mass galaxies at $2 5$ have a higher normalization, $n_{e,0}=211^{+36}_{-31}\ {\rm cm^{-3}}$, showing that individual-doublet samples select a denser subset. Stacking archival JWST spectra therefore provides a direct route to measuring the average gas density of low-mass galaxies below the individual-doublet detection threshold.

astro-ph.GA

Stellar Surface Density Modulates MgII Cool-gas Outflow Absorption in DESI Star-forming Galaxies

Galaxy outflows are usually ordered by stellar mass and star-formation rate (SFR), but the same feedback budget may couple differently to gas in diffuse and compact galaxies. We use Dark Energy Spectroscopic Instrument (DESI) Data Release 1 stacked spectra of massive star-forming galaxies at $0.35<z<1.0$ to test whether stellar surface density, $\Sigma_\star=M_\star/(2\pi R_e^2)$, is an independent empirical coordinate of down-the-barrel singly ionized magnesium (MgII) cool-gas absorption. In AGN-clean samples matched in stellar mass, and in a stricter sample matched in both stellar mass and a Balmer-line SFR proxy, the MgII outflow equivalent width (EW) rises monotonically with $\Sigma_\star$ in every redshift bin. From the lowest to highest $\Sigma_\star$ tertile, EW increases by 0.37-0.61 Angstrom, while the absolute outflow velocity changes only weakly. DESI therefore shows that cool-gas outflow strength in massive star-forming galaxies is not set only by how much stellar mass or star formation a galaxy has, but also by how tightly the galaxy is built. The structural dependence points to changes in the absorbing velocity distribution and/or the effective covering fraction of cool outflowing gas.

astro-ph.GA

A DESI Calibration of the [O II]--[S II] Electron-density Offset in Integrated Star-forming Galaxies

The [O II]$\lambda\lambda3726,3729$ and [S II]$\lambda\lambda6716,6731$ doublets are widely used as low-ionization electron-density diagnostics in galaxy spectra and are often treated as interchangeable when only one of them is accessible. We test this assumption using the DESI DR1 Emission Line Catalog. For star-forming galaxies with fiducial emission lines, [O II] yields systematically higher electron densities than [S II], with a median offset of 0.228 dex. The binned median calibration is $\log n_e({\rm OII})=(0.752^{+0.182}_{-0.097})\log n_e({\rm SII}) +(0.832^{+0.231}_{-0.422})$. The offset is larger in galaxies with higher stellar mass, H$\alpha$ star-formation rate, dust attenuation, and $N2\equiv\log([{\rm NII}]\lambda6583/{\rm H}\alpha)$, an empirical gas-phase metallicity proxy, and smaller in galaxies with higher \(\log O_{32}\equiv\log\{[{\rm OIII}]\lambda5007/ ([{\rm OII}]\lambda3726+\lambda3729)\}\), an ionization proxy; no significant trend is found with specific star-formation rate. These trends are consistent with [O II] and [S II] sampling different low-ionization gas phases in integrated spectra, with [S II] more strongly weighted toward lower-density diffuse or outer gas. Our results show that [O II]- and [S II]-based densities should not be mixed without empirical calibration in studies of ISM pressure, nebular density, and their evolution across galaxy samples and redshift.

astro-ph.GA

FiCA: Feed-forward instant Gaussian Codec Avatars from a Single Portrait Image

We introduce FiCA, a Feed-forward, instant Gaussian Codec Avatar generation pipeline that creates lifelike avatars from a single portrait image. Generating a photorealistic and drivable avatar from just a single image is significantly challenging due to the limited visual information available to accurately infer the 3D appearance and geometry of human heads. To address this, we develop a novel system that combines human-centric vision foundation models with a diffusion model. This system is designed to fully exploit partial visual observations to generate lifelike human avatars. Our proposed diffusion model learns a generative mapping from these partial observations to complete and authentic 3D mesh reconstruction. Additionally, we introduce a feed-forward mesh refinement network that enhances the fidelity and identity preservation of the generated avatars, eliminating the need for person-specific test-time optimization. By leveraging a universal prior model that decodes a generated mesh into a set of 3D Gaussians, we generate a photorealistic 3D Gaussian avatar, capable of being driven with novel expressions in real-time. Our experiments demonstrate that the avatars generated by our feed-forward approach faithfully represent diverse identities and surpass the visual quality of avatars produced by recent competing methods.

cs.CV

Understanding the Behaviors of Environment-aware Information Retrieval

Recent retrieval-augmented generation (RAG) approaches have demonstrated strong capability in handling complex queries, yet current research overlooks a critical challenge: different retrievers require fundamentally different query formulation strategies for optimal performance. In this work, we present the first systematic analysis of how LLMs can learn to adapt their query formulation strategies for different retrievers via reinforcement learning (RL). Our empirical study reveals that RL effectively teaches an LLM to tailor its queries to specific retriever characteristics. We discover that different retrievers exhibit surprisingly distinct optimal query styles (e.g., descriptive vs. question-like), suggesting strategies learned for one retriever ineffective for another. We further show that performance can be enhanced by incorporating retriever-specific human guidance and by scaling model size. To facilitate learning over multi-retrieval-step trajectories, we introduce a branching-based rollout technique that improves training stability. Our work provides the first empirical evidence and actionable insights for building truly retriever-aware RAG systems. Code and resources are available at https://github.com/LCO-Embedding/Envs-aware-Information-Retrieval.

cs.CL

Decouple before Integration: Test-time Synthesis of SFT and RLVR Task Vectors

SFT and RLVR represent two fundamental yet distinct paradigms for LLM post-training, each excelling in distinct dimensions. SFT expands knowledge breadth while RLVR enhances reasoning depth. Yet integrating these complementary strengths remains a formidable challenge. Sequential training can cause catastrophic forgetting, and joint optimization often suffers from severe gradient conflicts. We analyze SFT and RLVR through the lens of task vectors and reveal three structural properties behind these failures: a 30* magnitude disparity, 45* sign interference, and heterogeneous module-wise update distributions. These findings show SFT and RLVR are difficult to integrate directly, but they also suggest that the two paradigms modify partly complementary components of the model. Motivated by these observations, we propose Decoupled Test-time Synthesis (DoTS), a post-hoc framework allows SFT and RLVR checkpoints to be trained independently and synthesizes their capabilities only at inference time via task vector arithmetic, without updating model parameters. To reduce interference, DOTS applies selective sparsification with norm-preserving rescaling. It then uses Bayesian optimization on a small set of unlabeled queries to search for combination coefficients on the Pareto frontier of consistency and perplexity. Empirically, \ours matches or exceeds the performance of training-based SFT--RLVR integration methods across multiple mathematical reasoning benchmarks, incurring only $\sim$3\% of the computational cost. When applied to stronger post-trained checkpoints, DOTS surpasses SOTA models and generalizes to out-of-domain benchmarks without re-tuning. Code is available at https://github.com/chaohaoyuan/DoTS.

cs.LG

Agentic Fusion of Large Atomic and Language Models to Accelerate Superconductor Discovery

Artificial intelligence has accelerated materials discovery through high-throughput prediction and generation, yet the decision problem remains a formidable bottleneck. While current AI systems readily propose millions of candidates, navigating the decision regarding a viable experimental target requires resolving multi-dimensional judgments across atomic-scale numerical computation and high-level semantic reasoning. Here we present ElementsClaw, an agentic framework for materials discovery that orchestrates a suite of Large Atomic Model (LAM) tools finetuned from our proposed 1-billion-parameter model Elements for numerical computation, while leveraging Large Language Models (LLMs) for semantic reasoning. Applied to superconductors, ElementsClaw rediscovers 66 experimentally verified superconductors that are absent from the standard SuperCon3D database. Scaling to 2.4 million equilibrium crystals, ElementsClaw identifies 68,000 high-confidence candidates in just 28 GPU hours (https://developer.damo-academy.com/material), expanding known superconducting space by orders of magnitude compared to datasets curated over decades. Guided by the agent's reasoning, we experimentally synthesize and verify four novel superconductors: the motif-guided Zr$_3$ScRe$_8$ ($T_c$ = 6.5 K), the de novo generated HfZrRe$_4$ ($T_c$ = 5.9 K), the structurally reinterpreted Zr$_4$VRe$_7$ ($T_c$ = 3.5 K), and the database-latent Hf$_{21}$Re$_{25}$ ($T_c$ = 2.5 K). Together, our results establish a knowledge integrated, autonomously orchestrated, and experimentally grounded paradigm for materials discovery.

cs.LG

Large-scale Codec Avatars: The Unreasonable Effectiveness of Large-scale Avatar Pretraining

High-quality 3D avatar modeling faces a critical trade-off between fidelity and generalization. On the one hand, multi-view studio data enables high-fidelity modeling of humans with precise control over expressions and poses, but it struggles to generalize to real-world data due to limited scale and the domain gap between the studio environment and the real world. On the other hand, recent large-scale avatar models trained on millions of in-the-wild samples show promise for generalization across a wide range of identities, yet the resulting avatars are often of low-quality due to inherent 3D ambiguities. To address this, we present Large-Scale Codec Avatars (LCA), a high-fidelity, full-body 3D avatar model that generalizes to world-scale populations in a feedforward manner, enabling efficient inference. Inspired by the success of large language models and vision foundation models, we present, for the first time, a pre/post-training paradigm for 3D avatar modeling at scale: we pretrain on 1M in-the-wild videos to learn broad priors over appearance and geometry, then post-train on high-quality curated data to enhance expressivity and fidelity. LCA generalizes across hair styles, clothing, and demographics while providing precise, fine-grained facial expressions and finger-level articulation control, with strong identity preservation. Notably, we observe emergent generalization to relightability and loose garment support to unconstrained inputs, and zero-shot robustness to stylized imagery, despite the absence of direct supervision.

cs.CV

Lingshu-Cell: A generative cellular world model for transcriptome modeling toward virtual cells

Modeling cellular states and predicting their responses to perturbations are central challenges in computational biology and the development of virtual cells. Existing foundation models for single-cell transcriptomics provide powerful static representations, but they do not explicitly model the distribution of cellular states for generative simulation. Here, we introduce Lingshu-Cell, a masked discrete diffusion model that learns transcriptomic state distributions and supports conditional simulation under perturbation. By operating directly in a discrete token space that is compatible with the sparse, non-sequential nature of single-cell transcriptomic data, Lingshu-Cell captures complex transcriptome-wide expression dependencies across approximately 18,000 genes without relying on prior gene selection, such as filtering by high variability or ranking by expression level. Across diverse tissues and species, Lingshu-Cell accurately reproduces transcriptomic distributions, marker-gene expression patterns and cell-subtype proportions, demonstrating its ability to capture complex cellular heterogeneity. Moreover, by jointly embedding cell type or donor identity with perturbation, Lingshu-Cell can predict whole-transcriptome expression changes for novel combinations of identity and perturbation. It achieves leading performance on the Virtual Cell Challenge H1 genetic perturbation benchmark and in predicting cytokine-induced responses in human PBMCs. Together, these results establish Lingshu-Cell as a flexible cellular world model for in silico simulation of cell states and perturbation responses, laying the foundation for a new paradigm in biological discovery and perturbation screening.

q-bio.QM

SDSS-IV MaNGA: Distinct Structural Growth and Star Formation in Low and High Surface Brightness Disks

We analyze a clean sample of 1,118 late-type, face-on galaxies without AGN contamination from the MaNGA survey. Their photometric structures are quantified via two-component (bulge+disk) decompositions on deep $g$-band images from the DESI Legacy Survey. Using a disk central surface brightness of $\mu_{\rm 0,d,cor}$(g) = 22 $\pm$ 0.3 mag arcsec$^{-2}$ (corrected for inclination and cosmic dimming) as the classification threshold, we identify 159 low surface brightness (LSB) galaxies, 388 LSB candidates, and 571 high surface brightness (HSB) galaxies. LSB galaxies are predominantly low-mass ($M_\ast < 3 \times 10^{10}$ M$_\odot$), exhibiting 29\% larger effective radii, 15\% lower star formation rates (SFRs), and 12\% reduced gas-phase metallicities than HSB counterparts at comparable masses. These differences cause systematic offsets from standard scaling relations. Despite comparable gas content, LSB galaxies host older stellar populations, longer gas depletion times, and less efficient star formation. Spatially resolved analyses further reveal that LSB galaxies display centrally suppressed $\Sigma_{\rm SFR}$, flatter SFR gradients, and rising specific SFR profiles toward their outskirts. Together with steeper negative metallicity gradients, these trends suggest ongoing gas accretion fueling outer-disk star formation. Consistently, the outer regions of LSB galaxies exhibit stronger H$\delta_A$ absorption and lower D$_n$4000 indices, indicating fading A-star populations. Moreover, LSB galaxies show lower $\Sigma_{\ast}$ across all $R/R_e$ and more centrally depleted stellar mass profiles on an absolute radial scale, compared with HSB and large-size star-forming galaxies. Collectively, LSB galaxies represent a distinct population with slow evolution, inefficient star formation, and continued susceptibility to late-time gas accretion and peripheral star formation.

astro-ph.GA

Reference-guided Policy Optimization for Molecular Optimization via LLM Reasoning

Large language models (LLMs) benefit substantially from supervised fine-tuning (SFT) and reinforcement learning with verifiable rewards (RLVR) in reasoning tasks. However, these recipes perform poorly in instruction-based molecular optimization, where each data point typically provides only a single optimized reference molecule and no step-by-step optimization trajectory. We reveal that answer-only SFT on the reference molecules collapses reasoning, and RLVR provides sparse feedback under similarity constraints due to the model's lack of effective exploration, which slows learning and limits optimization. To encourage the exploration of new molecules while balancing the exploitation of the reference molecules, we introduce Reference-guided Policy Optimization (RePO), an optimization approach that learns from reference molecules without requiring trajectory data. At each update, RePO samples candidate molecules with their intermediate reasoning trajectories from the model and trains the model using verifiable rewards that measure property satisfaction under similarity constraints in an RL manner. Meanwhile, it applies reference guidance by keeping the policy's intermediate reasoning trajectory as context and training only the answer in a supervised manner. Together, the RL term promotes exploration, while the guidance term mitigates reward sparsity and stabilizes training by grounding outputs to references when many valid molecular edits exist. Across molecular optimization benchmarks, RePO consistently outperforms SFT and RLVR baselines (e.g., GRPO), achieving improvements on the optimization metric (Success Rate $\times$ Similarity), improving balance across competing objectives, and generalizing better to unseen instruction styles. Our code is publicly available at https://github.com/tmlr-group/RePO.

cs.LG

IBCircuit: Towards Holistic Circuit Discovery with Information Bottleneck

Circuit discovery has recently attracted attention as a potential research direction to explain the non-trivial behaviors of language models. It aims to find the computational subgraphs, also known as circuits, within the model that are responsible for solving specific tasks. However, most existing studies overlook the holistic nature of these circuits and require designing specific corrupted activations for different tasks, which is inaccurate and inefficient. In this work, we propose an end-to-end approach based on the principle of Information Bottleneck, called IBCircuit, to identify informative circuits holistically. IBCircuit is an optimization framework for holistic circuit discovery and can be applied to any given task without tediously corrupted activation design. In both the Indirect Object Identification (IOI) and Greater-Than tasks, IBCircuit identifies more faithful and minimal circuits in terms of critical node components and edge components compared to recent related work.

cs.LG

The First Systematic Survey of Stellar Halos in High-Inclination Galaxies Reveals Unusually Quiescent Merger Histories of Nearby Galaxies

Stellar halos are the only major stellar component of disk galaxies that lack systematic observational characterization, yet they encode critical information about galaxy merger histories. We present the first systematic census of stellar halos in a large, flux-limited sample of 169 high-inclination central galaxies with stellar masses 7.3 <= log Mstar/Msun <= 11.0 and redshift z < 0.1, using HSC-SSP Deep optical images. Stellar halos are detected in 93 galaxies, primarily through their low isophotal ellipticities in the outskirts, improving upon conventional methods of stellar halo identification. The halo detection rate reaches ~ 50% at log Mstar/Msun > 9.9 and >= 70% for Milky Way (MW)-mass galaxies. We derive halo surface brightness profiles, colors, and masses, finding that stellar halos generally follow power-law radial profiles. Higher-mass galaxies, on average, exhibit smaller power-law indices and larger halo mass fractions, indicating more extended halos and more active merger histories. A significant stellar halo color-mass correlation, driven mainly by the mass-metallicity relation, suggests dominance by a few massive accretion events. MW-mass galaxies have a median stellar halo fraction of 10% +/- 5%. Among nearby galaxies with halo measurements within 25 Mpc, two thirds (including the MW) lie below the mean stellar halo fraction-galaxy mass relation. Overall, the nearby galaxies show a median halo deficit of ~ 0.3 dex, implying unusually quiescent merger histories. We show that this deficit follows a broader trend in which typical halo fractions increase with heliocentric distance, tracking the gradual rise in matter density toward the cosmic average by z <= 0.07.

astro-ph.GA

Precedent-Informed Reasoning: Mitigating Overthinking in Large Reasoning Models via Test-Time Precedent Learning

Reasoning in Large Language Models (LLMs) often suffers from inefficient long chain-of-thought traces with redundant self-exploration and validation, which inflate computational costs and even degrade performance. Inspired by human reasoning patterns where people solve new problems by leveraging past related cases to constrain search spaces and reduce trial-and-error, we propose Precedent Informed Reasoning (PIR) transforming LRMs'reasoning paradigm from exhaustive self-exploration to guided learning from precedents. PIR addresses two key challenges: what precedents to adopt and how to utilize them. First, Adaptive Precedent Selection (APS) constructs, for each question and LRM, a compact set of precedents that are both semantically related and informative for the model. It ranks examples by a joint score with semantic similarity and model perplexity, then adapts the amount of precedents to maximize perplexity reduction. Second, Test-time Experience Internalization (TEI) is treated as the test-time learning on precedent-informed instruction, updating lightweight adapters to internalize solution patterns and use them as a prior during subsequent reasoning. Experiments across mathematical reasoning, scientific QA, and code generation demonstrate that PIR consistently shortens reasoning traces while maintaining or improving final accuracy across LLMs, yielding outstanding accuracy-efficiency trade-offs.

cs.AI