Searcharxiv⌕ Search

arXiv subjects

Yu Rong

Publications and source records attributed to Yu Rong.

At least 37 records · Page 2Linked to original sources

Large-scale Codec Avatars: The Unreasonable Effectiveness of Large-scale Avatar Pretraining

High-quality 3D avatar modeling faces a critical trade-off between fidelity and generalization. On the one hand, multi-view studio data enables high-fidelity modeling of humans with precise control over expressions and poses, but it struggles to generalize to real-world data due to limited scale and the domain gap between the studio environment and the real world. On the other hand, recent large-scale avatar models trained on millions of in-the-wild samples show promise for generalization across a wide range of identities, yet the resulting avatars are often of low-quality due to inherent 3D ambiguities. To address this, we present Large-Scale Codec Avatars (LCA), a high-fidelity, full-body 3D avatar model that generalizes to world-scale populations in a feedforward manner, enabling efficient inference. Inspired by the success of large language models and vision foundation models, we present, for the first time, a pre/post-training paradigm for 3D avatar modeling at scale: we pretrain on 1M in-the-wild videos to learn broad priors over appearance and geometry, then post-train on high-quality curated data to enhance expressivity and fidelity. LCA generalizes across hair styles, clothing, and demographics while providing precise, fine-grained facial expressions and finger-level articulation control, with strong identity preservation. Notably, we observe emergent generalization to relightability and loose garment support to unconstrained inputs, and zero-shot robustness to stylized imagery, despite the absence of direct supervision.

cs.CV↗

Halo Spin Depends on The Distance to Cosmic Filament

We employ a semi-analytical methodology to estimate the dark matter halo spin of HI-rich galaxies in the Arecibo Legacy Fast Alfa Survey and investigate the relationship between halo spin and the proximity of galaxies to cosmic filaments. We exclude galaxies with low HI signal-to-noise ratios, those potentially influenced by velocity dispersions, and those affiliated with galaxy clusters/groups. Additionally, we apply a mass-weighting technique to ensure consistent mass distribution across galaxy samples at varying distances from filaments. Our analysis reveals, for the first time, a subtle yet statistically significant correlation between halo spin and filament distance in observational data, indicating higher spins closer to filaments. This suggests that the tidal forces exerted by filaments may impact the spin of dark matter halos.

astro-ph.GA↗

Lingshu-Cell: A generative cellular world model for transcriptome modeling toward virtual cells

Modeling cellular states and predicting their responses to perturbations are central challenges in computational biology and the development of virtual cells. Existing foundation models for single-cell transcriptomics provide powerful static representations, but they do not explicitly model the distribution of cellular states for generative simulation. Here, we introduce Lingshu-Cell, a masked discrete diffusion model that learns transcriptomic state distributions and supports conditional simulation under perturbation. By operating directly in a discrete token space that is compatible with the sparse, non-sequential nature of single-cell transcriptomic data, Lingshu-Cell captures complex transcriptome-wide expression dependencies across approximately 18,000 genes without relying on prior gene selection, such as filtering by high variability or ranking by expression level. Across diverse tissues and species, Lingshu-Cell accurately reproduces transcriptomic distributions, marker-gene expression patterns and cell-subtype proportions, demonstrating its ability to capture complex cellular heterogeneity. Moreover, by jointly embedding cell type or donor identity with perturbation, Lingshu-Cell can predict whole-transcriptome expression changes for novel combinations of identity and perturbation. It achieves leading performance on the Virtual Cell Challenge H1 genetic perturbation benchmark and in predicting cytokine-induced responses in human PBMCs. Together, these results establish Lingshu-Cell as a flexible cellular world model for in silico simulation of cell states and perturbation responses, laying the foundation for a new paradigm in biological discovery and perturbation screening.

q-bio.QM↗

SDSS-IV MaNGA: Distinct Structural Growth and Star Formation in Low and High Surface Brightness Disks

We analyze a clean sample of 1,118 late-type, face-on galaxies without AGN contamination from the MaNGA survey. Their photometric structures are quantified via two-component (bulge+disk) decompositions on deep $g$-band images from the DESI Legacy Survey. Using a disk central surface brightness of $μ_{\rm 0,d,cor}$(g) = 22 $\pm$ 0.3 mag arcsec$^{-2}$ (corrected for inclination and cosmic dimming) as the classification threshold, we identify 159 low surface brightness (LSB) galaxies, 388 LSB candidates, and 571 high surface brightness (HSB) galaxies. LSB galaxies are predominantly low-mass ($M_\ast < 3 \times 10^{10}$ M$_\odot$), exhibiting 29\% larger effective radii, 15\% lower star formation rates (SFRs), and 12\% reduced gas-phase metallicities than HSB counterparts at comparable masses. These differences cause systematic offsets from standard scaling relations. Despite comparable gas content, LSB galaxies host older stellar populations, longer gas depletion times, and less efficient star formation. Spatially resolved analyses further reveal that LSB galaxies display centrally suppressed $Σ_{\rm SFR}$, flatter SFR gradients, and rising specific SFR profiles toward their outskirts. Together with steeper negative metallicity gradients, these trends suggest ongoing gas accretion fueling outer-disk star formation. Consistently, the outer regions of LSB galaxies exhibit stronger H$δ_A$ absorption and lower D$_n$4000 indices, indicating fading A-star populations. Moreover, LSB galaxies show lower $Σ_{\ast}$ across all $R/R_e$ and more centrally depleted stellar mass profiles on an absolute radial scale, compared with HSB and large-size star-forming galaxies. Collectively, LSB galaxies represent a distinct population with slow evolution, inefficient star formation, and continued susceptibility to late-time gas accretion and peripheral star formation.

astro-ph.GA↗

Frequency Autoregressive Image Generation with Continuous Tokens

Autoregressive (AR) models for image generation typically adopt a two-stage paradigm of vector quantization and raster-scan ``next-token prediction", inspired by its great success in language modeling. However, due to the huge modality gap, image autoregressive models may require a systematic reevaluation from two perspectives: tokenizer format and regression direction. In this paper, we introduce the frequency progressive autoregressive (\textbf{FAR}) paradigm and instantiate FAR with the continuous tokenizer. Specifically, we identify spectral dependency as the desirable regression direction for FAR, wherein higher-frequency components build upon the lower one to progressively construct a complete image. This design seamlessly fits the causality requirement for autoregressive models and preserves the unique spatial locality of image data. Besides, we delve into the integration of FAR and the continuous tokenizer, introducing a series of techniques to address optimization challenges and improve the efficiency of training and inference processes. We demonstrate the efficacy of FAR through comprehensive experiments on the ImageNet dataset and verify its potential on text-to-image generation.

cs.CV↗

Reference-guided Policy Optimization for Molecular Optimization via LLM Reasoning

Large language models (LLMs) benefit substantially from supervised fine-tuning (SFT) and reinforcement learning with verifiable rewards (RLVR) in reasoning tasks. However, these recipes perform poorly in instruction-based molecular optimization, where each data point typically provides only a single optimized reference molecule and no step-by-step optimization trajectory. We reveal that answer-only SFT on the reference molecules collapses reasoning, and RLVR provides sparse feedback under similarity constraints due to the model's lack of effective exploration, which slows learning and limits optimization. To encourage the exploration of new molecules while balancing the exploitation of the reference molecules, we introduce Reference-guided Policy Optimization (RePO), an optimization approach that learns from reference molecules without requiring trajectory data. At each update, RePO samples candidate molecules with their intermediate reasoning trajectories from the model and trains the model using verifiable rewards that measure property satisfaction under similarity constraints in an RL manner. Meanwhile, it applies reference guidance by keeping the policy's intermediate reasoning trajectory as context and training only the answer in a supervised manner. Together, the RL term promotes exploration, while the guidance term mitigates reward sparsity and stabilizes training by grounding outputs to references when many valid molecular edits exist. Across molecular optimization benchmarks, RePO consistently outperforms SFT and RLVR baselines (e.g., GRPO), achieving improvements on the optimization metric (Success Rate $\times$ Similarity), improving balance across competing objectives, and generalizing better to unseen instruction styles. Our code is publicly available at https://github.com/tmlr-group/RePO.

cs.LG↗

IBCircuit: Towards Holistic Circuit Discovery with Information Bottleneck

Circuit discovery has recently attracted attention as a potential research direction to explain the non-trivial behaviors of language models. It aims to find the computational subgraphs, also known as circuits, within the model that are responsible for solving specific tasks. However, most existing studies overlook the holistic nature of these circuits and require designing specific corrupted activations for different tasks, which is inaccurate and inefficient. In this work, we propose an end-to-end approach based on the principle of Information Bottleneck, called IBCircuit, to identify informative circuits holistically. IBCircuit is an optimization framework for holistic circuit discovery and can be applied to any given task without tediously corrupted activation design. In both the Indirect Object Identification (IOI) and Greater-Than tasks, IBCircuit identifies more faithful and minimal circuits in terms of critical node components and edge components compared to recent related work.

cs.LG↗

The First Systematic Survey of Stellar Halos in High-Inclination Galaxies Reveals Unusually Quiescent Merger Histories of Nearby Galaxies

Stellar halos are the only major stellar component of disk galaxies that lack systematic observational characterization, yet they encode critical information about galaxy merger histories. We present the first systematic census of stellar halos in a large, flux-limited sample of 169 high-inclination central galaxies with stellar masses 7.3 <= log Mstar/Msun <= 11.0 and redshift z < 0.1, using HSC-SSP Deep optical images. Stellar halos are detected in 93 galaxies, primarily through their low isophotal ellipticities in the outskirts, improving upon conventional methods of stellar halo identification. The halo detection rate reaches ~ 50% at log Mstar/Msun > 9.9 and >= 70% for Milky Way (MW)-mass galaxies. We derive halo surface brightness profiles, colors, and masses, finding that stellar halos generally follow power-law radial profiles. Higher-mass galaxies, on average, exhibit smaller power-law indices and larger halo mass fractions, indicating more extended halos and more active merger histories. A significant stellar halo color-mass correlation, driven mainly by the mass-metallicity relation, suggests dominance by a few massive accretion events. MW-mass galaxies have a median stellar halo fraction of 10% +/- 5%. Among nearby galaxies with halo measurements within 25 Mpc, two thirds (including the MW) lie below the mean stellar halo fraction-galaxy mass relation. Overall, the nearby galaxies show a median halo deficit of ~ 0.3 dex, implying unusually quiescent merger histories. We show that this deficit follows a broader trend in which typical halo fractions increase with heliocentric distance, tracking the gradual rise in matter density toward the cosmic average by z <= 0.07.

astro-ph.GA↗

Precedent-Informed Reasoning: Mitigating Overthinking in Large Reasoning Models via Test-Time Precedent Learning

Reasoning in Large Language Models (LLMs) often suffers from inefficient long chain-of-thought traces with redundant self-exploration and validation, which inflate computational costs and even degrade performance. Inspired by human reasoning patterns where people solve new problems by leveraging past related cases to constrain search spaces and reduce trial-and-error, we propose Precedent Informed Reasoning (PIR) transforming LRMs'reasoning paradigm from exhaustive self-exploration to guided learning from precedents. PIR addresses two key challenges: what precedents to adopt and how to utilize them. First, Adaptive Precedent Selection (APS) constructs, for each question and LRM, a compact set of precedents that are both semantically related and informative for the model. It ranks examples by a joint score with semantic similarity and model perplexity, then adapts the amount of precedents to maximize perplexity reduction. Second, Test-time Experience Internalization (TEI) is treated as the test-time learning on precedent-informed instruction, updating lightweight adapters to internalize solution patterns and use them as a prior during subsequent reasoning. Experiments across mathematical reasoning, scientific QA, and code generation demonstrate that PIR consistently shortens reasoning traces while maintaining or improving final accuracy across LLMs, yielding outstanding accuracy-efficiency trade-offs.

cs.AI↗

Towards Agentic Intelligence for Materials Science

The convergence of artificial intelligence and materials science presents a transformative opportunity, but achieving true acceleration in discovery requires moving beyond task-isolated, fine-tuned models toward agentic systems that plan, act, and learn across the full discovery loop. This survey advances a unique pipeline-centric view that spans from corpus curation and pretraining, through domain adaptation and instruction tuning, to goal-conditioned agents interfacing with simulation and experimental platforms. Unlike prior reviews, we treat the entire process as an end-to-end system to be optimized for tangible discovery outcomes rather than proxy benchmarks. This perspective allows us to trace how upstream design choices-such as data curation and training objectives-can be aligned with downstream experimental success through effective credit assignment. To bridge communities and establish a shared frame of reference, we first present an integrated lens that aligns terminology, evaluation, and workflow stages across AI and materials science. We then analyze the field through two focused lenses: From the AI perspective, the survey details LLM strengths in pattern recognition, predictive analytics, and natural language processing for literature mining, materials characterization, and property prediction; from the materials science perspective, it highlights applications in materials design, process optimization, and the acceleration of computational workflows via integration with external tools (e.g., DFT, robotic labs). Finally, we contrast passive, reactive approaches with agentic design, cataloging current contributions while motivating systems that pursue long-horizon goals with autonomy, memory, and tool use. This survey charts a practical roadmap towards autonomous, safety-aware LLM agents aimed at discovering novel and useful materials.

cond-mat.mtrl-sci↗

Updated Metallicity Diagnostics for Precision Oxygen Abundance Measurements in High-redshift Galaxies with JWST

Recent work has demonstrated that widely used strong-line oxygen abundance indicators, such as O3N2, $\rm R23$, and $\widehat{\rm R}$, suffer from large uncertainties when applied to high-redshift galaxies. We show that this loss of precision primarily arises because, at fixed \Oabund, galaxies span a wide dynamic range in ionization parameter and nitrogen enrichment. Here we develop updated indicators that explicitly incorporate both effects via the proxies O32 and N2O2. We define ${\rm R}_{\rm u}\equiv \rm R23+α_1 O32+α_2 N2O2$, $\widehat{\rm R}_{\rm u}\equiv \rm \widehat{R}+β_1 O32+β_2 N2O2$, and ${\rm O}_{\rm u}\equiv \rm O3N2+γ_1 O32+γ_2 N2O2$, and calibrate \Oabund~as low-order polynomials in each composite indicator. Applied to a JWST sample with $T_{\rm e}$-method abundances, the updated indicators substantially tighten the correlations with \Oabund, boosting adjusted coefficients of determination from $\mathbb{R}^2\lesssim 0$ (classical indicators) to $\mathbb{R}^2\gtrsim 0.5$ for the full sample and to $\sim 0.7$ at $z>2$. The residuals reveal a redshift evolution in the mapping between \Oabund, strong lines, ionization, and nitrogen enrichment, with a pivotal turning point near the cosmic noon ($z\sim 2$). Our calibrations provide a practical, physically grounded path to precise metallicity measurements in the JWST era and a firmer basis for quantifying early chemical enrichment and feedback.

astro-ph.GA↗

ParaFormer: A Generalized PageRank Graph Transformer for Graph Representation Learning

Graph Transformers (GTs) have emerged as a promising graph learning tool, leveraging their all-pair connected property to effectively capture global information. To address the over-smoothing problem in deep GNNs, global attention was initially introduced, eliminating the necessity for using deep GNNs. However, through empirical and theoretical analysis, we verify that the introduced global attention exhibits severe over-smoothing, causing node representations to become indistinguishable due to its inherent low-pass filtering. This effect is even stronger than that observed in GNNs. To mitigate this, we propose PageRank Transformer (ParaFormer), which features a PageRank-enhanced attention module designed to mimic the behavior of deep Transformers. We theoretically and empirically demonstrate that ParaFormer mitigates over-smoothing by functioning as an adaptive-pass filter. Experiments show that ParaFormer achieves consistent performance improvements across both node classification and graph classification tasks on 11 datasets ranging from thousands to millions of nodes, validating its efficacy. The supplementary material, including code and appendix, can be found in https://github.com/chaohaoyuan/ParaFormer.

cs.LG↗

From Macro to Micro: Benchmarking Microscopic Spatial Intelligence on Molecules via Vision-Language Models

This paper introduces the concept of Microscopic Spatial Intelligence (MiSI), the capability to perceive and reason about the spatial relationships of invisible microscopic entities, which is fundamental to scientific discovery. To assess the potential of Vision-Language Models (VLMs) in this domain, we propose a systematic benchmark framework MiSI-Bench. This framework features over 163,000 question-answer pairs and 587,000 images derived from approximately 4,000 molecular structures, covering nine complementary tasks that evaluate abilities ranging from elementary spatial transformations to complex relational identifications. Experimental results reveal that current state-of-the-art VLMs perform significantly below human level on this benchmark. However, a fine-tuned 7B model demonstrates substantial potential, even surpassing humans in spatial transformation tasks, while its poor performance in scientifically-grounded tasks like hydrogen bond recognition underscores the necessity of integrating explicit domain knowledge for progress toward scientific AGI. The datasets are available at https://huggingface.co/datasets/zongzhao/MiSI-bench.

cs.CV↗

EP241217a: a likely Type II GRB with an achromatic bump at z = 4.59

EP241217a is an X-ray transient detected by the Einstein Probe (EP) lasting for about 100 seconds and without accompanying $γ$-ray detection. The optical spectroscopy reveals the redshift of EP241217a is 4.59. By combining the $γ$-ray upper limit provided by GECAM-C, there is a considerable possibility that EP241217a is a typical Type II gamma-ray burst (GRB), but it is fainter than the detection threshold of any available $γ$-ray monitors (i.e., $E_{γ,{\rm iso}}\lesssim10^{53}$ erg). The X-ray light curve exhibits a plateau lasting for $\sim5\times10^4$ seconds. However, the joint analysis with optical data suggests the presence of an achromatic bump peaking at $\sim3\times10^4$ s after the trigger, indicating the actual duration of the X-ray plateau may be significantly shorter than it appears. To interpret the achromatic bump, we adopt the scenario of a mildly relativistic jet coasting in a wind-like medium and encountering a rapid density enhancement of the circumburst medium, which is likely induced by the the interaction of the progenitor's stellar wind and the interstellar medium. However, this model cannot fully explain observed data, and some issues do exist, e.g., the observed spectrum is harder than the model prediction. Consequently, we conclude that the scenario of a mildly relativistic jet coasting in the wind-like medium cannot explain all observed features of EP241217a. In addition, some alternative models commonly invoked to explain X-ray plateaus are discussed, but there are more or less issues when they are applied to EP241217a. Therefore, further theoretical modeling is encouraged to explore the origin of EP241217a.

astro-ph.HE↗

Quantifying the Relationship Between Galaxy Specific Star Formation Rate And Halo Spin For Star-forming Galaxies

Utilizing ALFALFA HI data, we investigate the relationship between specific star formation rate (sSFR) and halo spin across various star-forming galaxies. Our analysis reveals weak yet statistically significant positive correlation between sSFR and halo spin, irrespective of the galactic environment. This trend suggests that galaxies with higher spin parameters tend to host dynamically colder, gas-rich disks, sustaining elevated gas surface densities and prolonged star formation. These findings align with theoretical expectations of angular momentum-regulated gas accretion but highlight discrepancies with cosmological simulations, underscoring unresolved challenges in modeling baryonic feedback and star formation efficiency.

astro-ph.GA↗

The Cosmic Dance: Observational Detection of Coherent Spin in Galaxy Clusters

The spin of galaxy clusters encodes key information about their formation, dynamics, and the influence of large-scale structure. However, whether clusters possess statistically significant spin and how to measure it observationally remain open questions. Here, we present the first observational, statistical detection of coherent spin in galaxy clusters, using two samples of 2,170 and 1,329 systems with $M > 10^{14}\,M_\odot$, selected from two publicly available group catalogs (\citet{2017A&A...602A.100T} and \citet{2012ApJ...752...41Y}) constructed with two different algorithms and but both based primarily on SDSS galaxies. Cluster spin is quantified by identifying the orientation in the projected plane that maximizes the redshift difference ($ΔZ_{\rm max}$) between member galaxies in two regions divided by a trial axis. We find compelling statistical evidence for coherent rotation, as the observed $ΔZ_{\rm max}$ distribution departs markedly from the randomized controls, exhibiting pronounced deviations near $380\,\mathrm{km\,s^{-1}}$. Stacked visualizations confirm the spatial segregation of redshifted and blueshifted galaxies across the rotation axis. The radial profile of the rotational velocity indicates that it increases as a function of radius. The cluster rotation speed increases with mass, from $\sim360~\mathrm{km\,s}^{-1}$ at $10^{14} M_\odot$ to $\sim693~\mathrm{km\,s}^{-1}$ at $10^{15} M_\odot$. Additionally, cluster spin tends to align parallel with the central galaxy spin and perpendicular to the nearest cosmic filament, particularly in richer systems. These results reveal significant coherent spin in galaxy clusters, shaped by both internal dynamics and large-scale structure.

astro-ph.GA↗

The Rise of Parameter Specialization for Knowledge Storage in Large Language Models

Over time, a growing wave of large language models from various series has been introduced to the community. Researchers are striving to maximize the performance of language models with constrained parameter sizes. However, from a microscopic perspective, there has been limited research on how to better store knowledge in model parameters, particularly within MLPs, to enable more effective utilization of this knowledge by the model. In this work, we analyze twenty publicly available open-source large language models to investigate the relationship between their strong performance and the way knowledge is stored in their corresponding MLP parameters. Our findings reveal that as language models become more advanced and demonstrate stronger knowledge capabilities, their parameters exhibit increased specialization. Specifically, parameters in the MLPs tend to be more focused on encoding similar types of knowledge. We experimentally validate that this specialized distribution of knowledge contributes to improving the efficiency of knowledge utilization in these models. Furthermore, by conducting causal training experiments, we confirm that this specialized knowledge distribution plays a critical role in improving the model's efficiency in leveraging stored knowledge.

cs.CL↗

DiffSpectra: Molecular Structure Elucidation from Spectra using Diffusion Models

Molecular structure elucidation from spectra is a fundamental challenge in molecular science. Conventional approaches rely heavily on expert interpretation and lack scalability, while retrieval-based machine learning approaches remain constrained by limited reference libraries. Generative models offer a promising alternative, yet most adopt autoregressive architectures that overlook 3D geometry and struggle to integrate diverse spectral modalities. In this work, we present DiffSpectra, a generative framework that formulates molecular structure elucidation as a conditional generation process, directly inferring 2D and 3D molecular structures from multi-modal spectra using diffusion models. Its denoising network is parameterized by the Diffusion Molecule Transformer, an SE(3)-equivariant architecture for geometric modeling, conditioned by SpecFormer, a Transformer-based spectral encoder capturing multi-modal spectral dependencies. Extensive experiments demonstrate that DiffSpectra accurately elucidates molecular structures, achieving 40.76% top-1 and 99.49% top-10 accuracy. Its performance benefits substantially from 3D geometric modeling, SpecFormer pre-training, and multi-modal conditioning. To our knowledge, DiffSpectra is the first framework that unifies multi-modal spectral reasoning and joint 2D/3D generative modeling for de novo molecular structure elucidation.

cs.LG↗