SearcharxivSearch

arXiv subjects

Yurou Liu

Publications and source records attributed to Yurou Liu.

11 recordsLinked to original sources

Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning

Reinforcement learning with verifiable rewards without human-annotated data, often referred to as zero RL, has emerged as a powerful paradigm for eliciting chain-of-thought reasoning. However, due to computational constraints, existing studies are largely restricted to small models, leaving the training dynamics and emergent capabilities at a large scale unexplored. To meaningfully explore this frontier, we aim to elicit high-quality reasoning behaviors from the model. However, we find that naive scaling often suffers from poor readability, token redundancy, and a lack of adaptive reasoning depth. To address these challenges, we present a stable and efficient training pipeline, incorporating algorithmic and system optimizations such as clipped importance sampling, training-inference ratio correction, and mixed-precision control. Our experiments offer three key findings that validate the "bitter lesson" of scaling: (1) scaling to 1T parameters significantly enhances sample efficiency and performance ceilings; (2) the training process progresses sequentially through an initial discovery phase followed by a sharpening phase; and (3) the model spontaneously develops advanced cognitive behaviors, including anthropomorphism, structured formatting, self-verification, parallel reasoning, and context anxiety, rendering hand-crafted heuristics redundant. Evaluated on seven mathematical benchmarks, Ring-2.5-1T-Zero achieves competitive performance. Additionally, to assess CoT quality beyond final-answer correctness, we propose a structured evaluation framework across three dimensions: comprehensibility, reproducibility, and efficiency, where our model demonstrates clear advantages in producing structured and concise reasoning traces. By sharing our observed emergent phenomena, we hope to provide the community with deeper insights into scaling behaviors, particularly at the 1-trillion scale.

cs.CL

Speaking the Language of Science: Toward a General-Purpose Generative Foundation Model for the Natural Sciences

In this report, we present LOGOS (Language Of Generative Objects in Science), a scientific generative language model that unifies heterogeneous tasks across the natural sciences within a single autoregressive framework based on a shared scientific grammar. It encodes diverse scientific objects and their spatial interactions as token sequences over a common vocabulary. By representing spatial contact and constraint patterns as discrete tokens, the model captures complex structural interactions in a purely sequential manner, without relying on explicit coordinates or geometric neural networks. This unified representation enables a wide range of downstream tasks to be formulated consistently as next-token prediction in the same grammar space, creating strong alignment between continued multi-domain pre-training and downstream objectives. Across diverse tasks, LOGOS consistently matches or outperforms domain-specific baselines, providing preliminary evidence for the feasibility of "one model fits all" in the natural sciences. We train LOGOS models at different scales (1B, 3B, and 8B parameters) and find a consistent positive correlation between model size and performance. This suggests that the future of AI for Science (AI4S) may not lie in building an independent technical stack that is separated from large language models (LLMs). Instead, it may depend on deeply aligning scientific foundation models with LLMs through shared architectures, shared training paradigms, and shared inference infrastructure, so that LLMs can truly become a new entry point for AI4S. We release the model weights and associated resources to facilitate further research.

cs.CL

High-Eccentricity Tidal Migration Driven by Secular Chaos in Wide-Binary Systems

High-eccentricity tidal migration driven by a distant stellar companion offers a natural pathway for producing some hot Jupiters; yet, most theoretical work has relied on an idealized three-body configuration whose simplicity makes the problem especially tractable. In reality, many cold-Jupiter systems may host additional planets or substellar objects, whose interactions can dramatically alter the pathways to secularly excite extreme eccentricities. We investigate how secular chaos can drive high-eccentricity tidal migration in hierarchical ``3+1'' systems--stellar binaries hosting a planet and an additional intermediate companion orbiting the primary star. We show that the onset of secular chaos is regulated by the ratio of the von-Zeipel-Lidov-Kozai (ZLK) timescales of the inner and outer orbits $\mathcal{R}$. When $\mathcal{R}\sim 0.5-2$, most systems can undergo migration even when their mutual inclinations remain modest--below the $39.2^\circ$ critical angle for ZLK oscillations--with diffusion timescales spanning a broad range, up to thousands of inner orbit ZLK timescales. For larger mutual inclinations, secular migration operates over a much broader region of parameter space with $\mathcal{R} \sim 0.05-100$, but most evolutionary pathways become non-secular and potentially unstable--behavior recently identified as an alternative pathway to tidal migration. Our model predicts hot Jupiters in nearly polar orbits relative to both the host star's stellar equator (stellar obliquities $\sim 60^\circ-120^\circ$) and the orbits of the outer two companions. Future Gaia releases and long-term radial velocity campaigns are likely to uncover additional ``3+1'' systems, providing valuable opportunities to test this migration pathway.

astro-ph.EP

Chemistry and Isotope Ratios of Substellar Atmospheres in the $\beta$ Pictoris Young Moving Group and Vicinity

Measuring the chemical and isotopic compositions of gas giants and brown dwarfs provides insights into their formation pathways and birth environments. 2MASS J0249-0557 c is an L2-type planetary mass companion ($\sim 12 M_{\mathrm{Jup}}$) orbiting a pair of brown dwarfs in the $\beta$ Pic young moving group and vicinity. Its mass places it at the intersection of planets and brown dwarfs, making it an interesting target for constraining formation pathways at the planet-brown-dwarf boundary. Using high-resolution spectroscopic data of the planet acquired with CRIRES+ mounted on VLT, we conduct atmospheric retrieval with the radiative transfer code \texttt{petitRADTRANS} and the nested sampling tool PyMultiNest. We retrieve a C/O ratio of $0.57\pm0.01$, a metallicity of [M/H] = $0.18\pm0.05$, and a $^{12}$CO/$^{13}$CO ratio of $95^{+23}_{-17}$. We also retrieve atmospheric compositions for two benchmark brown dwarfs in the $\beta$ Pic YMG, 2MASSI J0443+0002 and SIPS J2000-7523, using CRIRES+ data and find consistent compositions. Together with 2MASS J0249-0557 c's wide separation from its host, its compositional consistency with benchmark brown dwarfs supports gravitational collapse in a star-like manner as its most likely formation mechanism. These results deliver a homogeneous comparison of three substellar members in the $\beta$ Pic YMG and vicinity. Their solar-like abundances provide a baseline for exoplanet members in the same moving group, such as $\beta$ Pic b, 51 Eri b, and AF Lep b, whose host stellar compositions are difficult to measure. Future comparisons of atmospheric compositions among this moving group offer the potential to distinguish between formation mechanisms for its planetary members.

astro-ph.EP

Probing RLVR training instability through the lens of objective-level hacking

Prolonged reinforcement learning with verifiable rewards (RLVR) has been shown to drive continuous improvements in the reasoning capabilities of large language models, but the training is often prone to instabilities, especially in Mixture-of-Experts (MoE) architectures. Training instability severely undermines model capability improvement, yet its underlying causes and mechanisms remain poorly understood. In this work, we introduce a principled framework for understanding RLVR instability through the lens of objective-level hacking. Unlike reward hacking, which arises from exploitable verifiers, objective-level hacking emerges from token-level credit misalignment and is manifested as system-level spurious signals in the optimization objective. Grounded in our framework, together with extensive experiments on a 30B MoE model, we trace the origin and formalize the mechanism behind a key pathological training dynamic in MoE models: the abnormal growth of the training-inference discrepancy, a phenomenon widely associated with instability but previously lacking a mechanistic explanation. These findings provide a concrete and causal account of the training dynamics underlying instabilities in MoE models, offering guidance for the design of stable RLVR algorithms.

cs.AI

Double Hot Jupiter Formation through Mirrored ZLK Migration in Binary Star Systems: The Case of WASP-94

To date, only a handful of binary star systems are known with at least one confirmed planet orbiting each star. Such systems, however, offer a unique perspective on the stochasticity intrinsic to planet formation and evolution -- particularly in twin binary star systems, which consist of near-equal-mass stars formed contemporaneously in the same birth environment. The WASP-94 system, which includes twin F-type stars, is a striking exemplar of such systems, containing two hot Jupiters: WASP-94 Ab is a transiting, spin-orbit misaligned giant planet with a 4-day orbital period, while WASP-94 Bb is non-transiting and has a tighter 2-day orbital period. In this work, we leverage N-body simulations to show that the current double hot Jupiter configuration of the WASP-94 system can be reproduced through mirrored von Zeipel-Lidov-Kozai migration. The upcoming Gaia astrometric data releases offer the potential to search for additional twin planetary systems, including double cold Jupiter systems that may serve as the progenitors for WASP-94-like configurations.

astro-ph.EP

Learning 3D Anisotropic Noise Distributions Improves Molecular Force Field Modeling

Coordinate denoising has emerged as a promising method for 3D molecular pretraining due to its theoretical connection to learning molecular force field. However, existing denoising methods rely on oversimplied molecular dynamics that assume atomic motions to be isotropic and homoscedastic. To address these limitations, we propose a novel denoising framework AniDS: Anisotropic Variational Autoencoder for 3D Molecular Denoising. AniDS introduces a structure-aware anisotropic noise generator that can produce atom-specific, full covariance matrices for Gaussian noise distributions to better reflect directional and structural variability in molecular systems. These covariances are derived from pairwise atomic interactions as anisotropic corrections to an isotropic base. Our design ensures that the resulting covariance matrices are symmetric, positive semi-definite, and SO(3)-equivariant, while providing greater capacity to model complex molecular dynamics. Extensive experiments show that AniDS outperforms prior isotropic and homoscedastic denoising models and other leading methods on the MD17 and OC22 benchmarks, achieving average relative improvements of 8.9% and 6.2% in force prediction accuracy. Our case study on a crystal and molecule structure shows that AniDS adaptively suppresses noise along the bonding direction, consistent with physicochemical principles. Our code is available at https://github.com/ZeroKnighting/AniDS.

cs.LG

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward

Recent advances in large reasoning models have leveraged reinforcement learning with verifiable rewards (RLVR) to improve reasoning capabilities. However, scaling these methods typically requires extensive rollout computation and large datasets, leading to high training costs and low data efficiency. To mitigate this issue, we propose DEPO, a Data-Efficient Policy Optimization pipeline that combines optimized strategies for both offline and online data selection. In the offline phase, we curate a high-quality subset of training samples based on diversity, influence, and appropriate difficulty. During online RLVR training, we introduce a sample-level explorability metric to dynamically filter samples with low exploration potential, thereby reducing substantial rollout computational costs. Furthermore, we incorporate a replay mechanism for under-explored samples to ensure adequate training, which enhances the model's final convergence performance. Experiments across five reasoning benchmarks show that DEPO consistently outperforms existing methods in both offline and online data selection scenarios. Notably, using only 20% of the training data, our approach achieves a 1.85 times speed-up on AIME24 and a 1.66 times speed-up on AIME25 compared to GRPO trained on the full dataset.

cs.LG

Collisional Fragmentation Support in TRACE

We present improved collision support for TRACE, a state-of-the-art hybrid integrator in REBOUND. TRACE now supports collisional fragmentation and can handle both removing and adding particles mid-timestep. We describe the back-end logic implemented for robust collision support, and compare TRACE's performance to other integrators including MERCURIUS on a large-N protoplanetary disk simulation with various collision prescriptions, a system which TRACE previously could not handle. TRACE matches the behavior of these integrators, while offering potentially vast speedups of over 70x. All updates described in this Note are available with the most recent public release of REBOUND.

astro-ph.EP

The Formation of Double Hot Jupiter Systems through von Zeipel-Lidov-Kozai Migration

The von Zeipel-Lidov-Kozai (ZLK) mechanism with tidal friction has been demonstrated as a promising avenue to generate hot Jupiters in stellar binary systems. Previous population studies of hot Jupiter formation have largely examined this mechanism in systems comprised of three bodies: two stars and one planet. However, because stars in a binary system form in similar environments with comparable metallicities, the formation of a single hot Jupiter in such a system may imply that the conditions are more likely met for the companion star, as well. We investigate the ZLK mechanism with tidal friction as a potential mechanism to produce double hot Jupiter systems in stellar binaries. Using N-body simulations, we characterize the evolution of two cold Jupiters, each orbiting one star in a binary system, undergoing mirrored ZLK migration. We then examine the robustness of this mechanism to asymmetries in stellar masses, planet masses, and planet orbital inclinations relative to the binary plane. We predict that, under the assumptions that (1) most hot Jupiters in binary star systems form through ZLK migration of primordially formed cold Jupiters and (2) if one star in a binary system forms a cold Jupiter, the second does as well, a comprehensive search could identify double hot Jupiters in up to ~9% of the close- to moderate- separation $a<2000$ AU) binary systems that already host a known hot Jupiter. We also argue that a blind search for ZLK-migrated double hot Jupiters should prioritize twin stellar binaries with pericenter approaches of a few hundred AU.

astro-ph.EP

Atomic and Subgraph-aware Bilateral Aggregation for Molecular Representation Learning

Molecular representation learning is a crucial task in predicting molecular properties. Molecules are often modeled as graphs where atoms and chemical bonds are represented as nodes and edges, respectively, and Graph Neural Networks (GNNs) have been commonly utilized to predict atom-related properties, such as reactivity and solubility. However, functional groups (subgraphs) are closely related to some chemical properties of molecules, such as efficacy, and metabolic properties, which cannot be solely determined by individual atoms. In this paper, we introduce a new model for molecular representation learning called the Atomic and Subgraph-aware Bilateral Aggregation (ASBA), which addresses the limitations of previous atom-wise and subgraph-wise models by incorporating both types of information. ASBA consists of two branches, one for atom-wise information and the other for subgraph-wise information. Considering existing atom-wise GNNs cannot properly extract invariant subgraph features, we propose a decomposition-polymerization GNN architecture for the subgraph-wise branch. Furthermore, we propose cooperative node-level and graph-level self-supervised learning strategies for ASBA to improve its generalization. Our method offers a more comprehensive way to learn representations for molecular property prediction and has broad potential in drug and material discovery applications. Extensive experiments have demonstrated the effectiveness of our method.

cs.LG