SearcharxivSearch

arXiv subjects

Lin Yan

Publications and source records attributed to Lin Yan.

At least 37 records · Page 2Linked to original sources

EvaLearn: Quantifying the Learning Capability and Efficiency of LLMs via Sequential Problem Solving

We introduce EvaLearn, a pioneering benchmark designed to evaluate large language models (LLMs) on their learning capability and efficiency in challenging tasks, a critical, yet underexplored aspect of model potential. EvaLearn contains 648 challenging problems across six task types, grouped into 182 sequences, each sequence dedicated to one task type. Diverging from most existing benchmarks that evaluate models in parallel, EvaLearn requires models to solve problems sequentially, allowing them to leverage the experience gained from previous solutions. EvaLearn provides five comprehensive automated metrics to evaluate models and quantify their learning capability and efficiency. We extensively benchmark nine frontier models and observe varied performance profiles: some models, such as Claude-3.7-sonnet, start with moderate initial performance but exhibit strong learning ability, while some models struggle to benefit from experience and may even show negative transfer. Moreover, we investigate model performance under two learning settings and find that instance-level rubrics and teacher-model feedback further facilitate model learning. Importantly, we observe that current LLMs with stronger static abilities do not show a clear advantage in learning capability across all tasks, highlighting that EvaLearn evaluates a new dimension of model performance. We hope EvaLearn provides a novel evaluation perspective for assessing LLM potential and understanding the gap between models and human capabilities, promoting the development of deeper and more dynamic evaluation approaches. All datasets, the automatic evaluation framework, and the results studied in this paper are available at the GitHub repository.

cs.CL

A JWST MIRI LRS Survey of 37 Massive Star-Forming Galaxies and AGN at Cosmic Noon -- Overview and First Results

We present a large spectroscopic survey with \textit{JWST}'s Mid-Infrared Instrument (MIRI) Low Resolution Spectrometer (LRS) targeting $37$ infrared-bright galaxies between $z=0.65-2.46$ with infrared luminosities $\log L_{\rm IR}/L_\odot>11.5$ and $\log M_*/M_\odot=10-11.5$. Targets were taken from a \textit{Spitzer} $24\,μ$m-selected sample with archival spectroscopy from the Infrared Spectrograph (IRS) and include a mix of star-forming galaxies and dust-obscured AGN. By combining IRS with the increased sensitivity of LRS, we expand the range of spectral features observed between $5-30\,μ$m for every galaxy in our sample. In this paper, we outline the sample selection, \textit{JWST} data reduction, 1D spectral extraction, and polycyclic aromatic hydrocarbon (PAH) feature measurements from $λ_{rest}=3.3-11.2\,μ$m. In the \textit{JWST} spectra, we detect PAH emission features at $3.3-5.3\,μ$m, as well as Paschen and Brackett lines. The $3.3\,μ$m feature can be as bright as $1\%$ of the $8-1000\,μ$m infrared luminosity and exhibits a tight correlation with the dust-obscured star-formation rate. We detect absorption features from CO gas, CO$_2$ ice, H$_2$O ice, and aliphatic dust. From the joint \textit{JWST} and \textit{Spitzer} analysis we find that the $11.3/3.3\,μ$m PAH ratios are on-average three times higher than that of local luminous, infrared galaxies. This is interpreted as evidence that the PAH grains are larger at $z\sim1-2$. The size distribution may be affected by coagulation of grains due to high gas densities and low temperatures. These conditions are supported by the observation of strong water ice absorption at $3.05\,μ$m, and can lower stellar radiative feedback as large PAHs transmit less energy per photon into the interstellar medium.

astro-ph.GA

The ALPINE-CRISTAL-JWST Survey: Revealing Less Massive Black Holes in High-Redshift Galaxies

We present a systematic search for broad-line active galactic nuclei (AGNs) in the ALPINE-CRISTAL-JWST sample of 18 star-forming galaxies ($M_\star>10^{9.5}~M_{\odot}$) at redshifts $z=4.4-5.7$. Using JWST/NIRSpec IFU, we identify 7 AGN candidates through the detection of broad \Ha\ emission lines from 33 aperture spectra centred on photometric peaks. These candidates include one highly robust AGN detection with FWHM $\sim$ 2800 \kms\ and six showing broad components with FWHM $\sim 600-1600$ \kms, with two in a merger system. We highlight that only broad-line detection is effective since these candidates uniformly lie within narrow emission-line ratio diagnostic diagrams where star-forming galaxies and AGNs overlap. The broad-line AGN fraction ranges from 5.9\% to 33\%, depending on the robustness of the candidates. Assuming that the majority are AGNs, the relatively high AGN fraction is likely due to targeting high-mass galaxies, where simulations demonstrate that broad-line detection is more feasible. Their black hole masses range from $10^6$ to $10^{7.5}~M_{\odot}$ with $0.1 \lesssim L_{\rm bol}/L_{\rm Edd}\lesssim 1$. Counter to previous JWST studies at high redshift that found overmassive black holes relative to their host galaxies, our candidates lie close to or below the local $M_{\rm BH}-M_\star$ scaling relations, thus demonstrating the effect of selection biases. This study provides new insights into AGN-host galaxy co-evolution at high redshift by identifying faint broad-line AGNs in galaxy samples, highlighting the importance of considering mass-dependent selection biases and the likelihood of a large population of AGNs being undermassive and just now being tapped by JWST.

astro-ph.GA

Risk-Sensitive RL for Alleviating Exploration Dilemmas in Large Language Models

Reinforcement Learning with Verifiable Rewards (RLVR) has proven effective for enhancing Large Language Models (LLMs) on complex reasoning tasks. However, existing methods suffer from an exploration dilemma: the sharply peaked initial policies of pre-trained LLMs confine standard RL algorithms to a narrow set of solutions, boosting single-solution accuracy (pass@1) but suppressing solution diversity and multi-solution performance (pass@k). As a result, RLVR often distills existing capabilities rather than discovering new reasoning strategies. To overcome this, we introduce a Risk-Sensitive Reinforcement Learning framework. Our approach employs a risk-seeking objective that interpolates between mean and maximum rewards, leading to a novel algorithm, Risk-Sensitive GRPO (RS-GRPO), which drives deeper exploration by amplifying learning from challenging prompts. Remarkably, RS-GRPO is simple to implement, requiring only minor code modifications. On six mathematical reasoning benchmarks and with five different LLMs, RS-GRPO consistently improves pass@k performance while maintaining or enhancing pass@1 accuracy.

cs.AI

SN 2023gpw: exploring the diversity and power sources of hydrogen-rich superluminous supernovae

We present our observations and analysis of SN 2023gpw, a hydrogen-rich superluminous supernova (SLSN II) with broad emission lines in its post-peak spectra. Unlike previously observed SLSNe II, its light curve suggests an abrupt drop during a solar conjunction between ~80 and ~180 d after the light-curve peak, possibly analogous to a normal hydrogen-rich supernova (SN). Spectra taken at and before the peak show hydrogen and helium `flash' emission lines attributed to early interaction with a dense confined circumstellar medium (CSM). A well-observed ultraviolet excess appears as these lines disappear, also as a result of CSM interaction. The blackbody photosphere expands roughly at the same velocity throughout the observations, indicating little or no bulk deceleration. This velocity is much higher than what is seen in spectral lines, suggesting asymmetry in the ejecta. The high total radiated energy ($\gtrsim9\times10^{50}$ erg) and aforementioned lack of bulk deceleration in SN 2023gpw are difficult to reconcile with a neutrino-driven SN simply combined with efficient conversion from kinetic energy to emission through interaction. This suggests an additional energy source such as a central engine. While magnetar-powered models qualitatively similar to SN 2023gpw exist, more modeling work is required to determine if they can reproduce the observed properties in combination with early interaction. The required energy might alternatively be provided by accretion onto a black hole created in the collapse of a massive progenitor star.

astro-ph.HE

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

The development of autonomous agents for graphical user interfaces (GUIs) presents major challenges in artificial intelligence. While recent advances in native agent models have shown promise by unifying perception, reasoning, action, and memory through end-to-end learning, open problems remain in data scalability, multi-turn reinforcement learning (RL), the limitations of GUI-only operation, and environment stability. In this technical report, we present UI-TARS-2, a native GUI-centered agent model that addresses these challenges through a systematic training methodology: a data flywheel for scalable data generation, a stabilized multi-turn RL framework, a hybrid GUI environment that integrates file systems and terminals, and a unified sandbox platform for large-scale rollouts. Empirical evaluation demonstrates that UI-TARS-2 achieves significant improvements over its predecessor UI-TARS-1.5. On GUI benchmarks, it reaches 88.2 on Online-Mind2Web, 47.5 on OSWorld, 50.6 on WindowsAgentArena, and 73.3 on AndroidWorld, outperforming strong baselines such as Claude and OpenAI agents. In game environments, it attains a mean normalized score of 59.8 across a 15-game suite-roughly 60% of human-level performance-and remains competitive with frontier proprietary models (e.g., OpenAI o3) on LMGame-Bench. Additionally, the model can generalize to long-horizon information-seeking tasks and software engineering benchmarks, highlighting its robustness across diverse agent tasks. Detailed analyses of training dynamics further provide insights into achieving stability and efficiency in large-scale agent RL. These results underscore UI-TARS-2's potential to advance the state of GUI agents and exhibit strong generalization to real-world interactive scenarios.

cs.AI

Flexible and Probabilistic Topology Tracking with Partial Optimal Transport

In this paper, we present a flexible and probabilistic framework for tracking topological features in time-varying scalar fields using merge trees and partial optimal transport. Merge trees are topological descriptors that record the evolution of connected components in the sublevel sets of scalar fields. We present a new technique for modeling and comparing merge trees using tools from partial optimal transport. In particular, we model a merge tree as a measure network, that is, a network equipped with a probability distribution, and define a notion of distance on the space of merge trees inspired by partial optimal transport. Such a distance offers a new and flexible perspective for encoding intrinsic and extrinsic information in the comparative measures of merge trees. More importantly, it gives rise to a partial matching between topological features in time-varying data, thus enabling flexible topology tracking for scientific simulations. Furthermore, such partial matching may be interpreted as probabilistic coupling between features at adjacent time steps, which gives rise to probabilistic tracking graphs. We derive a stability result for our distance and provide numerous experiments indicating the efficacy of our framework in extracting meaningful feature tracks.

cs.CG

SN 2021aaev: a Hydrogen-Rich Superluminous Supernova with Early Flash and Long-Lived Circumstellar Interaction in an Unusual Host Environment

We present photometric and spectroscopic observations of SN\,2021aaev, a hydrogen-rich, superluminous supernova with persistent (at least $\sim100$ days) narrow Balmer lines (SLSN-IIn) at redshift $z=0.1557$. We observed SN\,2021aaev to rise in $32.5 \pm 1.0$ days since first light and reach a peak absolute magnitude of $-21.46 \pm 0.01$ in the ATLAS $o$ band. The pre-peak spectra resemble those of typical SNe IIn with flash-ionization features arising from the interaction with a dense, confined circumstellar medium (CSM), albeit the flash timescale is longer than usual ($>20$ days). Post peak, the narrow emission lines evolve slowly, and the absence of ejecta features indicates strong deceleration by the CSM. The total radiated energy (about $1.41\times10^{51}$~ergs) is possible with a low-mass (1--$2\,M_{\odot}$) ejecta ploughing into a massive (9--$19\,M_{\odot}$), extended (outer radius $>1\times10^{16}$~cm) H-rich CSM, or alternatively, with magnetar-powered models. Interestingly, the host environment consists of a spiral galaxy with a red substructure in the south-eastern part, and the SN's exact location coincided with the quiescent red substructure (star-formation rate$=0.02^{+0.13}_{-0.02}\,M_{\odot}$~yr$^{-1}$). Given the atypical environment and the obscuring effect of the massive CSM, a thermonuclear (Type Ia-CSM) origin cannot be ruled out. Altogether, SN\,2021aaev is a compelling case to study the diversity of SLSN-IIn features and their host environment.

astro-ph.HE

SN 2023uqf: An Interacting Supernova Coincident with a High-Energy Neutrino

Astrophysical high-energy (TeV-PeV) neutrinos were first discovered in 2013, but their origin remains largely unknown. Here we present SN 2023uqf, a supernova found in coincidence with high-energy neutrino IC231004A, as part of a systematic optical follow-up program with the Zwicky Transient Facility. SN 2023uqf had a luminous and rapidly-evolving lightcurve, and spectroscopic observations indicated that the source was a Type Ibn supernova. Spectroscopic signatures confirm ongoing interaction between the supernova ejecta and a dense circumstellar medium, as expected for high-energy neutrino production in a core-collapse supernova. Given the rare nature of Type Ibn supernovae, SN 2023uqf is unlikely to have been discovered by chance over the course of our program (p=0.3%). Our discovery of SN 2023uqf provides the first observational evidence to support long-held theories that interacting supernovae can serve as cosmic hadron accelerators.

astro-ph.HE

Seed Diffusion: A Large-Scale Diffusion Language Model with High-Speed Inference

We present Seed Diffusion Preview, a large-scale language model based on discrete-state diffusion, offering remarkably fast inference speed. Thanks to non-sequential, parallel generation, discrete diffusion models provide a notable speedup to mitigate the inherent latency of token-by-token decoding, as demonstrated recently (e.g., Mercury Coder, Gemini Diffusion). Seed Diffusion Preview achieves an inference speed of 2,146 token/s over H20 GPUs while maintaining competitive performance across a sweep of standard code evaluation benchmarks, significantly faster than contemporary Mercury and Gemini Diffusion, establishing new state of the art on the speed-quality Pareto frontier for code models.

cs.CL

Twin peaks: SN 2021uvy and SN 2022hgk in the landscape of double-peaked stripped envelope supernovae

In recent years, a class of stripped-envelope supernovae (SESNe) showing two distinct light-curve peaks has emerged, where the first peak cannot be attributed to shock cooling emission. Such peculiar SNe are often studied individually, explained by a combination of powering mechanisms, but are rarely discussed broadly as a group. In this paper, we attempt to form a picture of the landscape of double-peaked SESNe and their powering mechanisms by adding two more objects -- SN 2021uvy and SN 2022hgk. SN 2021uvy is a broad, luminous SN Ib with an unusually long first peak rise and constant color evolution with rising photospheric temperature during the second peak. Though its first peak resembles SN 2019stc, their second peaks differ, making SN 2021uvy unique. SN 2022hgk shows photometric similarity to SN 2019cad and spectroscopic similarity to SN 2005bf, both proposed to be powered by a double-nickel distribution in their ejecta. We analyze their light curves and colors, compare them with a sample of double-peaked SESNe from the ZTF archive, and analyze the light curve parameters of the sample. We observe a correlation (p-value~0.025) between the peak absolute magnitudes of the first and second peaks. No single definitive powering mechanism applies to the whole sample, as it shows variety in the photometric and spectroscopic properties. However, sub-groups of similarity exist that can be explained by mechanisms like the double-nickel distribution, magnetar central engine, interaction, and fallback accretion. We also map out the duration between the peaks ($Δt^{21}$) vs the difference between peak absolute magnitudes ($ΔM^{21}$) as a phase-space that could potentially delineate the most promising powering mechanisms for the double-peaked SESNe.

astro-ph.HE

Truncated Proximal Policy Optimization

Recently, test-time scaling Large Language Models (LLMs) have demonstrated exceptional reasoning capabilities across scientific and professional tasks by generating long chains-of-thought (CoT). As a crucial component for developing these reasoning models, reinforcement learning (RL), exemplified by Proximal Policy Optimization (PPO) and its variants, allows models to learn through trial and error. However, PPO can be time-consuming due to its inherent on-policy nature, which is further exacerbated by increasing response lengths. In this work, we propose Truncated Proximal Policy Optimization (T-PPO), a novel extension to PPO that improves training efficiency by streamlining policy update and length-restricted response generation. T-PPO mitigates the issue of low hardware utilization, an inherent drawback of fully synchronized long-generation procedures, where resources often sit idle during the waiting periods for complete rollouts. Our contributions are two-folds. First, we propose Extended Generalized Advantage Estimation (EGAE) for advantage estimation derived from incomplete responses while maintaining the integrity of policy learning. Second, we devise a computationally optimized mechanism that allows for the independent optimization of the policy and value models. By selectively filtering prompt and truncated tokens, this mechanism reduces redundant computations and accelerates the training process without sacrificing convergence performance. We demonstrate the effectiveness and efficacy of T-PPO on AIME 2024 with a 32B base model. The experimental results show that T-PPO improves the training efficiency of reasoning LLMs by up to 2.5x and outperforms its existing competitors.

cs.AI

PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier

Large Language Models (LLMs) have demonstrated impressive capabilities in complex reasoning tasks, yet they still struggle to reliably verify the correctness of their own outputs. Existing solutions to this verification challenge often depend on separate verifier models or require multi-stage self-correction training pipelines, which limit scalability. In this paper, we propose Policy as Generative Verifier (PAG), a simple and effective framework that empowers LLMs to self-correct by alternating between policy and verifier roles within a unified multi-turn reinforcement learning (RL) paradigm. Distinct from prior approaches that always generate a second attempt regardless of model confidence, PAG introduces a selective revision mechanism: the model revises its answer only when its own generative verification step detects an error. This verify-then-revise workflow not only alleviates model collapse but also jointly enhances both reasoning and verification abilities. Extensive experiments across diverse reasoning benchmarks highlight PAG's dual advancements: as a policy, it enhances direct generation and self-correction accuracy; as a verifier, its self-verification outperforms self-consistency.

cs.CL

DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Inference scaling empowers LLMs with unprecedented reasoning ability, with reinforcement learning as the core technique to elicit complex reasoning. However, key technical details of state-of-the-art reasoning LLMs are concealed (such as in OpenAI o1 blog and DeepSeek R1 technical report), thus the community still struggles to reproduce their RL training results. We propose the $\textbf{D}$ecoupled Clip and $\textbf{D}$ynamic s$\textbf{A}$mpling $\textbf{P}$olicy $\textbf{O}$ptimization ($\textbf{DAPO}$) algorithm, and fully open-source a state-of-the-art large-scale RL system that achieves 50 points on AIME 2024 using Qwen2.5-32B base model. Unlike previous works that withhold training details, we introduce four key techniques of our algorithm that make large-scale LLM RL a success. In addition, we open-source our training code, which is built on the verl framework, along with a carefully curated and processed dataset. These components of our open-source system enhance reproducibility and support future research in large-scale LLM RL.

cs.LG

Seed1.5-Thinking: Advancing Superb Reasoning Models with Reinforcement Learning

We introduce Seed1.5-Thinking, capable of reasoning through thinking before responding, resulting in improved performance on a wide range of benchmarks. Seed1.5-Thinking achieves 86.7 on AIME 2024, 55.0 on Codeforces and 77.3 on GPQA, demonstrating excellent reasoning abilities in STEM and coding. Beyond reasoning tasks, the method demonstrates notable generalization across diverse domains. For instance, it surpasses DeepSeek R1 by 8% in win rate on non-reasoning tasks, indicating its broader applicability. Compared to other state-of-the-art reasoning models, Seed1.5-Thinking is a Mixture-of-Experts (MoE) model with a relatively small size, featuring 20B activated and 200B total parameters. As part of our effort to assess generalized reasoning, we develop two internal benchmarks, BeyondAIME and Codeforces, both of which will be publicly released to support future research. Model trial link: https://www.volcengine.com/experience/ark.

cs.CL

VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

We present VAPO, Value-based Augmented Proximal Policy Optimization framework for reasoning models., a novel framework tailored for reasoning models within the value-based paradigm. Benchmarked the AIME 2024 dataset, VAPO, built on the Qwen 32B pre-trained model, attains a state-of-the-art score of $\mathbf{60.4}$. In direct comparison under identical experimental settings, VAPO outperforms the previously reported results of DeepSeek-R1-Zero-Qwen-32B and DAPO by more than 10 points. The training process of VAPO stands out for its stability and efficiency. It reaches state-of-the-art performance within a mere 5,000 steps. Moreover, across multiple independent runs, no training crashes occur, underscoring its reliability. This research delves into long chain-of-thought (long-CoT) reasoning using a value-based reinforcement learning framework. We pinpoint three key challenges that plague value-based methods: value model bias, the presence of heterogeneous sequence lengths, and the sparsity of reward signals. Through systematic design, VAPO offers an integrated solution that effectively alleviates these challenges, enabling enhanced performance in long-CoT reasoning tasks.

cs.AI

A Unified Pairwise Framework for RLHF: Bridging Generative Reward Modeling and Policy Optimization

Reinforcement Learning from Human Feedback (RLHF) has emerged as a important paradigm for aligning large language models (LLMs) with human preferences during post-training. This framework typically involves two stages: first, training a reward model on human preference data, followed by optimizing the language model using reinforcement learning algorithms. However, current RLHF approaches may constrained by two limitations. First, existing RLHF frameworks often rely on Bradley-Terry models to assign scalar rewards based on pairwise comparisons of individual responses. However, this approach imposes significant challenges on reward model (RM), as the inherent variability in prompt-response pairs across different contexts demands robust calibration capabilities from the RM. Second, reward models are typically initialized from generative foundation models, such as pre-trained or supervised fine-tuned models, despite the fact that reward models perform discriminative tasks, creating a mismatch. This paper introduces Pairwise-RL, a RLHF framework that addresses these challenges through a combination of generative reward modeling and a pairwise proximal policy optimization (PPO) algorithm. Pairwise-RL unifies reward model training and its application during reinforcement learning within a consistent pairwise paradigm, leveraging generative modeling techniques to enhance reward model performance and score calibration. Experimental evaluations demonstrate that Pairwise-RL outperforms traditional RLHF frameworks across both internal evaluation datasets and standard public benchmarks, underscoring its effectiveness in improving alignment and model behavior.

cs.LG

Exploring Data Scaling Trends and Effects in Reinforcement Learning from Human Feedback

Reinforcement Learning from Human Feedback (RLHF) is crucial for aligning large language models with human preferences. While recent research has focused on algorithmic improvements, the importance of prompt-data construction has been overlooked. This paper addresses this gap by exploring data-driven bottlenecks in RLHF performance scaling, particularly reward hacking and decreasing response diversity. We introduce a hybrid reward system combining reasoning task verifiers (RTV) and a generative reward model (GenRM) to mitigate reward hacking. We also propose a novel prompt-selection method, Pre-PPO, to maintain response diversity and enhance learning effectiveness. Additionally, we find that prioritizing mathematical and coding tasks early in RLHF training significantly improves performance. Experiments across two model sizes validate our methods' effectiveness and scalability. Results show that RTV is most resistant to reward hacking, followed by GenRM with ground truth, and then GenRM with SFT Best-of-N responses. Our strategies enable rapid capture of subtle task-specific distinctions, leading to substantial improvements in overall RLHF performance. This work highlights the importance of careful data construction and provides practical methods to overcome performance barriers in RLHF.

cs.LG