SearcharxivSearch

arXiv subjects

Xing Jin

Publications and source records attributed to Xing Jin.

At least 19 recordsLinked to original sources

BAGEN: Are LLM Agents Budget-Aware?

While agents are increasingly spending more resources, today agent cost is mostly measured only after execution. A Budget-Aware Agent (BAGEN) should treat budget as an active control signal, rather than a passive cost metric. We first systematically define budget estimation as internal budgets (from agent computation) and external budgets (from agent actions). We then formalize budget-awareness as progressive interval estimation: at each step of a plan, an agent should predict an upper and lower bound on remaining budget, and alert when completion is unlikely. Scoring with a rollout-replay protocol, we find consistent failure patterns on four environments and five frontier agents: (1) strong agents do not necessarily have strong budget-awareness, with correlation r=0.35. (2) frontier models are consistently over-optimistic, continue spending on tasks that are unlikely to succeed, instead of alerting the user early. (3) budget-aware signal is actionable and trainable. Early stop saves 28-64% tokens on failed trajectories, and SFT+RL strengthens early stop and alert behavior. (4) precise interval calibration remains challenging, with interval coverage capping at 47% after SFT+RL. Project page: https://ragen-ai.github.io/bagen/

cs.LG

RAGEN-2: Reasoning Collapse in Agentic RL

RL training of multi-turn LLM agents is inherently unstable, and reasoning quality directly determines task performance. Entropy is widely used to track reasoning stability. However, entropy only measures diversity within the same input, and cannot tell whether reasoning actually responds to different inputs. In RAGEN-2, we find that even with stable entropy, models can rely on fixed templates that look diverse but are input-agnostic. We call this template collapse, a failure mode invisible to entropy and all existing metrics. To diagnose this failure, we decompose reasoning quality into within-input diversity (Entropy) and cross-input distinguishability (Mutual Information, MI), and introduce a family of mutual information proxies for online diagnosis. Across diverse tasks, mutual information correlates with final performance much more strongly than entropy, making it a more reliable proxy for reasoning quality. We further explain template collapse with a signal-to-noise ratio (SNR) mechanism. Low reward variance weakens task gradients, letting regularization terms dominate and erase cross-input reasoning differences. To address this, we propose SNR-Aware Filtering to select high-signal prompts per iteration using reward variance as a lightweight proxy. Across planning, math reasoning, web navigation, and code execution, the method consistently improves both input dependence and task performance.

cs.LG

Seed-Prover 1.5: Mastering Undergraduate-Level Theorem Proving via Learning from Experience

Large language models have recently made significant progress to generate rigorous mathematical proofs. In contrast, utilizing LLMs for theorem proving in formal languages (such as Lean) remains challenging and computationally expensive, particularly when addressing problems at the undergraduate level and beyond. In this work, we present \textbf{Seed-Prover 1.5}, a formal theorem-proving model trained via large-scale agentic reinforcement learning, alongside an efficient test-time scaling (TTS) workflow. Through extensive interactions with Lean and other tools, the model continuously accumulates experience during the RL process, substantially enhancing the capability and efficiency of formal theorem proving. Furthermore, leveraging recent advancements in natural language proving, our TTS workflow efficiently bridges the gap between natural and formal languages. Compared to state-of-the-art methods, Seed-Prover 1.5 achieves superior performance with a smaller compute budget. It solves \textbf{88\% of PutnamBench} (undergraduate-level), \textbf{80\% of Fate-H} (graduate-level), and \textbf{33\% of Fate-X} (PhD-level) problems. Notably, using our system, we solved \textbf{11 out of 12 problems} from Putnam 2025 within 9 hours. Our findings suggest that scaling learning from experience, driven by high-quality formal feedback, holds immense potential for the future of formal mathematical reasoning.

cs.CL

Hertz-Integral-Linewidth Lasers based on Portable Solid-state Microresonators

Optical reference resonators serve as a cornerstone in various scientific fields. In recent years, there has been an increasing demand for compact ultrastable reference resonators capable of operating in ambient environments, enabling applications beyond the laboratory, such as navigation, portable optical clocks, and remote sensing. Here, we present a compact ultrastable whispering-gallery-mode \ce{MgF2} reference resonator with a high loaded quality factor of $2.24\times 10^9$. The device is packaged in a compact form of 50$\times$77$\times$90 mm and supports stable optical coupling with polarization-maintaining fiber, which enables robust operation under ambient conditions. Laser stabilization using this resonator yields a phase noise of -105 dBc/Hz at a 10 kHz offset frequency, an integral linewidth of 4 Hz, and a fractional frequency stability of $2.5\times 10^{-14}$ at a 10 ms averaging time. With the high performance and rapid manufacturability, our work offers a promising solution for ultrastable optical frequency references beyond laboratory settings.

physics.optics

Communication-ready high-power soliton microcombs in highly-dispersive Fabry-Perot-microresonators

Microcombs generated in optical microresonators are widely regarded as promising light sources for next-generation communication systems, but the optical power available per comb line has so far fallen short of practical requirements. Here we introduce an integrated Fabry-P\'erot microresonator platform that overcomes fundamental dispersion-engineering constraints and enables bright soliton microcombs with unprecedented power per line. The resonator is defined by chirped Bragg gratings that provide exceptionally large anomalous group-velocity dispersion, allowing more than ten comb lines to reach the milliwatt level. These combs can be used directly in coherent communication systems without additional amplification, achieving an aggregate data rate of 2 Tb/s. Once integrated, our high-power soliton microcombs could be instantly ready for communications as well as a broad range of practical comb-based applications.

physics.optics

Synergetic Enhancement on Bulk and Grain Boundary Ionic Conduction of Mg Doped High-Entropy NASICON-Type Solid Electrolyte for Solid-State Na+ Batteries by Spray Flame Synthesis

All-solid-state sodium batteries represent a promising next-generation energy storage technology, owing to cost-effectiveness and enhanced safety. Among solid electrolytes for solid-state sodium batteries, NASICON-structured Na3Zr2Si2PO12 has emerged as a predominant candidate. However, its widespread implementation remains limited by suboptimal ionic conductivity in both bulk and grain boundary regions. In this study, we demonstrate a novel approach utilizing swirling spray flame synthesis to produce Mg-doped NASICON solid electrolyte nanoparticles. This method facilitates efficient doping and homogeneous mixing for scalable production, resulting in core-shell non-NASICON structures with nano-scale high-entropy mixing. Notably, the atomic migration distances achieved by flame synthesis are significantly reduced compared to conventional solid-state reactions, thereby enabling reactive sintering to preserve high sinterability of nanoparticles during post-treatment processes. High-temperature sintering yields dense NASICON-structured solid electrolytes. Among those, Mg0.25NZSP exhibits an optimal ionic conductivity of 1.91 mS/cm at room temperature and an activation energy of 0.200 eV. The enhancement mechanism can be attributed to incorporation into the NASICON phase and formation of a secondary phase. The low-melting-point secondary phase significantly improves grain boundary contact to enhance grain boundary conductivity. The process achieves simultaneous enhancement of both bulk and grain boundary conduction through a single-step procedure. Comparative analysis of sintering temperatures and ionic conductivities among NASICON solid electrolytes synthesized via different methods demonstrates flame-synthesized nanoparticles offer superior performance and reduced post-treatment costs, owing to their exceptional nano-scale sinterability and uniform elemental distribution.

cond-mat.mtrl-sci

Electrically-pumped soliton microcombs on thin-film lithium niobate

Thin-film lithium niobate (TFLN) has enabled efficient on-chip electro-optic modulation and frequency conversion for information processing and precision measurement. Extending these capabilities with optical frequency combs unlocks massively parallel operations and coherent optical-to-microwave transduction, which are achievable in TFLN microresonators via Kerr microcombs. However, fully integrated Kerr microcombs directly driven by semiconductor lasers remain elusive, which has delayed integration of these technologies. Here we demonstrate electrically pumped TFLN Kerr microcombs without optical amplification. With optimized laser-to-chip coupling and optical quality factors, we generate soliton microcombs at a 200 GHz repetition frequency with an optical span of 180 nm using only 25 mW of pump power. Moreover, self-injection locking enables turnkey initiation and substantially narrows the laser linewidth. Our work provides integrated comb sources for TFLN-based communicational, computational, and metrological applications.

physics.optics

Seed-Prover: Deep and Broad Reasoning for Automated Theorem Proving

LLMs have demonstrated strong mathematical reasoning abilities by leveraging reinforcement learning with long chain-of-thought, yet they continue to struggle with theorem proving due to the lack of clear supervision signals when solely using natural language. Dedicated domain-specific languages like Lean provide clear supervision via formal verification of proofs, enabling effective training through reinforcement learning. In this work, we propose \textbf{Seed-Prover}, a lemma-style whole-proof reasoning model. Seed-Prover can iteratively refine its proof based on Lean feedback, proved lemmas, and self-summarization. To solve IMO-level contest problems, we design three test-time inference strategies that enable both deep and broad reasoning. Seed-Prover proves $78.1\%$ of formalized past IMO problems, saturates MiniF2F, and achieves over 50\% on PutnamBench, outperforming the previous state-of-the-art by a large margin. To address the lack of geometry support in Lean, we introduce a geometry reasoning engine \textbf{Seed-Geometry}, which outperforms previous formal geometry engines. We use these two systems to participate in IMO 2025 and fully prove 5 out of 6 problems. This work represents a significant advancement in automated mathematical reasoning, demonstrating the effectiveness of formal verification with long chain-of-thought reasoning.

cs.AI

Spray flame synthesis of Y2O3-MgO nanoparticles for mid-infrared transparent nanocomposite ceramics

Spray flame synthesis offers a promising method for scalable production of homogeneously mixed Y2O3-MgO nanopowders as next-generation infrared-transparent window material, which has attracted significant attention owing to its excellent optical properties at high temperatures. However, systematic understanding of how flame synthesis parameters influence particle morphology, crystal phase, solid solubility, and subsequent ceramic performance remains insufficiently understood. In this study, we investigated the influence of precursor chemistry on particle crystal phase and examined the solid solubility of MgO in Y2O3 under different flame temperatures, demonstrating that the high-temperature conditions with O2 as dispersion gas allow up to 50 mol% MgO to fully dissolve into Y2O3, far exceeding the equilibrium solubility limit of 7 mol% at the eutectic temperature (2100{\deg}C) and near-zero at room temperature. Furthermore, we systematically evaluated how powder characteristics and sintering parameters-including powder deagglomeration methods, vacuum sintering temperature, hot isostatic pressing (HIP) temperature, and initial powder characteristics-affect ceramic microstructures and infrared transmittance. Despite cracking induced by phase transformation and finer particle sizes, ceramics fabricated from oxygen-synthesized monoclinic-dominated powders exhibited superior near-infrared transmittance (56.2% at 1550 nm), attributed to enhanced atomic mixing and effective grain boundary pinning. After optimization, pure cubic phase powders produced intact and crack-free ceramics with outstanding mid-infrared transparency, achieving a maximum transmittance of 84.6% and an average transmittance of 82.3% in 3-5 um range.

cond-mat.mtrl-sci

Isogeometric contact analysis in subsea umbilical and power cables

Subsea umbilical and power cables contain a large number of contact interfaces between different geometries and materials. These complex interactions rise significant challenges for accurately considering contact surface properties by using traditional analytical solutions or finite element methods. These properties have been identified as the most sensitive parameters when performing the numerical simulation for stress analysis. Therefore, it is essential to apply a novel approach for contact analysis which improves the accuracy and efficiency for predicting contact properties. This paper presents an isogeometric analysis (IGA) approach addressing contact problems in dynamic umbilicals and power cables. Firstly, this isogeometric contact algorithm is formulated in MATLAB as a tool including the geometry description, contact detection and penalty function. Secondly, the contact interface between a steel tube and an outer sheath in an dynamic umbilical is established by this IGA contact algorithm and validated against that in ABAQUS for proving the accuracy and efficiency of IGA. Finally, the effects of element refinement, geometrical description, penalty factor on the accuracy, efficiency and stability of IGA are discussed.

math.NA

RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Training large language models (LLMs) as interactive agents presents unique challenges including long-horizon decision making and interacting with stochastic environment feedback. While reinforcement learning (RL) has enabled progress in static tasks, multi-turn agent RL training remains underexplored. We propose StarPO (State-Thinking-Actions-Reward Policy Optimization), a general framework for trajectory-level agent RL, and introduce RAGEN, a modular system for training and evaluating LLM agents. Our study on four stylized environments reveals three core findings. First, our agent RL training shows a recurring mode of Echo Trap where reward variance cliffs and gradient spikes; we address this with StarPO-S, a stabilized variant with trajectory filtering, critic incorporation, and gradient stabilization. Second, we find the shaping of RL rollouts would benefit from diverse initial states, medium interaction granularity and more frequent sampling. Third, we show that without fine-grained, reasoning-aware reward signals, agent reasoning hardly emerge through multi-turn RL and they may show shallow strategies or hallucinated thoughts. Code and environments are available at https://github.com/RAGEN-AI/RAGEN.

cs.LG

Heimdall: test-time scaling on the generative verification

An AI system can create and maintain knowledge only to the extent that it can verify that knowledge itself. Recent work on long Chain-of-Thought reasoning has demonstrated great potential of LLMs on solving competitive problems, but their verification ability remains to be weak and not sufficiently investigated. In this paper, we propose Heimdall, the long CoT verification LLM that can accurately judge the correctness of solutions. With pure reinforcement learning, we boost the verification accuracy from 62.5% to 94.5% on competitive math problems. By scaling with repeated sampling, the accuracy further increases to 97.5%. Through human evaluation, Heimdall demonstrates impressive generalization capabilities, successfully detecting most issues in challenging math proofs, the type of which is not included during training. Furthermore, we propose Pessimistic Verification to extend the functionality of Heimdall to scaling up the problem solving. It calls Heimdall to judge the solutions from a solver model and based on the pessimistic principle, selects the most likely correct solution with the least uncertainty. Taking DeepSeek-R1-Distill-Qwen-32B as the solver model, Pessimistic Verification improves the solution accuracy on AIME2025 from 54.2% to 70.0% with 16x compute budget and to 83.3% with more compute budget. With the stronger solver Gemini 2.5 Pro, the score reaches 93.0%. Finally, we prototype an automatic knowledge discovery system, a ternary system where one poses questions, another provides solutions, and the third verifies the solutions. Using the data synthesis work NuminaMath for the first two components, Heimdall effectively identifies problematic records within the dataset and reveals that nearly half of the data is flawed, which interestingly aligns with the recent ablation studies from NuminaMath.

cs.AI

Seed1.5-Thinking: Advancing Superb Reasoning Models with Reinforcement Learning

We introduce Seed1.5-Thinking, capable of reasoning through thinking before responding, resulting in improved performance on a wide range of benchmarks. Seed1.5-Thinking achieves 86.7 on AIME 2024, 55.0 on Codeforces and 77.3 on GPQA, demonstrating excellent reasoning abilities in STEM and coding. Beyond reasoning tasks, the method demonstrates notable generalization across diverse domains. For instance, it surpasses DeepSeek R1 by 8% in win rate on non-reasoning tasks, indicating its broader applicability. Compared to other state-of-the-art reasoning models, Seed1.5-Thinking is a Mixture-of-Experts (MoE) model with a relatively small size, featuring 20B activated and 200B total parameters. As part of our effort to assess generalized reasoning, we develop two internal benchmarks, BeyondAIME and Codeforces, both of which will be publicly released to support future research. Model trial link: https://www.volcengine.com/experience/ark.

cs.CL

Power-efficient ultra-broadband soliton microcombs in resonantly-coupled microresonators

The drive to miniaturize optical frequency combs for practical deployment has spotlighted microresonator solitons as a promising chip-scale candidate. However, these soliton microcombs could be very power-hungry when their span increases, especially with fine comb spacings. As a result, realizing an octave-spanning comb at microwave repetition rates for direct optical-microwave linkage is considered not possible for photonic integration due to the high power requirements. Here, we introduce the concept of resonant-coupling to soliton microcombs to reduce pump consumption significantly. Compared to conventional waveguide-coupled designs, we demonstrate (i) a threefold increase in spectral span for high-power combs and (ii) up to a tenfold reduction in repetition frequency for octave-spanning operation. This configuration is compatible with laser integration and yields reliable, turnkey soliton generation. By eliminating the long-standing pump-power bottleneck, microcombs will soon become readily available for portable optical clocks, massively parallel data links, and field-deployable spectrometers.

physics.optics

Compact Turnkey Soliton Microcombs at Microwave Rates via Wafer-Scale Fabrication

Soliton microcombs generated in nonlinear microresonators facilitate the photonic integration of timing, frequency synthesis, and astronomical calibration functionalities. For these applications, low-repetition-rate soliton microcombs are essential as they establish a coherent link between optical and microwave signals. However, the required pump power typically scales with the inverse of the repetition rate, and the device footprint scales with the inverse of square of the repetition rate, rendering low-repetition-rate soliton microcombs challenging to integrate within photonic circuits. This study designs and fabricates silicon nitride microresonators on 4-inch wafers with highly compact form factors. The resonator geometries are engineered from ring to finger and spiral shapes to enhance integration density while attaining quality factors over 10^7. Driven directly by an integrated laser, soliton microcombs with repetition rates below 10 GHz are demonstrated via turnkey initiation. The phase noise performance of the synthesized microwave signals reaches -130 dBc/Hz at 100 kHz offset frequency for 10 GHz carrier frequencies. This work enables the high-density integration of soliton microcombs for chip-based microwave photonics and spectroscopy applications.

physics.optics

Soliton microcombs in X-cut LiNbO3 microresonators

Chip-scale integration of optical frequency combs, particularly soliton microcombs, enables miniaturized instrumentation for timekeeping, ranging, and spectroscopy. Although soliton microcombs have been demonstrated on various material platforms, realizing complete comb functionality on photonic chips requires the co-integration of high-speed modulators and efficient frequency doublers, features that are available in a monolithic form on X-cut thin-film lithium niobate (TFLN). However, the pronounced Raman nonlinearity associated with extraordinary light in this platform has so far precluded soliton microcomb generation. Here, we report the generation of transverse-electric-polarized soliton microcombs with a 25 GHz repetition rate in high-Q microresonators on X-cut TFLN chips. By precisely orienting the racetrack microresonator relative to the optical axis, we mitigate Raman nonlinearity and enable soliton formation under continuous-wave laser pumping. Moreover, the soliton microcomb spectra are extended to 350 nm with pulsed laser pumping. This work expands the capabilities of TFLN photonics and paves the way for the monolithic integration of fast-tunable, self-referenced microcombs.

physics.optics

Flaming-hot Initiation with Regular Execution Sampling for Large Language Models

Since the release of ChatGPT, large language models (LLMs) have demonstrated remarkable capabilities across various domains. A key challenge in developing these general capabilities is efficiently sourcing diverse, high-quality data. This becomes especially critical in reasoning-related tasks with sandbox checkers, such as math or code, where the goal is to generate correct solutions to specific problems with higher probability. In this work, we introduce Flaming-hot Initiation with Regular Execution (FIRE) sampling, a simple yet highly effective method to efficiently find good responses. Our empirical findings show that FIRE sampling enhances inference-time generation quality and also benefits training in the alignment stage. Furthermore, we explore how FIRE sampling improves performance by promoting diversity and analyze the impact of employing FIRE at different positions within a response.

cs.LG

Process Supervision-Guided Policy Optimization for Code Generation

Reinforcement learning (RL) with unit test feedback has enhanced large language models' (LLMs) code generation, but relies on sparse rewards provided only after complete code evaluation, limiting learning efficiency and incremental improvements. When generated code fails all unit tests, no learning signal is received, hindering progress on complex tasks. To address this, we propose a Process Reward Model (PRM) that delivers dense, line-level feedback on code correctness during generation, mimicking human code refinement and providing immediate guidance. We explore various strategies for training PRMs and integrating them into the RL framework, finding that using PRMs both as dense rewards and for value function initialization significantly boosts performance. Our experimental results also highlight the effectiveness of PRMs in enhancing RL-driven code generation, especially for long-horizon scenarios.

cs.AI