SearcharxivSearch

arXiv subjects

Wen Sun

Publications and source records attributed to Wen Sun.

At least 19 recordsLinked to original sources

Tilting Billingsley's model toward a giant prime: two phase transitions

Billingsley's theorem states that the normalized logarithms of the prime factors of a uniform random integer in $[1,x]$ converge to the Poisson--Dirichlet law $PD(1)$, while weighting an integer $n$ by the generalized divisor function $d_\theta(n)$ gives $PD(\theta)$. We ask what remains of this picture when the integer is also rewarded according to its largest prime factor $P^+(n)$. For fixed $\theta,\beta>0$ and $\gamma\ge0$, we sample $N_x=n\le x$ with probability proportional to $d_\theta(n)\exp\{\beta H_\gamma(n)\}$, where $H_\gamma(n)=(\log P^+(n))^\gamma$ for $\gamma>0$, while $H_0(1)=0$ and $H_0(n)=1$ for $n\ge2$. Two phase transitions occur. At $\gamma=0$ the $PD(\theta)$ partition survives, whereas every fixed $\gamma>0$ forces one prime to carry asymptotically all logarithmic mass. The second transition, at $\gamma=1$, concerns the cofactor $R_x=N_x/P^+(N_x)$. With $a_x=\beta\gamma(\log x)^{\gamma-1}$, its law is asymptotic in total variation to $Q_{a_x}(m)=d_\theta(m)m^{-1-a_x}/\zeta(1+a_x)^\theta$. Thus, for $0<\gamma<1$, $a_x\log R_x$ has a $Gamma(\theta,1)$ limit, with shape $\theta$ and rate one, and the normalized logarithmic prime factors of $R_x$ have an independent $PD(\theta)$ limit; at $\gamma=1$, $R_x$ has a nondegenerate discrete limit; and for $\gamma>1$, $R_x=1$ with high probability. We also determine the joint limits involving $N_x/x$ and the normalizing constants, which include the classical Alladi--Erd\H{o}s asymptotic as a special case.

math.PR

Existence and Stability of Dancing Equilibria in Asymmetric Kuramoto Networks

We study nonzero-frequency phase-locked motions in asymmetrically coupled Kuramoto networks. Such motions are relative equilibria with fixed phase differences and a nonzero common angular velocity, and we call them dancing equilibria. Their existence requires all coupling sums to have the same nonzero value. We show that neither symmetric coupling nor an acyclic associated digraph can support a dancing equilibrium. We introduce structurally equitable and $q$-twisted state equitable partitions and prove a partition-based criterion for the resulting class-constant profiles, with standard labeled $q$-twisted profiles recovered from singleton partitions. For the forward $m$-neighbor model, we characterize existence by an exact indivisibility criterion. Stability is studied modulo the common phase-shift direction. For general directed networks, strong connectivity and edgewise phase differences in $\left(-\pi/2,\pi/2\right)$ imply local orbital exponential stability and yield an explicit positively invariant set contained in the local basin of attraction. For arbitrary twisted indices, this contraction argument gives a low-winding stability regime with explicit positively invariant neighborhoods. For each existing $q$-twisted branch of the forward model, a discrete Fourier transform criterion yields local orbital exponential stability when all nonzero Fourier-mode factors are positive and nonlinear instability when at least one is negative. In the unstable case, the proof constructs explicit escaping real Fourier perturbations. We further derive additional explicit stability and instability ranges for arbitrary twisted indices in terms of constants $N$, $m$, and $q$. For the first two twisted branches, sharper arguments yield a first-mode transition criterion for $q=1$ and a complete finite-size classification for $q=2$, with the degenerate case in each branch handled separately.

math.DS

KREL: Automatic Medical Coding via Knowledge-Guided Reasoning over Clinical Evidence with LLMs

Automatic Medical Coding (AMC), which assigns standardized International Classification of Diseases (ICD) codes to clinical notes, is essential for medical reimbursement, quality reporting, and clinical research. Existing pre-trained language model (PLM)-based methods typically formulate AMC as an extreme multi-label classification problem over a predefined code set, while recent large language model (LLM)-based approaches instead frame it as generation or multi-step reasoning. However, key challenges remain, including the extreme length of clinical notes that hinders effective interpretation, the vast ICD label space, and complex coding rules that are not explicitly captured by LLMs. In this work, we propose Knowledge-Guided Reasoning over Clinical Evidence with LLMs (KREL), a framework that leverages LLMs for clinical text understanding and reasoning while integrating external ICD coding guidelines as structured knowledge. This design enables tight coupling between domain knowledge and LLM reasoning, reducing hallucinations and improving compliance with coding standards. Experiments on benchmark datasets show that KREL consistently outperforms strong PLM-based and state-of-the-art LLM-based baselines.

cs.CL

Uniform High Order Factorial Moment Bounds for the Critical Erd\H{o}s-R\'enyi Component Process

We give a finite $n$ enumerative derivation of the local surplus marked point process form of Aldous's critical window limit for the Erd\H{o}s-R\'enyi random graph. For $p_n=n^{-1}+\lambda n^{-4/3}$, let $\Xi_n$ place an atom at the rescaled size and surplus of each component. Exact component enumeration yields the limiting factorial correlation densities and the uniform bound \[ \mathbb E[(\Xi_n(K))_q]\le C_K^q e^{-c_Kq^3}, \qquad q\ge1, \] for every compact marked window $K$, uniformly in admissible $n$. The cubic order is optimal for the limiting factorial measures, and the estimate persists after summing over all surpluses on compact size intervals. It yields overcrowding and local exponential moment bounds, quantitative truncation of the finite $n$ Laplace functional expansion, and local point process convergence. Using the Janson-Spencer Palm description and a classical all excess estimate, we also recover ordered $\ell^2$ component size convergence.

math.PR

Dynamical Equations for Poisson Galton--Watson Trees and Component Densities of Sparse Inhomogeneous Random Graphs

We study Poisson Galton--Watson trees on a standard Borel type space when the offspring kernel is multiplied by a scalar parameter. On finite trees, we identify the Radon--Nikodym derivative between two parameter values and show that it remains measurable after projection to the total progeny measure. Under a uniform bound on the offspring intensities, differentiation yields exact differential and integral equations for the projected laws without irreducibility, reversibility, or a positive eigenfunction. With an additional positive eigenfunction bounded above and away from zero, we relate these equations to an infinite spinal tree, uniform pruning, the Doob transform, and the Aldous--Pitman ascension process. For a uniformly bounded offspring kernel, we also prove uniform exponential integrability of the total progeny throughout the spectrally subcritical regime. As an application, under the graphical-kernel assumptions of Bollobas, Janson and Riordan, the number $K_n$ of connected components satisfies $K_n/n \to {\mathbb E}_{\pi}[1/T_u]$ in probability and in $L^1$, where $T_u$ is the total progeny of the associated branching process and $1/\infty=0$. If $q_u(x)$ is its extinction probability from type $x$, re-rooting and extinction duality give the explicit limit $$ \int_S q_u(x)\,\pi(dx) - {u\over 2}\int_{S\times S}\kappa(x,y)q_u(x)q_u(y)\,\pi(dx)\pi(dy). $$ This extends the finite-type and compact-continuous formulas to the full BJR graphical-kernel setting, allowing separable noncompact type spaces and kernels that may be unbounded or reducible.

math.PR

Macroscopic Feynman Cycles and Poisson--Kingman Universality in Bose Condensation

We prove a canonical limit theorem for the macroscopic Feynman cycles of finite-volume ideal Bose gases. Cycles carry marks in a general Polish space $\mathsf{M}$, encoding spatial, geometric, spectral, or internal data. After removing a deterministic background density $\rho_{\mathrm{bg}}$, the marked macroscopic cycle process converges in the canonical ensemble to a marked Poisson--Kingman bridge of total mass $\rho - \rho_{\mathrm{bg}}$. The bridge is constructed from a marked Poisson point process with intensity $x^{-1}\eta_x(dm)\,dx$, conditioned on total mass~$\rho - \rho_{\mathrm{bg}}$, where the kernel $x \mapsto \eta_x$ and its total-mass profile $\phi(x) = \eta_x(\mathsf{M})$ are determined by the low-energy spectral data visible on the scale $j \sim V_L$. When $\phi$ is constant, the bridge reduces to a Gamma bridge and the ranked cycle lengths follow the Poisson--Dirichlet law. We verify this for the ideal Bose gas in dimension $d > 2$ under periodic, Dirichlet, and Neumann boundary conditions: in all three cases $\phi \equiv 1$ and the ranked lengths converge to $\mathrm{PD}(0,1)$, while the mark kernels distinguish the three models through their winding, killed-bridge, and reflected-bridge geometry. When $\phi$ is not constant, the bridge is no longer Gamma and the ranked lengths are not Poisson--Dirichlet. As a concrete example, a critical double-well potential whose tunnelling splitting satisfies $V_L \Delta_L \to \gamma$ gives $\phi_\gamma(x) = 1 + e^{-\beta\gamma x}$; more generally, a finite-type visible spectrum with $Q$ components yields $\phi(x) = \sum_{r=1}^{Q} \theta_r e^{-\beta\lambda_r x}$. These results identify Poisson--Kingman bridges as the canonical universality class for marked macroscopic Bose cycles, with the visible low-energy spectrum selecting the particular bridge.

math.PR

Agentic evolution of physically constrained foundation models

Artificial intelligence increasingly drives automated scientific discovery, yet contemporary generalist agents lack physical grounding, frequently hallucinating hardware-incompatible designs. Here, we present a physically grounded, multi-agent discovery engine that autonomously architects hardware-compliant computing systems. Anchored by an Evolutionary Knowledge Graph structuring past scientific innovations, the framework extracts an "algorithmic Chain-of-Thought" to transform blind stochastic search into directed structural evolution. Applied to the extreme testbed of foundation model deployment, the engine evolved two hardware-aware compression methodologies surpassing human-engineered heuristics: Q-Enhance mitigates long-context accuracy loss in dense models, and MoE-Salient-AQ outperforms state-of-the-art manual sparse Mixture-of-Experts designs by 3.7% at sub-3-bit regimes. Utilizing a bandwidth-efficient Sensitivity Profile, we successfully deployed a massive 235-billion-parameter model onto a constrained dual-A100 server, reducing memory requirements by 75% with a marginal 0.64% accuracy degradation. By transforming unconstrained combinatorial search into knowledge-driven autonomy, this establishes a scalable hardware-software co-design paradigm for machine-driven discovery within strict physical boundaries.

cs.AI

CLIPer: Tailoring Diverse User Preference via Classifier-Guided Inference-Time Personalization

Personalized LLMs can significantly enhance user experiences by tailoring responses to preferences such as helpfulness, conciseness, and humor. However, fine-tuning models to address all possible combinations of user preferences is computationally expensive and impractical. In this paper, we introduce \textbf{CLIPer}(\textbf{Cl}assifier-guided \textbf{I}nference-time \textbf{Per}sonalization), a lightweight personalization approach that leverages a classifier model to steer LLM generation dynamically to different user preferences at inference time. Our method eliminates the need for extensive fine-tuning, inducing negligible additional computational overhead while enabling more controllable and nuanced personalization across single and multi-dimensional preferences. Comprehensive empirical analyses demonstrate the scalability and effectiveness of our approach in delivering personalized language generation.

cs.CL

$p1$: Better Prompt Optimization with Fewer Prompts

Prompt optimization improves language models without updating their weights by searching for a better system prompt, but its effectiveness varies widely across tasks. We study what makes a task amenable to prompt optimization. We show that the reward variance across different system prompts can be decomposed into two components: variance among responses, which captures generation stochasticity, and variance among system prompts, which captures differences in system prompt quality. Prompt optimization succeeds when variance among system prompts is sufficiently large, but fails when variance among responses dominates the variance of the system prompts. Surprisingly, we further show that scaling to more user prompts can hurt optimization by reducing variance among system prompts, especially on heterogeneous datasets where different user prompts favor different system prompts. Motivated by this insight, we propose $p1$, a simple user prompt filtering method that selects a small subset of user prompts with high variance across candidate system prompts. This subset of user prompts allows one to distinguish a good system prompt from a bad one, making system optimization easier. Experiments on reasoning benchmarks show that $p1$ substantially improves prompt optimization over training on the full dataset and outperforms strong baselines such as GEPA. Notably, training on only two prompts from AIME 24 yields a system prompt that generalizes well to other reasoning benchmarks.

cs.LG

KARL: Knowledge Agents via Reinforcement Learning

We present a system for training enterprise search agents via reinforcement learning that achieves state-of-the-art performance across a diverse suite of hard-to-verify agentic search tasks. Our work makes four core contributions. First, we introduce KARLBench, a multi-capability evaluation suite spanning six distinct search regimes, including constraint-driven entity search, cross-document report synthesis, tabular numerical reasoning, exhaustive entity retrieval, procedural reasoning over technical documentation, and fact aggregation over internal enterprise notes. Second, we show that models trained across heterogeneous search behaviors generalize substantially better than those optimized for any single benchmark. Third, we develop an agentic synthesis pipeline that employs long-horizon reasoning and tool use to generate diverse, grounded, and high-quality training data, with iterative bootstrapping from increasingly capable models. Fourth, we propose a new post-training paradigm based on iterative large-batch off-policy RL that is sample efficient, robust to train-inference engine discrepancies, and naturally extends to multi-task training with out-of-distribution generalization. Compared to Claude 4.6 and GPT 5.2, KARL is Pareto-optimal on KARLBench across cost-quality and latency-quality trade-offs, including tasks that were out-of-distribution during training. With sufficient test-time compute, it surpasses the strongest closed models. These results show that tailored synthetic data in combination with multi-task reinforcement learning enables cost-efficient and high-performing knowledge agents for grounded reasoning.

cs.AI

LLMs Can Learn to Reason Via Off-Policy RL

Reinforcement learning (RL) approaches for Large Language Models (LLMs) frequently use on-policy algorithms, such as PPO or GRPO. However, policy lag from distributed training architectures and differences between the training and inference policies break this assumption, making the data off-policy by design. To rectify this, prior work has focused on making this off-policy data appear more on-policy, either via importance sampling (IS), or by more closely aligning the training and inference policies by explicitly modifying the inference engine. In this work, we embrace off-policyness and propose a novel off-policy RL algorithm that does not require these modifications: Optimal Advantage-based Policy Optimization with Lagged Inference policy (OAPL). We show that OAPL outperforms GRPO with importance sampling on competition math benchmarks, and can match the performance of a publicly available coding model, DeepCoder, on LiveCodeBench, while using 3x fewer generations during training. We further empirically demonstrate that models trained via OAPL have improved test time scaling under the Pass@k metric. OAPL allows for efficient, effective post-training even with lags of more than 400 gradient steps between the training and inference policies, 100x more off-policy than prior approaches.

cs.LG

Step-DeepResearch Technical Report

As LLMs shift toward autonomous agents, Deep Research has emerged as a pivotal metric. However, existing academic benchmarks like BrowseComp often fail to meet real-world demands for open-ended research, which requires robust skills in intent recognition, long-horizon decision-making, and cross-source verification. To address this, we introduce Step-DeepResearch, a cost-effective, end-to-end agent. We propose a Data Synthesis Strategy Based on Atomic Capabilities to reinforce planning and report writing, combined with a progressive training path from agentic mid-training to SFT and RL. Enhanced by a Checklist-style Judger, this approach significantly improves robustness. Furthermore, to bridge the evaluation gap in the Chinese domain, we establish ADR-Bench for realistic deep research scenarios. Experimental results show that Step-DeepResearch (32B) scores 61.4% on Scale AI Research Rubrics. On ADR-Bench, it significantly outperforms comparable models and rivals SOTA closed-source models like OpenAI and Gemini DeepResearch. These findings prove that refined training enables medium-sized models to achieve expert-level capabilities at industry-leading cost-efficiency.

cs.CL

Step-GUI Technical Report

Recent advances in multimodal large language models unlock unprecedented opportunities for GUI automation. However, a fundamental challenge remains: how to efficiently acquire high-quality training data while maintaining annotation reliability? We introduce a self-evolving training pipeline powered by the Calibrated Step Reward System, which converts model-generated trajectories into reliable training signals through trajectory-level calibration, achieving >90% annotation accuracy with 10-100x lower cost. Leveraging this pipeline, we introduce Step-GUI, a family of models (4B/8B) that achieves state-of-the-art GUI performance (8B: 80.2% AndroidWorld, 48.5% OSWorld, 62.6% ScreenShot-Pro) while maintaining robust general capabilities. As GUI agent capabilities improve, practical deployment demands standardized interfaces across heterogeneous devices while protecting user privacy. To this end, we propose GUI-MCP, the first Model Context Protocol for GUI automation with hierarchical architecture that combines low-level atomic operations and high-level task delegation to local specialist models, enabling high-privacy execution where sensitive data stays on-device. Finally, to assess whether agents can handle authentic everyday usage, we introduce AndroidDaily, a benchmark grounded in real-world mobile usage patterns with 3146 static actions and 235 end-to-end tasks across high-frequency daily scenarios (8B: static 89.91%, end-to-end 52.50%). Our work advances the development of practical GUI agents and demonstrates strong potential for real-world deployment in everyday digital interactions.

cs.CV

Friend or Foe: How LLMs' Safety Mind Gets Fooled by Intent Shift Attack

Large language models (LLMs) remain vulnerable to jailbreaking attacks despite their impressive capabilities. Investigating these weaknesses is crucial for robust safety mechanisms. Existing attacks primarily distract LLMs by introducing additional context or adversarial tokens, leaving the core harmful intent unchanged. In this paper, we introduce ISA (Intent Shift Attack), which obfuscates LLMs about the intent of the attacks. More specifically, we establish a taxonomy of intent transformations and leverage them to generate attacks that may be misperceived by LLMs as benign requests for information. Unlike prior methods relying on complex tokens or lengthy context, our approach only needs minimal edits to the original request, and yields natural, human-readable, and seemingly harmless prompts. Extensive experiments on both open-source and commercial LLMs show that ISA achieves over 70% improvement in attack success rate compared to direct harmful prompts. More critically, fine-tuning models on only benign data reformulated with ISA templates elevates success rates to nearly 100%. For defense, we evaluate existing methods and demonstrate their inadequacy against ISA, while exploring both training-free and training-based mitigation strategies. Our findings reveal fundamental challenges in intent inference for LLMs safety and underscore the need for more effective defenses. Our code and datasets are available at https://github.com/NJUNLP/ISA.

cs.CL

Higher Satisfaction, Lower Cost: A Technical Report on How LLMs Revolutionize Meituan's Intelligent Interaction Systems

Enhancing customer experience is essential for business success, particularly as service demands grow in scale and complexity. Generative artificial intelligence and Large Language Models (LLMs) have empowered intelligent interaction systems to deliver efficient, personalized, and 24/7 support. In practice, intelligent interaction systems encounter several challenges: (1) Constructing high-quality data for cold-start training is difficult, hindering self-evolution and raising labor costs. (2) Multi-turn dialogue performance remains suboptimal due to inadequate intent understanding, rule compliance, and solution extraction. (3) Frequent evolution of business rules affects system operability and transferability, constraining low-cost expansion and adaptability. (4) Reliance on a single LLM is insufficient in complex scenarios, where the absence of multi-agent frameworks and effective collaboration undermines process completeness and service quality. (5) The open-domain nature of multi-turn dialogues, lacking unified golden answers, hampers quantitative evaluation and continuous optimization. To address these challenges, we introduce WOWService, an intelligent interaction system tailored for industrial applications. With the integration of LLMs and multi-agent architectures, WOWService enables autonomous task management and collaborative problem-solving. Specifically, WOWService focuses on core modules including data construction, general capability enhancement, business scenario adaptation, multi-agent coordination, and automated evaluation. Currently, WOWService is deployed on the Meituan App, achieving significant gains in key metrics, e.g., User Satisfaction Metric 1 (USM 1) -27.53% and User Satisfaction Metric 2 (USM 2) +25.51%, demonstrating its effectiveness in capturing user needs and advancing personalized service.

cs.CL

Expressive Value Learning for Scalable Offline Reinforcement Learning

Reinforcement learning (RL) is a powerful paradigm for learning to make sequences of decisions. However, RL has yet to be fully leveraged in robotics, principally due to its lack of scalability. Offline RL offers a promising avenue by training agents on large, diverse datasets, avoiding the costly real-world interactions of online RL. Scaling offline RL to increasingly complex datasets requires expressive generative models such as diffusion and flow matching. However, existing methods typically depend on either backpropagation through time (BPTT), which is computationally prohibitive, or policy distillation, which introduces compounding errors and limits scalability to larger base policies. In this paper, we consider the question of how to develop a scalable offline RL approach without relying on distillation or backpropagation through time. We introduce Expressive Value Learning for Offline Reinforcement Learning (EVOR): a scalable offline RL approach that integrates both expressive policies and expressive value functions. EVOR learns an optimal, regularized Q-function via flow matching during training. At inference-time, EVOR performs inference-time policy extraction via rejection sampling against the expressive value function, enabling efficient optimization, regularization, and compute-scalable search without retraining. Empirically, we show that EVOR outperforms baselines on a diverse set of offline RL tasks, demonstrating the benefit of integrating expressive value learning into offline RL.

cs.LG

Prompt Curriculum Learning for Efficient LLM Post-Training

We introduce Prompt Curriculum Learning (PCL), a lightweight reinforcement learning (RL) algorithm that selects intermediate-difficulty prompts using a learned value model to post-train language models. Since post-training LLMs via RL remains sensitive to batching and prompt selection strategies, we first conduct a series of systematic experiments where we (1) determine the optimal training batch size that balances generation efficiency and gradient quality and (2) establish the importance of focusing on prompts of intermediate difficulty for the policy. We build upon these results to design PCL, which identifies prompts of intermediate difficulty for the current policy in an on-policy manner by using a value model that is concurrently updated based on the current policy. By focusing on informative prompts that yield high effective ratios, PCL achieves either the highest performance or requires significantly less time to reach comparable performance to its counterparts. Compared to rollout-based filtering methods, PCL avoids costly rollouts and achieves $12.1\times$ and $16.9\times$ faster speed on identifying intermediate-difficulty prompts when training on MATH and DeepScaleR, respectively. We further demonstrate that our value model accurately predicts prompt difficulty and allows PCL to focus on progressively more challenging prompts during RL. Our results present a new methodology that delivers improved tradeoff between upper-bound performance and efficiency for reasoning-focused RL.

cs.LG

SDGO: Self-Discrimination-Guided Optimization for Consistent Safety in Large Language Models

Large Language Models (LLMs) excel at various natural language processing tasks but remain vulnerable to jailbreaking attacks that induce harmful content generation. In this paper, we reveal a critical safety inconsistency: LLMs can more effectively identify harmful requests as discriminators than defend against them as generators. This insight inspires us to explore aligning the model's inherent discrimination and generation capabilities. To this end, we propose SDGO (Self-Discrimination-Guided Optimization), a reinforcement learning framework that leverages the model's own discrimination capabilities as a reward signal to enhance generation safety through iterative self-improvement. Our method does not require any additional annotated data or external models during the training phase. Extensive experiments demonstrate that SDGO significantly improves model safety compared to both prompt-based and training-based baselines while maintaining helpfulness on general benchmarks. By aligning LLMs' discrimination and generation capabilities, SDGO brings robust performance against out-of-distribution (OOD) jailbreaking attacks. This alignment achieves tighter coupling between these two capabilities, enabling the model's generation capability to be further enhanced with only a small amount of discriminative samples. Our code and datasets are available at https://github.com/NJUNLP/SDGO.

cs.CL