SearcharxivSearch

arXiv subjects

Song Jiang

Publications and source records attributed to Song Jiang.

At least 19 recordsLinked to original sources

Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training

Reasoning LLMs-as-Judges, which can benefit from inference-time scaling, provide a promising path for extending the success of reasoning models to non-verifiable domains where the output correctness/quality cannot be directly checked. However, while reasoning judges have shown better performance on static evaluation benchmarks, their effectiveness in actual policy training has not been systematically examined. Therefore, we conduct a rigorous study to investigate the actual impact of non-reasoning and reasoning judges in reinforcement-learning-based LLM alignment. Our controlled synthetic setting, where a "gold-standard" judge (gpt-oss-120b) provides preference annotations to train smaller judges, reveals key differences between non-reasoning and reasoning judges: non-reasoning judges lead to reward hacking easily, while reasoning judges can lead to policies that achieve strong performance when evaluated by the gold-standard judge. Interestingly, we find that the reasoning-judge-trained policies achieve such strong performance by learning to generate highly effective adversarial outputs that can also score well on popular benchmarks such as Arena-Hard by deceiving other LLM-judges. Combined with our further analysis, our study highlights both important findings and room for improvements for applying (reasoning) LLM-judges in non-verifiable LLM post-training.

cs.AI

On the large time behavior of the 2D inhomogeneous incompressible viscous flows

This paper studies the two-dimensional inhomogeneous Navier--Stokes equations governing stratified flows in a bounded domain under a gravitational potential \(f\). Our main results are as follows. First, we provide a rigorous characterization of steady states, proving that under the Dirichlet condition \(\mathbf{u}|_{\partial \Omega} = \mathbf{0}\), all admissible equilibria are hydrostatic and satisfy \(\nabla p_s = -\rho_s \nabla f\). Second, through a perturbative analysis around arbitrary hydrostatic profiles, we show that despite possible transient growth induced by the Rayleigh--Taylor mechanism, the system always relaxes to a hydrostatic equilibrium. Third, we identify a necessary and sufficient condition on the initial density perturbation for convergence to a linear hydrostatic density profile of the form \(\rho_s = -\gamma f + \beta\), with \(\gamma > 0\) and \(\beta > 0\). Finally, we establish improved regularity estimates for strong solutions corresponding to initial data in the Sobolev space \(H^3(\Omega)\).

math.AP

Stability and bifurcation of 2D viscous primitive equations with full diffusion

This paper investigates the stability and bifurcation of the two-dimensional viscous primitive equations with full diffusion under thermal forcing. The system governs perturbations about a motionless basic state with a linear temperature profile in a periodic channel, where the temperature is fixed at $T_0$ and $T_1$ on the bottom and upper boundaries, respectively. Through a rigorous analysis of three distinct thermal regimes, we identify a critical temperature difference $T_c$ that fundamentally dictates the system's dynamical transitions. Our main contributions are fourfold. Firstly, in the subcritical case $T_0 - T_1 < T_c$, we use energy methods to establish the global nonlinear stability in $H^2$-norm, proving that perturbations decay exponentially. Secondly, precisely at the critical threshold $T_0 - T_1 = T_c$, we prove not only the nonlinear stability in $H^1$-norm but also the asymptotic convergence of all solutions to zero, leveraging spectral and dynamical systems theory. Finally, in the supercritical regime $T_0 - T_1 > T_c$, a bootstrap argument reveals that the basic state is nonlinearly unstable across all $L^p$-norms for $1 \leq p \leq \infty$. Finally, near the critical point, the dynamics are first reduced to a two-dimensional system on a center manifold. This reduced system then undergoes a supercritical bifurcation, generating a countable family of stable steady states that are organized into a local ring attractor. This work closes a significant gap in the stability analysis of the thermally driven primitive equations, establishing a rigorous mathematical foundation for understanding the formation of convection cells in large-scale geophysical flows.

math.AP

From Solving to Verifying: A Unified Objective for Robust Reasoning in LLMs

The reasoning capabilities of large language models (LLMs) have been significantly improved through reinforcement learning (RL). Nevertheless, LLMs still struggle to consistently verify their own reasoning traces. This raises the research question of how to enhance the self-verification ability of LLMs and whether such an ability can further improve reasoning performance. In this work, we propose GRPO-Verif, an algorithm that jointly optimizes solution generation and self-verification within a unified loss function, with an adjustable hyperparameter controlling the weight of the verification signal. Experimental results demonstrate that our method enhances self-verification capability while maintaining comparable performance in reasoning.

cs.LG

AI based signage classification for linguistic landscape studies

Linguistic Landscape (LL) research traditionally relies on manual photography and annotation of public signages to examine distribution of languages in urban space. While such methods yield valuable findings, the process is time-consuming and difficult for large study areas. This study explores the use of AI powered language detection method to automate LL analysis. Using Honolulu Chinatown as a case study, we constructed a georeferenced photo dataset of 1,449 images collected by researchers and applied AI for optical character recognition (OCR) and language classification. We also conducted manual validations for accuracy checking. This model achieved an overall accuracy of 79%. Five recurring types of mislabeling were identified, including distortion, reflection, degraded surface, graffiti, and hallucination. The analysis also reveals that the AI model treats all regions of an image equally, detecting peripheral or background texts that human interpreters typically ignore. Despite these limitations, the results demonstrate the potential of integrating AI-assisted workflows into LL research to reduce such time-consuming processes. However, due to all the limitations and mis-labels, we recognize that AI cannot be fully trusted during this process. This paper encourages a hybrid approach combining AI automation with human validation for a more reliable and efficient workflow.

cs.LG

SPG: Sandwiched Policy Gradient for Masked Diffusion Language Models

Diffusion large language models (dLLMs) are emerging as an efficient alternative to autoregressive models due to their ability to decode multiple tokens in parallel. However, aligning dLLMs with human preferences or task-specific rewards via reinforcement learning (RL) is challenging because their intractable log-likelihood precludes the direct application of standard policy gradient methods. While prior work uses surrogates like the evidence lower bound (ELBO), these one-sided approximations can introduce significant policy gradient bias. To address this, we propose the Sandwiched Policy Gradient (SPG) that leverages both an upper and a lower bound of the true log-likelihood. Experiments show that SPG significantly outperforms baselines based on ELBO or one-step estimation. Specifically, SPG improves the accuracy over state-of-the-art RL methods for dLLMs by 3.6% in GSM8K, 2.6% in MATH500, 18.4% in Countdown and 27.0% in Sudoku.

cs.CL

Large Reasoning Models Learn Better Alignment from Flawed Thinking

Large reasoning models (LRMs) "think" by generating structured chain-of-thought (CoT) before producing a final answer, yet they still lack the ability to reason critically about safety alignment and are easily biased when a flawed premise is injected into their thought process. We propose RECAP (Robust Safety Alignment via Counter-Aligned Prefilling), a principled reinforcement learning (RL) method for post-training that explicitly teaches models to override flawed reasoning trajectories and reroute to safe and helpful responses. RECAP trains on a mixture of synthetically generated counter-aligned CoT prefills and standard prompts, requires no additional training cost or modifications beyond vanilla reinforcement learning from human feedback (RLHF), and substantially improves safety and jailbreak robustness, reduces overrefusal, and preserves core reasoning capability -- all while maintaining inference token budget. Extensive analysis shows that RECAP-trained models engage in self-reflection more frequently and remain robust under adaptive attacks, preserving safety even after repeated attempts to override their reasoning.

cs.LG

Non-implosion mechanism of 3D incompessible Euler equations

This paper studies the non-implosion mechanism for the 3D incompressible Euler equations. We prove that vorticity blows up in finite time, whereas the $L^p_T L^\infty_{loc}$ $(p\in[1,\infty))$ norm of the velocity field remains bounded. Moreover, under an appropriate assumption on the scaling index, the exponent $p$ can be taken to be infinite. The proof is based on the introduction of a refined framework, the new observations for the null structure of transport term, and stability analysis of the self-similar model.

math.AP

VQualA 2025 Challenge on Visual Quality Comparison for Large Multimodal Models: Methods and Results

This paper presents a summary of the VQualA 2025 Challenge on Visual Quality Comparison for Large Multimodal Models (LMMs), hosted as part of the ICCV 2025 Workshop on Visual Quality Assessment. The challenge aims to evaluate and enhance the ability of state-of-the-art LMMs to perform open-ended and detailed reasoning about visual quality differences across multiple images. To this end, the competition introduces a novel benchmark comprising thousands of coarse-to-fine grained visual quality comparison tasks, spanning single images, pairs, and multi-image groups. Each task requires models to provide accurate quality judgments. The competition emphasizes holistic evaluation protocols, including 2AFC-based binary preference and multi-choice questions (MCQs). Around 100 participants submitted entries, with five models demonstrating the emerging capabilities of instruction-tuned LMMs on quality assessment. This challenge marks a significant step toward open-domain visual quality reasoning and comparison and serves as a catalyst for future research on interpretable and human-aligned quality evaluation systems.

cs.CV

Data-driven optimized high-order WENO schemes with low-dissipation and low-dispersion

Classical high-order weighted essentially non-oscillatory (WENO) schemes are designed to achieve optimal convergence order for smooth solutions and to maintain non-oscillatory behaviors for discontinuities. However, their spectral properties are not optimal, which limits the ability to capture high-frequency waves and small-scale features. In this paper, we propose a data-driven optimized method to improve the spectral properties of the WENO schemes. By analyzing the approximate dispersion relation (ADR), the spectral error of the schemes can be bounded by the reconstructed errors of a series of trigonometric functions with different wavenumbers. Therefore, we propose the new schemes WENO5-JS/Z-NN that introduce a compensation term parameterized by a neural network to the weight function of the WENO5-JS/Z schemes. The neural network is trained such that the generated weights can minimize the reconstructed errors over a large number of spatial stencils, and furthermore, improve the spectral accuracy. Meanwhile, the Total Variation Diminishing (TVD) constraint and anti-dissipation penalization are incorporated into the loss function to enhance the shock-capturing capability and preserve stability in simulating high-frequency waves. Compared to WENO5-JS/Z, our schemes maintain the ability to capture discontinuities while providing higher resolution for fine-scale flow features. The ADR indicates that the new schemes can match the exact spectrum more accurately over a broader range of wavenumbers.

math.NA

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference

Scaling inference for large language models (LLMs) is increasingly constrained by limited GPU memory, especially due to growing key-value (KV) caches required for long-context generation. While existing approaches offload KV caches to CPU memory or apply sparse attention to reduce GPU load, they often underutilize CPU compute resources and compromise accuracy. We present HGCA, a hybrid CPU-GPU attention mechanism that enables scalable, high-throughput LLM inference with near-full attention quality. HGCA performs dense attention on recently generated KV entries retained in GPU memory and parallel sparse attention on selected, salient KV entries in CPU memory. The attention outputs are efficiently merged using log-sum-exp fusion, minimizing PCIe transfer overhead. HGCA also introduces a finegrained, per-head sparsification strategy optimized for CPU execution, preserving contextual relevance while reducing computation. Our implementation seamlessly integrates into existing LLM frameworks without requiring model retraining. Experiments across diverse models and workloads show that HGCA achieves superior scalability, supports longer sequences and larger batch sizes, and outperforms existing sparse attention baselines in both performance and accuracy -- all on commodity GPU hardware.

cs.LG

FUSE: Measure-Theoretic Compact Fuzzy Set Representation for Taxonomy Expansion

Taxonomy Expansion, which models complex concepts and their relations, can be formulated as a set representation learning task. The generalization of set, fuzzy set, incorporates uncertainty and measures the information within a semantic concept, making it suitable for concept modeling. Existing works usually model sets as vectors or geometric objects such as boxes, which are not closed under set operations. In this work, we propose a sound and efficient formulation of set representation learning based on its volume approximation as a fuzzy set. The resulting embedding framework, Fuzzy Set Embedding (FUSE), satisfies all set operations and compactly approximates the underlying fuzzy set, hence preserving information while being efficient to learn, relying on minimum neural architecture. We empirically demonstrate the power of FUSE on the task of taxonomy expansion, where FUSE achieves remarkable improvements up to 23% compared with existing baselines. Our work marks the first attempt to understand and efficiently compute the embeddings of fuzzy sets.

cs.LG

100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability?

Long-context capability is considered one of the most important abilities of LLMs, as a truly long context-capable LLM enables users to effortlessly process many originally exhausting tasks -- e.g., digesting a long-form document to find answers vs. directly asking an LLM about it. However, existing real-task-based long-context evaluation benchmarks have two major shortcomings. First, benchmarks like LongBench often do not provide proper metrics to separate long-context performance from the model's baseline ability, making cross-model comparison unclear. Second, such benchmarks are usually constructed with fixed input lengths, which limits their applicability across different models and fails to reveal when a model begins to break down. To address these issues, we introduce a length-controllable long-context benchmark and a novel metric that disentangles baseline knowledge from true long-context capabilities. Experiments demonstrate the superiority of our approach in effectively evaluating LLMs.

cs.CL

Large-time behavior of solutions to the Boussinesq equations with partial dissipation and influence of rotation

This paper investigates the stability and large-time behavior of solutions to the rotating Boussinesq system under the influence of a general gravitational potential $\Psi$, which is widely used to model the dynamics of stratified geophysical fluids on the $f-$plane. Our main results are threefold: First, by imposing physically realistic boundary conditions and viscosity constraints, we prove that the solutions of the system smust necessarily take the following steady-state form $(\rho,u,v,w,p)=(\rho_s,0,v_s,0, p_s)$. These solutions are characterized by both geostrophic balance, given by $fv_s-\frac{\partial p_s}{\partial x}=\rho_s\frac{\partial \Psi}{\partial x}$ and hydrostatic balance, expressed as $-\frac{\partial p_s}{\partial z}=\rho_s\frac{\partial \Psi}{\partial z}$. Second, we establish that any steady-state solution satisfying the conditions $\nabla \rho_s=\delta (x,z)\nabla \Psi$ with $v_s(x,z)=a_0x+a_1$ is linearly unstable when the conditions $\delta(x,z)|_{(x_0,z_0)}>0$ and $(f+\alpha_0)\leq 0$ are simultaneously satisfied. This instability under the condition $\delta(x,z)|_{(x_0,z_0)}>0$ corresponds to the well-known Rayleigh-Taylor instability. Third, although the inherent Rayleigh-Taylor instability could potentially amplify the velocity around unstable steady-state solutions (heavier density over lighter one), we rigorously demonstrate that for any sufficiently smooth initial data, the solutions of the system asymptotically converge to a neighborhood of a steady-state solution in which both the zonal and vertical velocity components vanish. Finally, under a moderate additional assumption, we demonstrate that the system converges to a specific steady-state solution. In this state, the density profile is given by $\rho=-\gamma \Psi+\beta$, where $\gamma$ and $\beta$ are positive constants, and the meridional velocity $v$ depends solely and linearly on $x$ variable.

math.AP

SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Large language model (LLM) agents need to perform multi-turn interactions in real-world tasks. However, existing multi-turn RL algorithms for optimizing LLM agents fail to perform effective credit assignment over multiple turns while leveraging the generalization capabilities of LLMs and it remains unclear how to develop such algorithms. To study this, we first introduce a new benchmark, ColBench, where an LLM agent interacts with a human collaborator over multiple turns to solve realistic tasks in backend programming and frontend design. Building on this benchmark, we propose a novel RL algorithm, SWEET-RL (RL with Step-WisE Evaluation from Training-time information), that uses a carefully designed optimization objective to train a critic model with access to additional training-time information. The critic provides step-level rewards for improving the policy model. Our experiments demonstrate that SWEET-RL achieves a 6% absolute improvement in success and win rates on ColBench compared to other state-of-the-art multi-turn RL algorithms, enabling Llama-3.1-8B to match or exceed the performance of GPT4-o in realistic collaborative content creation.

cs.LG

Structure stability of steady supersonic shear flow with inflow boundary conditions

We study the existence and zero viscous limit of smooth solutions to steady compressible Navier-Stokes equations near plane shear flow between two moving parallel walls. Under the assumption $0<L\ll1$, we prove that for any plane supersonic shear flow $\mathbf{U}^0=(\mu(x_2),0)$, there exist smooth solutions near $\mathbf{U}^0$ to steady compressible Navier-Stokes equations in a 2-dimension domain $\Omega=(0,L)\times (0,2)$. Moreover, based on the uniform-in-$\varepsilon$ estimates, we establish the zero viscosity limit of the solutions obtained above to the solutions of the steady Euler equations.

math.AP

On the global stability and large time behavior of solutions of the Boussinesq equations

We study the two dimensional viscous Boussinesq equations, which model stratified flows in a circular domain under the influence of a general gravitational potential $f$. First, we show that the Boussinesq equations admit steady-state solutions only in the form of hydrostatic equilibria, $(\mathbf{u},\rho,p) = (0, \rho_s, p_s)$, where the pressure gradient satisfies $\nabla p_s = -\rho_s \nabla f$. Moreover, the relation between $\rho_s$ and $f$ is constrained by $(\partial_y \rho_s, -\partial_x \rho_s) \cdot (\partial_x f, \partial_y f) = 0$, which allows us to write $\nabla \rho_s = h(x,y) \nabla f$ for some scalar function $h(x,y)$. Second, we prove that any hydrostatic equilibrium $(0, \rho_s, p_s)$ is linearly unstable if $h(x_0, y_0) > 0$ at some point $(x, y) = (x_0, y_0)$. This instability coincides with the classical Rayleigh--Taylor instability. Third, by employing a series of regularity estimates, we reveal that although the presence of the Rayleigh--Taylor instability makes perturbations around the unstable equilibrium grow exponentially in time, the system ultimately converges to a state of hydrostatic equilibrium. The analysis is carried out for perturbations about an arbitrary hydrostatic equilibrium, covering both stable and unstable configurations. Finally, we derive a necessary and sufficient condition on the initial density perturbation under which the density converges to a profile of the form $-\gamma f + \beta$ with constants $\gamma, \beta > 0$. This result underscores the system's inherent tendency to settle into a hydrostatic state, even in the presence of Rayleigh--Taylor instability.

math.AP

NaturalReasoning: Reasoning in the Wild with 2.8M Challenging Questions

Scaling reasoning capabilities beyond traditional domains such as math and coding is hindered by the lack of diverse and high-quality questions. To overcome this limitation, we introduce a scalable approach for generating diverse and challenging reasoning questions, accompanied by reference answers. We present NaturalReasoning, a comprehensive dataset comprising 2.8 million questions that span multiple domains, including STEM fields (e.g., Physics, Computer Science), Economics, Social Sciences, and more. We demonstrate the utility of the questions in NaturalReasoning through knowledge distillation experiments which show that NaturalReasoning can effectively elicit and transfer reasoning capabilities from a strong teacher model. Furthermore, we demonstrate that NaturalReasoning is also effective for unsupervised self-training using external reward models or self-rewarding. To foster future work, we publicly release NaturalReasoning at https://huggingface.co/datasets/facebook/natural_reasoning.

cs.CL