SearcharxivSearch

arXiv subjects

Yucong Huang

Publications and source records attributed to Yucong Huang.

13 recordsLinked to original sources

Transitivity Meets Cyclicity: Explicit Preference Decomposition for Dynamic Large Language Model Alignment

Standard RLHF relies on transitive scalar rewards, failing to capture the cyclic nature of human preferences. While some approaches like the General Preference Model (GPM) address this, we identify a theoretical limitation: their implicit formulation entangles hierarchy with cyclicity, failing to guarantee dominant solutions. To address this, we propose the Hybrid Reward-Cyclic (HRC) model, which utilizes game-theoretic decomposition to explicitly disentangle preferences into orthogonal transitive (scalar) and cyclic (vector) components. Complementing this, we introduce Dynamic Self-Play Preference Optimization (DSPPO), which treats alignment as a time-varying game to progressively guide the policy toward the Nash equilibrium. Synthetic data experiments further validate HRC's structural superiority in mixed transitive--cyclic settings, where HRC converges faster and achieves higher accuracy than GPM. Experiments on RewardBench 2 demonstrate that HRC consistently improves over both BT and GPM baselines (e.g., +1.23% on Gemma-2B-it). In particular, its superior performance in the Ties domain empirically validates the model's robustness in handling complex, non-strict preferences. Extensive downstream evaluations on AlpacaEval 2.0, Arena-Hard-v0.1, and MT-Bench confirm the efficacy of our framework. Notably, when using Gemma-2B-it as the base preference model, HRC+DSPPO achieves a peak length-controlled win-rate of 44.75% on AlpacaEval 2.0 and 46.8% on Arena-Hard-v0.1, significantly outperforming SPPO baselines trained with BT or GPM. Our code is publicly available at https://github.com/lab-klc/Hybrid-Reward-Cyclic.

cs.CL

CreativeGame:Toward Mechanic-Aware Creative Game Generation

Large language models can generate plausible game code, but turning this capability into \emph{iterative creative improvement} remains difficult. In practice, single-shot generation often produces brittle runtime behavior, weak accumulation of experience across versions, and creativity scores that are too subjective to serve as reliable optimization signals. A further limitation is that mechanics are frequently treated only as post-hoc descriptions, rather than as explicit objects that can be planned, tracked, preserved, and evaluated during generation. This report presents \textbf{CreativeGame}, a multi-agent system for iterative HTML5 game generation that addresses these issues through four coupled ideas: a proxy reward centered on programmatic signals rather than pure LLM judgment; lineage-scoped memory for cross-version experience accumulation; runtime validation integrated into both repair and reward; and a mechanic-guided planning loop in which retrieved mechanic knowledge is converted into an explicit mechanic plan before code generation begins. The goal is not merely to produce a playable artifact in one step, but to support interpretable version-to-version evolution. The current system contains 71 stored lineages, 88 saved nodes, and a 774-entry global mechanic archive, implemented in 6{,}181 lines of Python together with inspection and visualization tooling. The system is therefore substantial enough to support architectural analysis, reward inspection, and real lineage-level case studies rather than only prompt-level demos. A real 4-generation lineage shows that mechanic-level innovation can emerge in later versions and can be inspected directly through version-to-version records. The central contribution is therefore not only game generation, but a concrete pipeline for observing progressive evolution through explicit mechanic change.

cs.AI

Fusion-PSRO: Nash Policy Fusion for Policy Space Response Oracles

For solving zero-sum games involving non-transitivity, a useful approach is to maintain a policy population to approximate the Nash Equilibrium (NE). Previous studies have shown that the Policy Space Response Oracles (PSRO) algorithm is an effective framework for solving such games. However, current methods initialize a new policy from scratch or inherit a single historical policy in Best Response (BR), missing the opportunity to leverage past policies to generate a better BR. In this paper, we propose Fusion-PSRO, which employs Nash Policy Fusion to initialize a new policy for BR training. Nash Policy Fusion serves as an implicit guiding policy that starts exploration on the current Meta-NE, thus providing a closer approximation to BR. Moreover, it insightfully captures a weighted moving average of past policies, dynamically adjusting these weights based on the Meta-NE in each iteration. This cumulative process further enhances the policy population. Empirical results on classic benchmarks show that Fusion-PSRO achieves lower exploitability, thereby mitigating the shortcomings of previous research on policy initialization in BR.

cs.GT

On the stability of the spherically symmetric solution to an inflow problem for an isentropic model of compressible viscous fluid

We investigate an inflow problem for the multi-dimensional isentropic compressible Navier-Stokes equations. The fluid under consideration occupies the exterior domain of unit ball, $Ω=\{x\in\mathbb{R}^n\,\vert\, |x|\ge 1\}$, and a constant stream of mass is flowing into the domain from the boundary $\partialΩ=\{|x|=1\}$. It is shown in Hashimoto-Matsumura(2021) that if the fluid velocity at the far-field is assumed to be zero, then there exists a unique spherically symmetric stationary solution, denoted as $(\tildeρ,\tilde{u})(r)$ with $r\equiv |x|$. In this paper, we show that either $\tildeρ$ is monotone increasing or $\tildeρ$ attains a unique global minimum, and this is classified by the boundary condition of density. In addition, we also derive a set of spatial decay rates for $(\tildeρ,\tilde{u})$ which allows us to prove the time-asymptotic stability of $(\tildeρ,\tilde{u})$ using the energy method. More specifically, we prove this under small initial perturbation on $(\tildeρ,\tilde{u})$, provided that the density at the far-field is supposed to be strictly positive but suitably small, in other words, the far-field state of the fluid is not vacuum but suitably rarefied. The main difficulty for the proof is the boundary terms that appears in the a-priori estimates. We resolve this issue by reformulating the problem in Lagrangian coordinate system.

math.AP

Asymptotic Stability of 3D Out-flowing Compressible Viscous Fluid under Non-Spherical Perturbation

We study an outflow problem for the $3$-dimensional isentropic compressible Navier-Stokes equations. The fluid under consideration occupies the exterior domain of the unit ball $Ω=\{x\in\mathbb{R}^3\,\vert\, |x|\ge 1\}$ and it is flowing out from the unit ball $Ω$ at a constant speed $|u_b|$, in the normal direction to the boundary surface $\partialΩ$. The existence of a unique spherically symmetric stationary solution $(\tildeρ,\tilde{u})$ is obtained by I.~Hashimoto and A.~Matsumura in 2021, provided that the fluid velocity at the far-field is assumed to be zero, and $|u_b|$ is sufficiently small. Subsequently, authors of the present article prove in 2024 that $(\tildeρ,\tilde{u})$ is time-asymptotically stable under large spherically symmetric initial perturbations in the suitable Sobolev norm. The main purpose of the present paper is to investigate the case when the initial perturbations are possibly non-spherically symmetric. We show that $(\tildeρ,\tilde{u})$ remains asymptotically stable in time, under general small initial perturbations in the $H^3$-norm.

math.AP

A Monotonicity formula for almost self-similar suitable weak solutions to the stationary Navier-Stokes equations in $\mathbb R^5$

In this paper we show that a suitable weak solution to the stationary Navier-Stokes system in $\mathbb R^5$, cannot behave like a self-similar function of degree negative one if the lower limit of the local Reynolds number is finite. To prove the result we develop a method that uses a monotonicity formula approach, classification of homogenous solutions to the incompressible Euler equations in $\mathbb R^5$, and a projection theorem.

math.AP

Conflux-PSRO: Effectively Leveraging Collective Advantages in Policy Space Response Oracles

Policy Space Response Oracle (PSRO) with policy population construction has been demonstrated as an effective method for approximating Nash Equilibrium (NE) in zero-sum games. Existing studies have attempted to improve diversity in policy space, primarily by incorporating diversity regularization into the Best Response (BR). However, these methods cause the BR to deviate from maximizing rewards, easily resulting in a population that favors diversity over performance, even when diversity is not always necessary. Consequently, exploitability is difficult to reduce until policies are fully explored, especially in complex games. In this paper, we propose Conflux-PSRO, which fully exploits the diversity of the population by adaptively selecting and training policies at state-level. Specifically, Conflux-PSRO identifies useful policies from the existing population and employs a routing policy to select the most appropriate policies at each decision point, while simultaneously training them to enhance their effectiveness. Compared to the single-policy BR of traditional PSRO and its diversity-improved variants, the BR generated by Conflux-PSRO not only leverages the specialized expertise of diverse policies but also synergistically enhances overall performance. Our experiments on various environments demonstrate that Conflux-PSRO significantly improves the utility of BRs and reduces exploitability compared to existing methods.

cs.GT

The Dirichlet-Neumann Operator for Taylor's Cone

The aim of this paper is to analyse the Dirichlet-Neumann operator in axially symmetric conical domains. We provide a constructive treatment of the generic singularity at the vertex by using a new coordinate system that maps the conical domain to a strip. Building upon the paradifferential theory, we then establish our main Sobolev estimates. We also find the shape derivative, the linearization formula, and the cancellation property for the Dirichlet-Neumann operator. Our results can be viewed as the first step towards establishing the mathematical framework for the perturbations of Taylor's cone which appears in the jet break-up control.

math.AP

The Well-posedness of Cylindrical Jets with Surface Tension

In 1879 Rayleigh \cite{Rayleigh} studied the stability of infinite cylindrical jets, inspired by the experiments of Plateau \cite{Plateau}. The principal question that Rayleigh asked is: under what circumstances the jet is stable, for small displacements. In this paper we show that the jet flow is well-posed in short time if the initial condition belongs to some Sobolev space, and the initial jet boundary remains uniformly bounded away from the axis of symmetry. This will be proved by the method of paradifferential calculus and paralinearization. The salient feature of these results is that no smallness assumption is imposed on the initial condition.

math.AP

Large-time behaviour of the spherically symmetric solution to an outflow problem for isentropic model of compressible viscous fluid

We study the large time behaviour of a spherically symmetric motion of out-flowing isentropic and compressible viscous gas. The fluid occupies an unbounded exterior domain in $\mathbb{R}^n \; (n \ge 2)$, and it flows out from an inner sphere centred at the origin of radius $r=1$. The unique existence of a stationary solution satisfying the outflow boundary condition has been obtained by I. Hashimoto and A. Matsumura in 2021. The main aim of present paper is to show that this stationary solution becomes a time asymptotic state to the initial boundary value problem with the same boundary and spatial asymptotic conditions. Here, the initial data is chosen arbitrarily large if it belongs to the suitable weighted Sobolev space. The main strategy is to approximate the unbounded exterior problem by solving a sequence of outflow-inflow initial boundary value problems posed in finite annular domain. Then the solution is obtained as a limit of these approximate solutions. The key argument for the stability theorem is based on the derivation of a-priori estimates in the weighted Sobolev space, executed under the Lagrangian coordinate. The essential step of the proof is to obtain the point-wise upper and lower bound for the density. It is derived through employing a representation formula of the density with the aid of the weighted energy method.

math.AP

Global Spherically Symmetric Solutions of the Multidimensional Full Compressible Navier-Stokes Equations with Large Data

We establish the global-in-time existence of solutions of the Cauchy problem for the full Navier-Stokes equations for compressible heat-conducting flow in multidimensions with initial data that are large, discontinuous, spherically symmetric, and away from the vacuum. The solutions obtained here are of global finite total relative-energy including the origin, while cavitation may occur as balls centred at the origin of symmetry for which the interfaces between the fluid and the vacuum must be upper semi-continuous in space-time in the Eulerian coordinates. On any region strictly away from the possible vacuum, the velocity and specific internal energy are Hölder continuous, and the density has a uniform upper bound. To achieve these, our main strategy is to regard the Cauchy problem as the limit of a series of carefully designed initial-boundary value problems that are formulated in finite annular regions. For such approximation problems, we can derive uniform {\it a-priori} estimates that are independent of both the inner and outer radii of the annuli considered in the spherically symmetric Lagrangian coordinates. The entropy inequality is recovered after taking the limit of the outer radius to infinity by using Mazur's lemma and the convexity of the entropy function, which is required for the limit of the inner radius tending to zero. Then the global weak solutions of the original problem are attained via careful compactness arguments applied to the approximate solutions in the Eulerian coordinates.

math.AP

Globally Optimal Departure Rates for Several Groups of Drivers

The first part of this paper contains a brief introduction to conservation law models of traffic flow on a network of roads. Globally optimal solutions and Nash equilibrium solutions are reviewed, with several groups of drivers sharing different cost functions. In the second part we consider a globally optimal set of departure rates, for different groups of drivers but on a single road. Necessary conditions are proved, which lead to a practical algorithm for computing the optimal solution.

math.AP