SearcharxivSearch

arXiv subjects

Xiang Yu

Publications and source records attributed to Xiang Yu.

At least 37 records · Page 2Linked to original sources

Mean Field Game of Controls with State Reflections: Existence and Limit Theory

This paper studies mean field game (MFG) of controls by featuring the state-control joint distribution and the reflected state process at an exogenous stochastic reflection boundary. We contribute to the literature with a customized relaxed formulation and some new compactification arguments to establish the existence of a Markovian mean field equilibrium (MFE) in the weak sense. We consider an enlarged canonical space, utilizing the dynamic Skorokhod mapping, to accommodate the stochastic reflection boundary process. A fixed-point argument on the extended space using an extension transformation technique is developed to tackle challenges from the joint measure flow of the state and the relaxed control that may not be continuous. Furthermore, the bidirectional connections between the MFG and the $N$-player game are established. We first show that any weak limit of empirical measures induced by $\boldsymbolε$-Nash equilibria in $N$-player games must be supported exclusively on the set of relaxed mean field equilibria, analogous to the propagation of chaos in mean field control problems. We then prove the convergence result that a Markovian MFE in the weak sense can be approximated by a sequence of constructed $\boldsymbolε$-Nash equilibria in the weak sense in $N$-player games when $N$ tends to infinity.

math.OC

Equilibrium for Time-inconsistent Mean Field Games: A Systematic Analysis by Entropy Regularization

This paper studies the existence and approximation of equilibria for general time-inconsistent mean field game (MFG) problems in continuous time. To handle the intricate nonlocal equilibrium Hamilton-Jacobi-Bellman (EHJB) system arising from initial-time dependence, such as non-exponential discounting, we develop a vanishing entropy regularization approach. Using entropy regularization, we first characterize the regularized equilibrium through a coupled exploratory equilibrium HJB (EEHJB) equation and a law-dependent stochastic differential equation. By exploiting Schauder fixed-point arguments and tailored parabolic regularity estimates in a suitable functional space involving both value functions and measure flows, we establish the global existence of regularized equilibria under mild assumptions. We then establish convergence as the entropy regularization vanishes. By employing compactness arguments, Young measure techniques, and a duality tool for divergence-form Fokker-Planck equations, we prove that the regularized equilibria converge, up to subsequences, to a mean-field equilibrium of the original MFG. Furthermore, under entropy regularization, we propose a policy iteration algorithm and establish its convergence under short-time-horizon and weak-terminal-interaction conditions.

math.OC

Mean Field Competition of Optimal Switching: The Vanishing Entropy Regularization Approach

This paper studies a type of rank-based mean field game in which competing agents strategically switch among multiple effort regimes. We propose an entropy regularized auxiliary problem where the switching decisions are randomized to the control of transition probability for a continuous-time finite-state Markov chain. We first establish the existence of regularized equilibrium in this auxiliary problem. Assuming the convexity of reward scheme, we then prove that the equilibrium is unique and can be approximated by a fictitious play iteration scheme. Furthermore, as the entropy regularization vanishes, we establish the convergence analysis of the regularized equilibrium towards the relaxed equilibrium in the original MFG of optimal switching. The uniqueness of the population ranking distribution under the relaxed equilibrium is also obtained given a strictly convex reward scheme.

math.OC

Mean-field games with rough common noise: the compactification approach

We study mean-field game (MFG) problems with rough common noise, in which the representative state dynamics are governed by a controlled rough stochastic differential equation driven by an idiosyncratic Brownian motion and a deterministic rough-path signal that affects the whole population. Within this new framework, we introduce a canonical weak formulation based on relaxed controls and rough martingale problems. We prove the existence of a pathwise mean-field equilibrium by developing new compactification tools that accommodate rough integration and differ substantially from classical compactification arguments in the literature. Finally, we discuss the relationship between the pathwise problem and the classical MFG problem with randomized Brownian common noise. Using the notion of a pathwise admissible set, we recast mean-field game problems with common noise as optimization problems over an extended space of probability measures. We establish an equivalent characterization of Carmona-Delarue-Lacker's weak equilibrium and, as an application, give an alternative proof of strong equilibrium without first establishing pathwise uniqueness.

math.PR

Interpretation of $Ω(2012)$ as a $Ξ(1530)K$ molecular state

We investigate the mass and strong decay properties of the $Ω(2012)$ resonance using QCD sum rules, assuming it to be an S-wave $Ξ(1530)\bar{K}$ molecular pentaquark state with $I(J^{P})= 0(\frac{3}{2}^{-})$. A unified interpolating current is constructed, and the two-point correlation functions and three-point functions are calculated up to dimension-13 and 10 condensate terms in the OPE series, respectively. The negative-parity contribution is isolated by employing parity-projected sum rules. The two-body strong decays to $Ξ^0 K^-$ and $Ξ^- \bar{K}^0$ are studied via their three-point correlation functions. Our analysis yields a mass of $2.00 \pm 0.15~\mathrm{GeV}$ and a total two-body decay width of $Γ= 0.96^{+0.79}_{-0.41}~\mathrm{MeV}$ for the $Ξ(1530)\bar{K}$ molecular state. The ratio of the two-body decay branching fractions is obtained as $\mathcal{R}^{Ξ^- \bar{K}^0}_{Ξ^0 K^-} = 0.85$. These results are compatible with the experimental data for the $Ω(2012)$ within uncertainties and support its interpretation as a $Ξ(1530)\bar{K}$ molecular pentaquark state.

hep-ph

AlpsBench: An LLM Personalization Benchmark for Real-Dialogue Memorization and Preference Alignment

As Large Language Models (LLMs) evolve into lifelong AI assistants, LLM personalization has become a critical frontier. However, progress is currently bottlenecked by the absence of a gold-standard evaluation benchmark. Existing benchmarks either overlook personalized information management that is critical for personalization or rely heavily on synthetic dialogues, which exhibit an inherent distribution gap from real-world dialogue. To bridge this gap, we introduce AlpsBench, An LLM PerSonalization benchmark derived from real-world human-LLM dialogues. AlpsBench comprises 2,500 long-term interaction sequences curated from WildChat, paired with human-verified structured memories that encapsulate both explicit and implicit personalization signals. We define four pivotal tasks - personalized information extraction, updating, retrieval, and utilization - and establish protocols to evaluate the entire lifecycle of memory management. Our benchmarking of frontier LLMs and memory-centric systems reveals that: (i) models struggle to reliably extract latent user traits; (ii) memory updating faces a performance ceiling even in the strongest models; (iii) retrieval accuracy declines sharply in the presence of large distractor pools; and (iv) while explicit memory mechanisms improve recall, they do not inherently guarantee more preference-aligned or emotionally resonant responses. AlpsBench aims to provide a comprehensive framework.

cs.CL

Continuous-time q-learning for mean-field control with common noise, part-I: Theoretical foundations

This paper investigates the continuous-time counterpart of the Q-function for entropy-regularized mean-field control (MFC) with controlled common noise, coined as q-function by Jia and Zhou (2023) in the single agent's model. We first show that, under discretely sampled actions, the value function in the exploratory formulation converges to the one in the relaxed control formulation as the time grid refines. Leveraging the relaxed control formulation, we derive the exploratory Hamilton-Jacobi-Bellman (HJB) equation, in which the controlled common noise gives rise to an additional nonlinear functional of policy, rendering the policy iteration intricate. Under certain concavity condition, we establish the existence and uniqueness of the optimal one-step policy iteration via a first-order condition using the partial linear functional derivative with respect to policy. The policy improvement at each iteration is verified by relating to an entropy-regularized optimization problem over the space of policies. In the mean-field setting, we introduce the integrated q-function (Iq-function) defined on the state distribution and the policy, and it is shown that an optimal policy is identified as a two-layer fixed point to the argmax operator of the Iq-function. Finally, we provide the explicit characterization of an optimal policy as a Gaussian distribution in the general linear-quadratic (LQ) setting.

math.OC

Continuous-time q-learning for mean-field control with common noise, part-II: q-learning algorithms

This paper is a continuation work of Ren et al. (2026) aiming to further devise q-learning algorithms for mean-field control (MFC) with controlled common noise. Based on the relaxed control formulation, we first establish the martingale condition of the value function and the Iq-function by evaluating along the conditional state distributions generated by all test policies. As the data in the relaxed control formulation are not observable in practice, we quantify the error incurred when they are replaced by the observable ones in the exploratory formulation under discretely sampled actions. This, together with a two-layer fixed point characterization of an optimal policy in Ren et al. (2026), allows us to propose several algorithms including the Actor-Critic q-learning algorithm, in which the policy is updated in the Actor-step based on the iteration rule induced by the improved Iq-function, and the value function and Iq-function are updated in the Critic-step based on the martingale orthogonality condition using the data from the exploratory formulation. We also establish the convergence of the inner iterations in the Actor-step in an infinite-horizon linear quadratic (LQ) framework. In two examples, within and beyond LQ framework, our q-learning algorithms are implemented with satisfactory performance.

math.OC

Constrained mean-field control with singular controls: Existence, stochastic maximum principle and constrained FBSDE

This paper studies a class of mean-field control (MFC) problems with singular controls under general dynamic state-control-law constraints. We first propose a customized relaxed control formulation to cope with the dynamic mixed constraints and establish the existence of an optimal control using compactification argument in the proper canonical spaces to accommodate singular controls. To further characterize the optimal pair of regular and singular controls, we treat the controlled McKean-Vlasov process as an infinite-dimensional equality constraint and recast the MFC problem as an optimization problem on canonical spaces with constraints on Banach space, allowing us to derive the stochastic maximum principle (SMP) and a class of constrained BSDE using a new Lagrange multipliers method. Additionally, we investigate the uniqueness and the stability result of the solution to the constrained FBSDE associated with the constrained MFC with singular controls.

math.OC

Extended mean-field control problems with Poissonian common noise: Stochastic maximum principle and Hamiltonian-Jacobi-Bellman equation

This paper studies mean-field control problems with state-control joint law dependence and Poissonian common noise. We develop the stochastic maximum principle (SMP) and establish its connection to the Hamiltonian-Jacobi-Bellman (HJB) equation on the Wasserstein space. The presence of the conditional joint law and its discontinuity under Poissonian common noise bring new technical challenges. To develop the SMP when the control domain is not necessarily convex, we first consider a strong relaxed control formulation that allows us to perform the first-order variation. We propose the technique of extension transformation to overcome the compatibility issues arising from the joint law in the relaxed control formulation. By further establishing the equivalence between the relaxed control and the strict control formulations, we obtain the SMP for the original problem with strict controls. In the part to investigate the HJB equation, we formulate an equivalent controlled Fokker-Planck problem subjecting to a controlled measure-valued dynamics with Poisson jumps, which allows us to derive the HJB equation of the original problem under open-loop strict controls. We also establish the connection between the SMP and the HJB equation.

math.OC

Extended mean-field control under constraints: The generalized Fritz-John conditions and Lagrangian method

This paper studies mean-field control with joint law dependence under dynamic expectation constraints and/or dynamic state-control-law constraints. We pioneer the establishment of the stochastic maximum principle (SMP) and the derivation of the backward SDE (BSDE) from the perspective of constrained optimization using the method of Lagrangian multipliers. We first propose to embed the constrained mean-field control (C-MFC) with joint-law dependence into some abstract optimization problems with constraints on Banach spaces, for which we develop the generalized Fritz-John (FJ) optimality conditions. We then prove the stochastic maximum principle (SMP) for C-MFC by transforming the FJ conditions into an equivalent stochastic first-order condition associated with a general type of constrained forward-backward SDEs (FBSDEs). Contrary to the existing literature, we treat the McKean-Vlasov SDE as an infinite-dimensional equality constraint such that the BSDE induced by the FJ first-order optimality conditions can be interpreted as the generalized Lagrange multiplier. We also employ the methodology to stochastic control and mean field game problems under dynamic constraints.

math.OC

Mean Field Game of Optimal Tracking Portfolio

This paper studies the mean field game (MFG) problem arising from a large population competition in fund management, featuring a new type of relative performance via the benchmark tracking. In the $n$-player model, each agent aims to minimize the expected largest shortfall of the wealth with reference to the benchmark process, which is modeled by a linear combination of the population's average wealth process and a market index process. With a continuum of agents, we formulate the MFG problem with a reflected state process. We establish the existence of the mean field equilibrium (MFE) using the partial differential equation (PDE) approach. Firstly, by applying the dual transform, the best response control of the representative agent can be characterized in analytical form in terms of a dual reflected diffusion process. As a novel contribution, we verify the consistency condition of the MFE in separated domains with the help of the duality relationship and properties of the dual process. Moreover, based on the MFE, we construct an approximate Nash equilibrium for the $n$-player game when the number $n$ is sufficiently large.

math.OC

Equilibrium under Time-Inconsistency: A New Existence Theory by Vanishing Entropy Regularization

This paper develops a framework for establishing the existence of solutions to the equilibrium Hamilton-Jacobi-Bellman (EHJB) equation arising in time-inconsistent stochastic control problems. The time-inconsistency in our setting arises from the initial-time dependence such as the non-exponential discounting. The classical approach typically relates the existence of equilibrium to the classical solution of the EHJB, whose existence is still an open problem under general model assumptions. We resolve this challenge by building on a vanishing entropy regularization approach. Using fixed-point arguments, we first establish the existence of classical solutions to the exploratory equilibrium Hamilton-Jacobi-Bellman Equation (EEHJB) by deriving a series of delicate PDE estimates for the solution and its derivatives. Building on these estimates for the solution of the EEHJB and its derivatives, we then conduct a rigorous convergence analysis under suitable norms as the entropy regularization vanishes. Our main result shows that solutions of the EEHJB converge to a strong solution of the original EHJB, corresponding to the limit of the regularized equilibria. This convergence yields a verification argument ensuring that the limiting relaxed equilibrium indeed constitutes an equilibrium for the original time-inconsistent control problem. We thus establish the well-posedness of the EHJB and the existence of equilibria in diffusion models under time-inconsistency, without resorting to conventional stringent regularity assumptions of the EHJB.

math.OC

PepThink-R1: LLM for Interpretable Cyclic Peptide Optimization with CoT SFT and Reinforcement Learning

Designing therapeutic peptides with tailored properties is hindered by the vastness of sequence space, limited experimental data, and poor interpretability of current generative models. To address these challenges, we introduce PepThink-R1, a generative framework that integrates large language models (LLMs) with chain-of-thought (CoT) supervised fine-tuning and reinforcement learning (RL). Unlike prior approaches, PepThink-R1 explicitly reasons about monomer-level modifications during sequence generation, enabling interpretable design choices while optimizing for multiple pharmacological properties. Guided by a tailored reward function balancing chemical validity and property improvements, the model autonomously explores diverse sequence variants. We demonstrate that PepThink-R1 generates cyclic peptides with significantly enhanced lipophilicity, stability, and exposure, outperforming existing general LLMs (e.g., GPT-5) and domain-specific baseline in both optimization success and interpretability. To our knowledge, this is the first LLM-based peptide design framework that combines explicit reasoning with RL-driven property control, marking a step toward reliable and transparent peptide optimization for therapeutic discovery.

cs.LG

Policy Iteration Achieves Regularized Equilibrium under Time Inconsistency

For a general entropy-regularized time-inconsistent stochastic control problem, we propose a policy iteration algorithm (PIA) and establish its convergence to an equilibrium policy with an exponential convergence rate. The design of the PIA is based on a coupled system of non-local partial differential equations, called the exploratory equilibrium Hamilton--Jacobi--Bellman (EEHJB) equation. As opposed to the standard time-consistent case, policy improvement fails in general and the target value function (now an equilibrium value function) is not even known to exist a priori. To overcome these, we prove that the value functions generated by the PIA form a Cauchy sequence in a specialized Banach space, hence admit a limit, and the rate of convergence is exponential, on the strength of the Bismut--Elworthy--Li formula of stochastic representation. The limiting value function is shown to fulfill the EEHJB equation, which induces an equilibrium policy in a Gibbs form. Such convergence in value additionally implies uniform convergence of the generated policies to the equilibrium policy, again with an exponential rate. As a byproduct, the PIA gives a constructive proof of the global existence and uniqueness of a classical solution to our general EEHJB equation, whose well-posedness has not been explored in the literature.

math.OC

Optimal consumption under loss-averse multiplicative habit-formation preferences

This paper studies a loss-averse version of the multiplicative habit formation preference and the corresponding optimal investment and consumption strategies over an infinite horizon. The agent's consumption preference is depicted by a general S-shaped utility function of her consumption-to-habit ratio. By considering the concave envelope of the S-shaped utility and the associated dual value function, we provide a thorough analysis of the HJB equation for the concavified problem via studying a related nonlinear free boundary problem. Based on established properties of the solution to this free boundary problem, we obtain the optimal consumption and investment policies in feedback form. Some new and technical verification arguments are developed to cope with generality of the utility function. The equivalence between the original problem and the concavified problem readily follows from the structure of the feedback controls. We also discuss some quantitative properties of the optimal policies, complemented by illustrative numerical examples and their financial implications.

q-fin.MF

Linking Knowledge to Care: Knowledge Graph-Augmented Medical Follow-Up Question Generation

Clinical diagnosis is time-consuming, requiring intensive interactions between patients and medical professionals. While large language models (LLMs) could ease the pre-diagnostic workload, their limited domain knowledge hinders effective medical question generation. We introduce a Knowledge Graph-augmented LLM with active in-context learning to generate relevant and important follow-up questions, KG-Followup, serving as a critical module for the pre-diagnostic assessment. The structured medical domain knowledge graph serves as a seamless patch-up to provide professional domain expertise upon which the LLM can reason. Experiments demonstrate that KG-Followup outperforms state-of-the-art methods by 5% - 8% on relevant benchmarks in recall.

cs.CL

Retrieval-Augmented Foundation Models for Matched Molecular Pair Transformations to Recapitulate Medicinal Chemistry Intuition

Matched molecular pairs (MMPs) capture the local chemical edits that medicinal chemists routinely use to design analogs, but existing ML approaches either operate at the whole-molecule level with limited edit controllability or learn MMP-style edits from restricted settings and small models. We propose a variable-to-variable formulation of analog generation and train a foundation model on large-scale MMP transformations (MMPTs) to generate diverse variables conditioned on an input variable. To enable practical control, we develop prompting mechanisms that let the users specify preferred transformation patterns during generation. We further introduce MMPT-RAG, a retrieval-augmented framework that uses external reference analogs as contextual guidance to steer generation and generalize from project-specific series. Experiments on general chemical corpora and patent-specific datasets demonstrate improved diversity, novelty, and controllability, and show that our method recovers realistic analog structures in practical discovery scenarios.

cs.LG