SearcharxivSearch

arXiv subjects

Xiang Yu

Publications and source records attributed to Xiang Yu.

At least 55 records · Page 3Linked to original sources

Continuous-time q-Learning for Jump-Diffusion Models under Tsallis Entropy

This paper studies the continuous-time reinforcement learning in jump-diffusion models by featuring the q-learning (the continuous-time counterpart of Q-learning) under Tsallis entropy regularization. Contrary to the Shannon entropy, the general form of Tsallis entropy renders the optimal policy not necessarily a Gibbs measure. Herein, the Lagrange multiplier and KKT condition are needed to ensure that the learned policy is a probability density function. As a consequence, the characterization of the optimal policy using the q-function also involves a Lagrange multiplier. In response, we establish the martingale characterization of the q-function and devise two q-learning algorithms depending on whether the Lagrange multiplier can be derived explicitly or not. In the latter case, we consider different parameterizations of the optimal q-function and the optimal policy, and update them alternatively in an Actor-Critic manner. We also study two numerical examples, namely, an optimal liquidation problem in dark pools and a non-LQ control problem. It is interesting to see therein that the optimal policies under the Tsallis entropy regularization can be characterized explicitly, which are distributions concentrated on some compact support. The satisfactory performance of our q-learning algorithms is illustrated in each example.

math.OC

Learning-based Observer for Coupled Disturbance

Achieving high-precision control for robotic systems is hindered by the low-fidelity dynamical model and external disturbances. Especially, the intricate coupling between internal uncertainties and external disturbances further exacerbates this challenge. This study introduces an effective and convergent algorithm enabling accurate estimation of the coupled disturbance via combining control and learning philosophies. Concretely, by resorting to Chebyshev series expansion, the coupled disturbance is firstly decomposed into an unknown parameter matrix and two known structures dependent on system state and external disturbance respectively. A regularized least squares algorithm is subsequently formalized to learn the parameter matrix using historical time-series data. Finally, a polynomial disturbance observer is specifically devised to achieve a high-precision estimation of the coupled disturbance by utilizing the learned portion. The proposed algorithm is evaluated through extensive simulations and real flight tests. We believe this work can offer a new pathway to integrate learning approaches into control frameworks for addressing longstanding challenges in robotic applications.

cs.RO

Case-Guided Sequential Assay Planning in Drug Discovery

Optimally sequencing experimental assays in drug discovery is a high-stakes planning problem under severe uncertainty and resource constraints. A primary obstacle for standard reinforcement learning (RL) is the absence of an explicit environment simulator or transition data $(s, a, s')$; planning must rely solely on a static database of historical outcomes. We introduce the Implicit Bayesian Markov Decision Process (IBMDP), a model-based RL framework designed for such simulator-free settings. IBMDP constructs a case-guided implicit model of transition dynamics by forming a nonparametric belief distribution using similar historical outcomes. This mechanism enables Bayesian belief updating as evidence accumulates and employs ensemble MCTS planning to generate stable policies that balance information gain toward desired outcomes with resource efficiency. We validate IBMDP through comprehensive experiments. On a real-world central nervous system (CNS) drug discovery task, IBMDP reduced resource consumption by up to 92\% compared to established heuristics while maintaining decision confidence. To rigorously assess decision quality, we also benchmarked IBMDP in a synthetic environment with a computable optimal policy. Our framework achieves significantly higher alignment with this optimal policy than a deterministic value iteration alternative that uses the same similarity-based model, demonstrating the superiority of our ensemble planner. IBMDP offers a practical solution for sequential experimental design in data-rich but simulator-poor domains.

cs.LG

Transformer-Based Approach for Automated Functional Group Replacement in Chemical Compounds

Functional group replacement is a pivotal approach in cheminformatics to enable the design of novel chemical compounds with tailored properties. Traditional methods for functional group removal and replacement often rely on rule-based heuristics, which can be limited in their ability to generate diverse and novel chemical structures. Recently, transformer-based models have shown promise in improving the accuracy and efficiency of molecular transformations, but existing approaches typically focus on single-step modeling, lacking the guarantee of structural similarity. In this work, we seek to advance the state of the art by developing a novel two-stage transformer model for functional group removal and replacement. Unlike one-shot approaches that generate entire molecules in a single pass, our method generates the functional group to be removed and appended sequentially, ensuring strict substructure-level modifications. Using a matched molecular pairs (MMPs) dataset derived from ChEMBL, we trained an encoder-decoder transformer model with SMIRKS-based representations to capture transformation rules effectively. Extensive evaluations demonstrate our method's ability to generate chemically valid transformations, explore diverse chemical spaces, and maintain scalability across varying search sizes.

cs.LG

Incremental equations in curvature-dependent surface elasticity

We develop a general incremental framework for hyperelastic solids whose surfaces exhibit both stretch-dependent and curvature-dependent elastic behavior. Building upon a variational formulation of curvature-dependent surface elasticity, we derive compact governing equations expressed in a coordinate-free Lagrangian setting that remain valid for arbitrary geometries. Linearization about an arbitrarily large finite deformation yields incremental bulk and surface balance laws that closely resemble the classical small-on-large theory, but are now extended to include surface-curvatureinduced stresses. The applicability of the general theory is demonstrated by analyzing the onset of periodic beading in a soft cylindrical substrate coated with a surface layer exhibiting stretching- or curvature-dependent behavior, illustrating how surface stretching and bending effects influence instability thresholds for both compressible and incompressible bulk. This unified formulation thus provides a foundation for studying stability phenomena in elasto-capillary systems where surface curvature plays a critical mechanical role.

math-ph

Optimizing Control-Friendly Trajectories with Self-Supervised Residual Learning

Real-world physics can only be analytically modeled with a certain level of precision for modern intricate robotic systems. As a result, tracking aggressive trajectories accurately could be challenging due to the existence of residual physics during controller synthesis. This paper presents a self-supervised residual learning and trajectory optimization framework to address the aforementioned challenges. At first, unknown dynamic effects on the closed-loop model are learned and treated as residuals of the nominal dynamics, jointly forming a hybrid model. We show that learning with analytic gradients can be achieved using only trajectory-level data while enjoying accurate long-horizon prediction with an arbitrary integration step size. Subsequently, a trajectory optimizer is developed to compute the optimal reference trajectory with the residual physics along it minimized. It ends up with trajectories that are friendly to the following control level. The agile flight of quadrotors illustrates that by utilizing the hybrid dynamics, the proposed optimizer outputs aggressive motions that can be precisely tracked.

cs.RO

Unified Meta-Representation and Feedback Calibration for General Disturbance Estimation

Precise control in modern robotic applications is always an open issue due to unknown time-varying disturbances. Existing meta-learning-based approaches require a shared representation of environmental structures, which lack flexibility for realistic non-structural disturbances. Besides, representation error and the distribution shifts can lead to heavy degradation in prediction accuracy. This work presents a generalizable disturbance estimation framework that builds on meta-learning and feedback-calibrated online adaptation. By extracting features from a finite time window of past observations, a unified representation that effectively captures general non-structural disturbances can be learned without predefined structural assumptions. The online adaptation process is subsequently calibrated by a state-feedback mechanism to attenuate the learning residual originating from the representation and generalizability limitations. Theoretical analysis shows that simultaneous convergence of both the online learning error and the disturbance estimation error can be achieved. Through the unified meta-representation, our framework effectively estimates multiple rapidly changing disturbances, as demonstrated by quadrotor flight experiments. See the project page for video, supplementary material and code: https://nonstructural-metalearn.github.io.

cs.RO

Equilibrium Portfolio Selection under Utility-Variance Analysis of Log Returns in Incomplete Markets

This paper investigates a time-inconsistent portfolio selection problem in the incomplete mar ket model, integrating expected utility maximization with risk control. The objective functional balances the expected utility and variance on log returns, giving rise to time inconsistency and motivating the search of a time-consistent equilibrium strategy. We characterize the equilibrium via a coupled quadratic backward stochastic differential equation (BSDE) system and establish the existence theory in two special cases: (i)the two Brownian motions driven the price dynamics and the factor process are independent with $ρ= 0$; (ii) the trading strategy is constrained to be bounded. For the general case with correlation coefficient $ρ\neq 0$, we introduce the notion of an approximate time-consistent equilibrium. Employing the solution structure from the equilibrium in the case $ρ= 0$, we can construct an approximate time-consistent equilibrium in the general case with an error of order $O(ρ^2)$. Numerical examples and financial insights are also presented based on deep learning algorithms.

q-fin.PM

Surface elasticity effect on Plateau-Rayleigh instability in soft solids

Soft solids exhibit instability and develop surface undulations due to surface effects, a phenomenon known as the elastic Plateau-Rayleigh (PR) instability, driven by the interplay of surface and bulk elasticity. Previous studies on the PR instability in solids mainly focused on the case of constant surface tension and ignored the effect of surface elasticity. It has been shown by experiments that the surface effects in solid-like materials depend both on the surface tension and surface elasticity, but little is known about the role of the latter in the elasto-capillary instabilities in soft solids. Here, we conduct an in-depth exploration of the effect of surface elasticity on the PR instability in an elastic cylinder by coupling theoretical and numerical methods. We derive an asymptotically consistent one-dimensional (1d) model to characterize the PR instability from three-dimensional (3d) nonlinear bulk-surface elasticity, and develop a new finite-element (FE) scheme for simulating 3d deformations of the bulk-surface system. The initiation and evolution of the PR instability are obtained analytically with the aid of the 1d model. The 1d results are further validated by the 3d FE simulations. By synthesizing the 1d analytic solutions and 3d numerical results, the effects of surface elasticity, surface compressibility, surface tension, axial force and geometrical size on the PR instability are thoroughly elucidated. Our results can be applied to calibrate surface parameters for solid-like materials and develop constitutive models for elastic surfaces.

cond-mat.soft

A one-dimensional model for axisymmetric deformations of an inflated hyperelastic tube of finite wall thickness

We derive a one-dimensional (1d) model for the analysis of bulging or necking in an inflated hyperelastic tube of {\it finite wall thickness} from the three-dimensional finite elasticity theory by applying the dimension reduction methodology proposed by Audoly and Hutchinson (J. Mech. Phys. Solids, 97, 2016). The 1d model makes it much easier to characterize fully nonlinear axisymmetric deformations of a thick-walled tube using simple numerical schemes such as the finite difference method. The new model recovers the diffuse interface model for analyzing bulging in a membrane tube and the 1d model for investigating necking in a stretched solid cylinder as two limiting cases. It is consistent with, but significantly refines, the exact linear and weakly nonlinear bifurcation analyses. Comparisons with finite element simulations show that for the bulging problem, the 1d model is capable of describing the entire bulging process accurately, from initiation, growth, to propagation. The 1d model provides a stepping stone from which similar 1d models can be derived and used to study other effects such as anisotropy and electric loading, and other phenomena such as rupture.

cond-mat.soft

Mean-Field Game of Relative Performance Portfolio for Two Populations with Poisson Common Noise

This paper studies the mean field game (MFG) and N-player game on relative performance portfolio management with two heterogeneous populations. In addition to the Brownian idiosyncratic and common noise, the first population invests in assets driven by idiosyncratic Poisson jump risk, while the second population invests in assets subject to Poisson common noise. We establish the characterization of the mean-field equilibrium (MFE) in MFG with two populations as well as the Nash equilibrium in the $N_1+N_2$-player game. Furthermore, we prove the convergence of the Nash equilibrium in the $N_1+N_2$-player game to the MFE as the number of players in two populations tends to infinity. We also discuss some impacts on MFE by the Poisson idiosyncratic risk and Poisson common noise in the context of relative performance, compensated by some numerical examples and financial implications.

math.OC

Mean Field Game with Reflected Jump Diffusion Dynamics: A Linear Programming Approach

This paper develops a linear programming approach for mean field games with reflected jump-diffusion dynamics. We first prove the equivalence between the mean field equilibria in the linear programming formulation and those in the weak relaxed control formulation under some measurability and growth conditions on model coefficients. Building upon the characterization of the occupation measure in the equivalence result, we further establish the existence of linear programming mean field equilibria under fairly general conditions on model coefficients. Finally, a numerical example is presented to illustrate the computation of a mean field equilibrium using the linear programming formulation.

math.OC

Unraveling $K(1690)$ as a pseudoscalar $ud\bar{d}\bar{s}$ tetraquark state

The recent observed $K (1690)$ has been identified as a supernumerary pseudoscalar resonance signal in the strange-meson spectrum predicted by quark model calculations. It is the best candidate of a strange crypto-exotic state. In this work, we systematically study the hadron masses of $ud\bar{d}\bar{s}$ tetraquark states with $J^P = 0^-$ in the method of QCD sum rules (QCDSR). For ten interpolating currents, we calculate the correlation functions up to dimension-8 nonperturbative condensates. To calculate the tri-gluon condensate, we comprehensively consider the contributions from different operators with and without covariant derivatives. The infrared (IR) safety can be guaranteed for the completely calculated tri-gluon condensate by properly addressing the IR divergences in Feynman diagrams. It is demonstrated that the tri-gluon condensate provides significant contributions to the sum-rule analyses in these light tetraquark systems. Our results support the interpretation of $K (1690)$ resonance to be a pseudoscalar $ud\bar{d}\bar{s}$ tetraquark state.

hep-ph

An extended Merton problem with relaxed benchmark tracking

This paper studies Merton's problem in an extended formulation by incorporating the benchmark tracking on the wealth process. We consider a tracking formulation where the fund manager aims to maximize the trade-off between the expected utility of consumption and the expected largest shortfall of the wealth with reference to the benchmark level. Equivalently, the problem can be interpreted as a mixed stochastic control problem if a fictitious capital injection singular control is allowed, subjecting to the dynamic constraint that the wealth process compensated by the costly capital injection outperforms the benchmark at all times. By considering an auxiliary state process, we formulate an equivalent stochastic control problem with state reflections at zero. For general utility functions and Ito's diffusion benchmark process, we develop a convex duality theorem, new to the literature, to the auxiliary stochastic control problem with state reflections in which the dual process also exhibits reflections from above. For CRRA utility and geometric Brownian motion benchmark process, we further derive the optimal portfolio and consumption in feedback form using the new duality theorem, allowing us to discuss some interesting financial implications induced by the additional risk-taking from the capital injection and the goal of tracking.

math.OC

On time-consistent equilibrium stopping under aggregation of diverse discount rates

This paper studies a central planner's decision making on behalf of a group of members with diverse discount rates. In the context of optimal stopping, we work with an aggregation preference to incorporate all discount rates via an attitude function that reflects the aggregation rule chosen by the central planner. The problem formulation is also applicable to single agent's stopping problem with uncertain discount rate, where our aggregation preference coincides with the conventional smooth ambiguity preference. The resulting optimal stopping problem is time inconsistent, for which we develop an iterative approach using consistent planning and characterize all time-consistent mild equilibria as fixed points of an operator in the setting of one-dimensional diffusion processes. We provide some sufficient conditions on the underlying models and the attitude function such that the smallest mild equilibrium attains the optimal equilibrium. In addition, we show that the optimal equilibrium is a weak equilibrium. When the sufficient condition of the attitude function is violated, we illustrate by various examples that the characterization of the optimal equilibrium may differ significantly from some existing results for a single agent, which now sensitively depends on the attitude function and the diversity distribution of discount rates within the group.

q-fin.MF

Major-Minor Mean Field Game of Stopping: An Entropy Regularization Approach

This paper studies a discrete-time major-minor mean field game of stopping where the major player can choose either an optimal control or stopping time. We look for the relaxed equilibrium as a randomized stopping policy, which is formulated as a fixed point of a set-valued mapping, whose existence is challenging by direct arguments. To overcome the difficulties caused by the presence of a major player, we propose to study an auxiliary problem by considering entropy regularization in the major player's problem while formulating the minor players' optimal stopping problems as linear programming over occupation measures. We first show the existence of regularized equilibria as fixed points of some simplified set-valued operator using the Kakutani-Fan-Glicksberg fixed-point theorem. Next, we prove that the regularized equilibrium converges as the regularization parameter $λ$ tends to 0, and the limit corresponds to a fixed point of the original operator, thereby confirming the existence of a relaxed equilibrium in the original mean field game problem. We also extend this entropy regularization method to the mean-field game problem where the minor players choose optimal controls.

math.OC

Optimal consumption under relaxed benchmark tracking and consumption drawdown constraint

This paper studies an optimal consumption problem with both relaxed benchmark tracking and consumption drawdown constraint, leading to a stochastic control problem with dynamic state-control constraints. In our relaxed tracking formulation, it is assumed that the fund manager can strategically inject capital to the fund account such that the total capital process always outperforms the benchmark process, which is described by a geometric Brownian motion. We first transform the original regular-singular control problem with state-control constraints into an equivalent regular control problem with a reflected state process and consumption drawdown constraint. By utilizing the dual transform and the optimal consumption behavior, we then turn to study the linear dual PDE with both Neumann boundary condition and free boundary condition in a piecewise manner across different regions. Using the smoothfit principle and the super-contact condition, we derive the closed-form solution of the dual PDE, and obtain the optimal investment and consumption in feedback form. We then prove the verification theorem on optimality by some novel arguments with the aid of an auxiliary reflected dual process and some technical estimations. Some numerical examples and financial insights are also presented.

math.OC