SearcharxivSearch

arXiv subjects

Haibin Xie

Publications and source records attributed to Haibin Xie.

5 recordsLinked to original sources

Koopman Dreamer: Spectrally Constrained Latent Dynamics for Stable World-Model Imagination

Latent world models improve sample efficiency in continuous control by optimizing policies over imagined latent trajectories, but common neural transitions offer limited direct control over modal persistence and error accumulation in long rollouts. We propose Koopman Dreamer, a Dreamer-style world model with a spectrally constrained deterministic latent dynamics core. Its Koopman-inspired backbone uses two-dimensional rotation--scaling blocks with bounded radii to represent damping, rotation, and near-periodic modes. Linear and low-rank bilinear action terms capture global and state-dependent control effects, while stochastic-state modulation supplies local correction information. To reduce the mismatch between posterior-conditioned training and prior-only imagination, the model combines posterior-conditioned EMA teacher targets with one-step consistency, multi-step rollout, and open-loop observation-prediction objectives. We further derive a multi-step rollout-error bound that separates amplification by the spectral backbone and bilinear interaction from the additive effects of stochastic-state mismatch and modeling residuals, clarifying the trade-off between error attenuation and long-term information retention. Experimental results on proprioceptive continuous-control tasks from the DeepMind Control Suite and UAV-LiDAR autonomous navigation demonstrate that Koopman Dreamer improves the stability of long-horizon latent rollouts and achieves stronger closed-loop control performance on tasks that rely on high-quality multi-step imagination.

cs.LG

Long-term User Engagement Optimization through Model-agnostic Downstream Rewards Learning

As recommender systems mature in the past few years, their optimization objectives have evolved from a primary focusing on short-term behavioral signals to a broader emphasis on long-term user engagement and retention. However, directly optimizing retention is difficult because return signals are sparse, delayed, and only partially attributable to earlier recommendations. Prior work has addressed this challenge with sequential modeling and reinforcement learning, but these approaches typically require task specific reward engineering, substantial computational overhead, and surface specific implementations that are difficult to generalize. In this paper, we present a unified, model-agnostic downstream reward framework for optimizing long-term user value in large-scale recommendation systems. First, we formulate the downstream reward learning problem and develop an offline screening framework to identify session level behaviors that are both observable early and predictive of future retention. We then propose several model-agnostic downstream rewards signals derived from observed user action patterns across multiple sources. We further discuss the engineering effort to productionize the proposed rewards derivations and challenges we faced when adding them to our ranking models. Online A/B experiments demonstrate consistent improvements in engagement and retention-related metrics, and the framework has been deployed across multiple Pinterest surfaces, including Homefeed, Related Pins, Search, and Notifications.

cs.LG

Model-Based Safe Reinforcement Learning with Time-Varying State and Control Constraints: An Application to Intelligent Vehicles

Recently, safe reinforcement learning (RL) with the actor-critic structure for continuous control tasks has received increasing attention. It is still challenging to learn a near-optimal control policy with safety and convergence guarantees. Also, few works have addressed the safe RL algorithm design under time-varying safety constraints. This paper proposes a safe RL algorithm for optimal control of nonlinear systems with time-varying state and control constraints. In the proposed approach, we construct a novel barrier force-based control policy structure to guarantee control safety. A multi-step policy evaluation mechanism is proposed to predict the policy's safety risk under time-varying safety constraints and guide the policy to update safely. Theoretical results on stability and robustness are proven. Also, the convergence of the actor-critic implementation is analyzed. The performance of the proposed algorithm outperforms several state-of-the-art RL algorithms in the simulated Safety Gym environment. Furthermore, the approach is applied to the integrated path following and collision avoidance problem for two real-world intelligent vehicles. A differential-drive vehicle and an Ackermann-drive one are used to verify offline deployment and online learning performance, respectively. Our approach shows an impressive sim-to-real transfer capability and a satisfactory online control performance in the experiment.

cs.LG

Freshness-Optimal Caching for Information Updating Systems with Limited Cache Storage Capacity

In this paper, we investigate a cache updating system with a server containing $N$ files, $K$ relays and $M$ users. The server keeps the freshest versions of the files which are updated with fixed rates. Each relay can download the fresh files from the server in a certain period of time. Each user can get the fresh files from any relay as long as the relay has stored the fresh versions of the requested files. Due to the limited storage capacity and updating capacity of each relay, different cache designs will lead to different average freshness of all updating files at users. In order to keep the average freshness as large as possible in the cache updating system, we formulate an average freshness-optimal cache updating problem (AFOCUP) to obtain an optimal cache scheme. However, because of the nonlinearity of the AFOCUP, it is difficult to seek out the optimal cache scheme. As a result, an linear approximate model is suggested by distributing the total update rates completely in accordance with the number of files in the relay in advance. Then we utilize the greedy algorithm to search the optimal cache scheme that is satisfied with the limited storage capacity of each relay. Finally, some numerical examples are provided to illustrate the performance of the approximate solution.

cs.IT

Timing the market: the economic value of price extremes

By decomposing asset returns into potential maximum gain (PMG) and potential maximum loss (PML) with price extremes, this study empirically investigated the relationships between PMG and PML. We found significant asymmetry between PMG and PML. PML significantly contributed to forecasting PMG but not vice versa. We further explored the power of this asymmetry for predicting asset returns and found it could significantly improve asset return predictability in both in-sample and out-of-sample forecasting. Investors who incorporate this asymmetry into their investment decisions can get substantial utility gains. This asymmetry remains significant even when controlling for macroeconomic variables, technical indicators, market sentiment, and skewness. Moreover, this asymmetry was found to be quite general across different countries.

q-fin.CP