SearcharxivSearch

arXiv subjects

Ziyue Chen

Publications and source records attributed to Ziyue Chen.

5 recordsLinked to original sources

Global linear convergence of entropy-regularized softmax policy gradient beyond tabular MDPs

We study the global convergence of policy gradient for infinite-horizon entropy-regularized Markov decision processes (MDPs) with continuous state and action spaces. We consider log-linear softmax policies with linear function approximation, which extend the tabular softmax parameterization while retaining a tractable policy class. Under $Q^\pi_\tau$-realizability for the regularized state-action value function, we first establish a non-uniform Polyak--{\L}ojasiewicz (P\L) inequality. The non-uniformity arises through degeneracy of constants associated with the policy geometry, namely the Fisher information matrix or an uncentered feature covariance matrix. We then identify two feature regimes under which this non-uniform constant can be bounded along the gradient flow. For full-affine-span features, we prove radial unboundedness of the KL regularizer and show that the smallest eigenvalue of the Fisher information matrix remains bounded below by an initialization-dependent positive constant. For simplex-valued features, we prove an analogous radial unboundedness result in the subspace orthogonal to the all-ones vector and obtain a uniform lower bound for the smallest eigenvalue of the uncentered covariance matrix. These results imply global linear convergence of the regularized objective along the gradient flow, i.e. suboptimality decaying as $\mathcal{O}(e^{-Ct})$ for some $C>0$. Our analysis extends the global convergence theory of entropy-regularized softmax policy gradient beyond the tabular setting of Agarwal et al. (2020); Bhandari and Russo (2024); Mei et al. (2020).

cs.LG

Curriculum Negative Mining For Temporal Networks

Temporal networks are effective in capturing the evolving interactions of networks over time, such as social networks and e-commerce networks. In recent years, researchers have primarily concentrated on developing specific model architectures for Temporal Graph Neural Networks (TGNNs) in order to improve the representation quality of temporal nodes and edges. However, limited attention has been given to the quality of negative samples during the training of TGNNs. When compared with static networks, temporal networks present two specific challenges for negative sampling: positive sparsity and positive shift. Positive sparsity refers to the presence of a single positive sample amidst numerous negative samples at each timestamp, while positive shift relates to the variations in positive samples across different timestamps. To robustly address these challenges in training TGNNs, we introduce Curriculum Negative Mining (CurNM), a model-aware curriculum learning framework that adaptively adjusts the difficulty of negative samples. Within this framework, we first establish a dynamically updated negative pool that balances random, historical, and hard negatives to address the challenges posed by positive sparsity. Secondly, we implement a temporal-aware negative selection module that focuses on learning from the disentangled factors of recently active edges, thus accurately capturing shifting preferences. Finally, the selected negatives are combined with annealing random negatives to support stable training. Extensive experiments on 12 datasets and 3 TGNNs demonstrate that our method outperforms baseline methods by a significant margin. Additionally, thorough ablation studies and parameter sensitivity experiments verify the usefulness and robustness of our approach.

cs.LG

Backward Stochastic Control System with Entropy Regularization

The entropy regularization is inspired by information entropy from machine learning and the ideas of exploration and exploitation in reinforcement learning, which appears in the control problem to design an approximating algorithm for the optimal control. This paper is concerned with the optimal exploratory control for backward stochastic system, generated by the backward stochastic differential equation and with the entropy regularization in its cost functional. We give the theoretical depict of the optimal relaxed control so as to lay the foundation for the application of such a backward stochastic control system to mathematical finance and algorithm implementation. For this, we first establish the stochastic maximum principle by convex variation method. Then we prove sufficient condition for the optimal control and demonstrate the implicit form of optimal control. Finally, the existence and uniqueness of the optimal control for backward linear-quadratic control problem with entropy regularization is proved by decoupling techniques.

math.OC

Distribution and Determinants of Correlation between PM2.5 and O3 in China Mainland: Dynamitic simil-Hu Lines

In recent years, China has made great efforts to control air pollution. During the governance process, it is found that fine particulate matter (PM2.5) and ozone (O3) change in the same trend among some areas and the opposite in others, which brings some difficulties to take measures in a planned way. Therefore, this study adopted multi-year and large-scale air quality data to explore the distribution of correlation between PM2.5 and O3, and proposed a concept called dynamic similar hu lines to replace the single fixed division in the previous research. Furthermore, this study discussed the causes of distribution patterns quantitatively with geographical detector and random forest. The causes included natural factors and anthropogenic factors. And these factors could be divided into three parts according to the characteristics of spatial distribution: broadly changing with longitude, changing with latitude, and having local characteristics. Overall, regions with relatively more densely population, higher GDP, lower altitude, higher humidity, higher atmospheric pressure, higher surface temperature, less sunshine hours and more accumulated precipitation often corresponds to positive correlation coefficient between PM2.5 and O3, no matter in which season. The parts with opposite conditions that mentioned above are essentially negative correlation coefficient. And what's more, humidity, global surface temperature, air temperature and accumulated precipitation are four decisive factors to form the distribution of correlation between PM2.5 and O3. In general, collaborative governance of atmospheric pollutants should consider particular time and space background and also be based on the local actual socio-economic situations, geography and geomorphology, climate and meteorology and other comprehensive factors.

physics.ao-ph

On variance estimation for generalizing from a trial to a target population

Randomized controlled trials (RCTs) provide strong internal validity compared with observational studies. However, selection bias threatens the external validity of randomized trials. Thus, RCT results may not apply to either broad public policy populations or narrow populations, such as specific insurance pools. Some researchers use propensity scores (PSs) to generalize results from an RCT to a target population. In this scenario, a PS is defined as the probability of participating in the trial conditioning on observed covariates. We study a model-free inverse probability weighted estimator (IPWE) of the average treatment effect in a target population with data from a randomized trial. We present variance estimators and compare the performance of our method with that of model-based approaches. We examine the robustness of the model-free estimators to heterogeneous treatment effects.

stat.ME