SearcharxivSearch

arXiv subjects

Luca Serfilippi

Publications and source records attributed to Luca Serfilippi.

2 recordsLinked to original sources

Complexity-Regularized Proximal Policy Optimization

Policy gradient methods usually rely on entropy regularization to prevent premature convergence. However, maximizing entropy indiscriminately pushes the policy towards a uniform distribution, often overriding the reward signal if not optimally tuned. We propose replacing the standard entropy term with a self-regulating complexity term, defined as the product of Shannon entropy and disequilibrium, where the latter quantifies the distance from the uniform distribution. Unlike pure entropy, which favors maximal disorder, this complexity measure is zero for both fully deterministic and perfectly uniform distributions, i.e., it is strictly positive for systems that exhibit a meaningful interplay between order and randomness. These properties ensure the policy maintains beneficial stochasticity while reducing regularization pressure when the policy is highly uncertain, allowing learning to focus on reward optimization. We introduce Complexity-Regularized Proximal Policy Optimization (CR-PPO), a modification of PPO that leverages this dynamic. We empirically demonstrate that CR-PPO is significantly more robust to hyperparameter selection than entropy-regularized PPO, achieving consistent performance across orders of magnitude of regularization coefficients and remaining harmless when regularization is unnecessary, thereby reducing the need for expensive hyperparameter tuning.

cs.LG

Quantum Chemistry Driven Molecular Inverse Design with Data-free Reinforcement Learning

The inverse design of molecules has challenged chemists for decades. In the past years, machine learning and artificial intelligence have emerged as new tools to generate molecules tailoring desired properties, but with the limit of relying on models that are pretrained on large datasets. Here, we present a data-free generative model based on reinforcement learning and quantum mechanics calculations. To improve the generation, our software is based on a five-model reinforcement learning algorithm designed to mimic the syntactic rules of an original ASCII encoding based on the SMILES one, and here reported. The reinforcement learning generator is rewarded by on-the-fly quantum mechanics calculations within a computational routine addressing conformational sampling. We demonstrate that our software successfully generates new molecules with desired properties finding optimal solutions for problems with known solutions and (sub)optimal molecules for unexplored chemical (sub)spaces, jointly showing significant speed-up to a reference baseline.

physics.chem-ph