Searcharxiv⌕ Search

arXiv subjects

Paolo Brandimarte

Publications and source records attributed to Paolo Brandimarte.

3 recordsLinked to original sources

A Reinforcement Learning Method for Environments with Stochastic Variables: Post-Decision Proximal Policy Optimization with Dual Critic Networks

This paper presents Post-Decision Proximal Policy Optimization (PDPPO), a novel variation of the leading deep reinforcement learning method, Proximal Policy Optimization (PPO). The PDPPO state transition process is divided into two steps: a deterministic step resulting in the post-decision state and a stochastic step leading to the next state. Our approach incorporates post-decision states and dual critics to reduce the problem's dimensionality and enhance the accuracy of value function estimation. Lot-sizing is a mixed integer programming problem for which we exemplify such dynamics. The objective of lot-sizing is to optimize production, delivery fulfillment, and inventory levels in uncertain demand and cost parameters. This paper evaluates the performance of PDPPO across various environments and configurations. Notably, PDPPO with a dual critic architecture achieves nearly double the maximum reward of vanilla PPO in specific scenarios, requiring fewer episode iterations and demonstrating faster and more consistent learning across different initializations. On average, PDPPO outperforms PPO in environments with a stochastic component in the state transition. These results support the benefits of using a post-decision state. Integrating this post-decision state in the value function approximation leads to more informed and efficient learning in high-dimensional and stochastic environments.

cs.LG↗

Multivariate Lévy models: calibration and pricing

The goal of this paper is to investigate how the marginal and dependence structures of a variety of multivariate Lévy models affect calibration and pricing. To this aim, we study the approaches of Luciano and Semeraro (2010) and Ballotta and Bonfiglioli (2016) to construct multivariate processes. We explore several calibration methods that can be used to fine-tune the models, and that deal with the observed trade-off between marginal and correlation fit. We carry out a thorough empirical analysis to evaluate the ability of the models to fit market data, price exotic derivatives, and embed a rich dependence structure. By merging theoretical aspects with the results of the empirical test, we provide tools to make suitable decisions about the models and calibration techniques to employ in a real context.

q-fin.PR↗

Rolling horizon policies for multi-stage stochastic assemble-to-order problems

Assemble-to-order approaches deal with randomness in demand for end items by producing components under uncertainty, but assembling them only after demand is observed. Such planning problems can be tackled by stochastic programming, but true multistage models are computationally challenging and only a few studies apply them to production planning. Solutions based on two-stage models are often short-sighted and unable to effectively deal with non-stationary demand. A further complication may be the scarcity of available data, especially in the case of correlated and seasonal demand. In this paper, we compare different scenario tree structures. In particular, we enrich a two-stage formulation by introducing a piecewise linear approximation of the value of the terminal inventory, to mitigate the two-stage myopic behavior. We compare the out-of-sample performance of the resulting models by rolling horizon simulations, within a data-driven setting, characterized by seasonality, bimodality, and correlations in the distribution of end item demand. Computational experiments suggest the potential benefit of adding a terminal value function and illustrate interesting patterns arising from demand correlations and the level of available capacity. The proposed approach can provide support to typical MRP/ERP systems, when a two-level approach is pursued, based on master production and final assembly scheduling.

math.OC↗