arXiv · 2609.24605
Scalable Dynamic Pricing of Substitutable Products through Structure-Guided Policy Learning
Abstract
Problem definition: We study dynamic pricing of substitutable products with finite, product-specific inventories. Customer substitution couples pricing decisions across products, while the inventory state makes exact dynamic programming intractable at realistic scale. Methodology / results: We develop two MNL-guided policy-learning approaches that replace the dynamic program with a statistical mapping from inventory states to pricing decisions. The first learns prices directly, while the second learns inventory opportunity costs and converts them into prices using the optimal MNL pricing rule. Both policies are trained using decision-focused learning. We also propose an efficient method to generate anticipative customer-choice targets to train our policies. Managerial implications: We conduct an extensive numerical evaluation of our approaches. On small instances for which the optimal dynamic program can be computed, the learned policies achieve average optimality gaps below 0.4%. On larger airline-motivated instances, they consistently improve on the tested revenue-management benchmarks. The comparison between the two architectures also highlights the role of model structure: using the MNL pricing characterization is particularly effective when demand is well described by MNL, while directly learning prices provides greater flexibility under heterogeneous mixed-MNL demand.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yue Su, Antoine Désir, Axel Parmentier. 2026-09-21. Scalable Dynamic Pricing of Substitutable Products through Structure-Guided Policy Learning. https://arxiv.org/abs/2609.24605
Cite the original work for its findings. Save a collection to share your selection of sources.