Searcharxiv⌕ Search

arXiv · 2610.08409

Bayesian Machine Learning Methods For Large Scale Demand Estimation

Abstract

This work studies how Bayesian machine learning methods can be used for large-scale demand estimation with many product categories. I compare two model classes, a latent factorization model and a mixed logit model and two Bayesian estimation approaches, Markov Chain Monte Carlo (MCMC) and Variational Inference (VI). The analysis combines a simulation study with an application to supermarket scanner data. The results show that the latent factorization model benefits from information across categories and improves its predictive performance as the dimensionality of the choice environment increases, whereas the mixed logit model does not exhibit the same pattern. MCMC delivers the highest predictive accuracy but is computationally intensive. VI achieves slightly lower predictive performance while substantially reducing runtime. In the empirical application, VI also outperforms the mixed logit benchmark. These findings highlight a trade-off between accuracy and computational feasibility in multi-category demand estimation.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Anna B. Schmidt. 2026-10-06. Bayesian Machine Learning Methods For Large Scale Demand Estimation. https://arxiv.org/abs/2610.08409

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Identifying treatment effects on categorical outcomes in IV models

This paper studies population treatment effects when the outcome is unordered categorical, the treatment is binary, and the instrument is binary. I introduce an assumption called association similarity. For each instrument value, association similarity requires the odds-ratio association between potential treatment and potential outcome to be the same whether the outcome is evaluated under treatment or under no treatment. Under association similarity and a testable categorywise dominance condition, the full potential-outcome distributions are point identified. The identification is constructive and allows estimation using sample frequencies. I also consider weaker restrictions on departures from association similarity that lead to sharp partial identification. I illustrate the method with an application about the effect of insurance on health outcomes.

econ.EM↗

Robust Inference for Dyadic Data with Spatially Dependent Nodes

We develop inference for complete dyadic samples with spatial dependence governed by geographic distances between nodes, covering settings beyond the scope of existing dyadic inference theory. We propose spatial variance and corrected block jackknife estimators consistent in nondegenerate and degenerate Gaussian cases. Under degeneracy, spatial subsampling consistently estimates possibly non-Gaussian limits. Combining the jackknife with subsampling yields the max jackknifes-ubsampling (MJS) interval, which provides pointwise asymptotically exact coverage in both Gaussian cases and conservative coverage in the specified non-Gaussian case. Fixed-effect extensions show that estimating node effects can change the fast limiting distribution. Simulations and an empirical illustration are provided.

econ.EM↗

Vine Copula VAR:From Recursive Margins to Joint Forecast Inference

Joint-event forecasts often combine a dependence estimate based on past forecast errors with newly estimated marginal distributions. When each historical error retains the marginal fit available at its issue date, inference must account for an overlapping sequence of estimation errors. We derive their joint influence with the terminal forecast estimates in a stable Vine Copula VAR with normal innovation margins and a fixed, correctly specified Gaussian or positive Clayton vine. An intercept identity and the stable VAR filter reduce the historical correction to harmonically weighted innovation moments, while terminal slope uncertainty remains. The resulting covariance estimator gives asymptotically valid repeated-sample intervals for fixed one-sided event probabilities at the realized forecast state. In the Gaussian submodel, retaining issued transforms adds a positive semidefinite covariance term relative to refitting margins on the same observations. Monte Carlo simulations show that terminal-margin uncertainty is quantitatively more important than this additional term and that logit intervals improve lower-tail coverage in the designs studied. A real-time forecasting application to U.S. macroeconomic releases shows how marginal estimation contributes to uncertainty in predicted probabilities of joint contractions and identifies limitations of the stationary marginal model.

econ.EM↗