Searcharxiv⌕ Search

arXiv · 1405.5623

Stochastic variational inference for large-scale discrete choice models using adaptive batch sizes

Abstract

Discrete choice models describe the choices made by decision makers among alternatives and play an important role in transportation planning, marketing research and other applications. The mixed multinomial logit (MMNL) model is a popular discrete choice model that captures heterogeneity in the preferences of decision makers through random coefficients. While Markov chain Monte Carlo methods provide the Bayesian analogue to classical procedures for estimating MMNL models, computations can be prohibitively expensive for large datasets. Approximate inference can be obtained using variational methods at a lower computational cost with competitive accuracy. In this paper, we develop variational methods for estimating MMNL models that allow random coefficients to be correlated in the posterior and can be extended easily to large-scale datasets. We explore three alternatives: (1) Laplace variational inference, (2) nonconjugate variational message passing and (3) stochastic linear regression. Their performances are compared using real and simulated data. To accelerate convergence for large datasets, we develop stochastic variational inference for MMNL models using each of the above alternatives. Stochastic variational inference allows data to be processed in minibatches by optimizing global variational parameters using stochastic gradient approximation. A novel strategy for increasing minibatch sizes adaptively within stochastic variational inference is proposed.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Linda S. L. Tan. 2015-10-08. Stochastic variational inference for large-scale discrete choice models using adaptive batch sizes. https://doi.org/10.1007/s11222-015-9618-x

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Who Blocks Whom? Probabilistic Pass-Blocking Assignments for Evaluating Blockers and Pass Rushers in American Football

Historically, statistical analysis of offensive lineman has been hindered by the lack of easily measurable quantities. More recently, with the introduction of player tracking data new methodological advances are now possible. Using high-dimensional spatio-temporal data, we adapt the defensive-matchup hidden Markov model of \cite{franks2015characterizing} from basketball to football pass protection, producing frame-by-frame probabilistic assignments of each pass blocker to the rushers. We show how this probabilistic assignment is a usable modeling artifact that augments existing player-evaluation frameworks. We directly quantify the attention a rusher commands, upgrade adjusted plus-minus \citep{Macdonald+2012} from all-or-nothing stints to partial, continuous blocking credit in continuous time, yield block-shedding survival metrics, and measure the space a rusher generates for his teammates. Fit to the first eight weeks of the 2021 NFL season, the resulting metrics recover widely-recognized elite rushers and pass protectors and align with independent charting.

stat.AP↗

Geographic Disparities in Hospice Quality and Family Caregiver Experience: The Roles of Ownership, Social Vulnerability, and Workforce Capacity

Hospice quality should be interpreted in relation to both provider organization and the local conditions under which care is delivered. This study develops a provider-county performance assessment framework by linking national Centers for Medicare & Medicaid Services (CMS) hospice data and Consumer Assessment of Healthcare Providers and Systems (CAHPS) Hospice Survey outcomes with county measures of rurality, social vulnerability, health burden, and workforce and health-resource context. The adjusted analysis includes 2,928 providers in 1,078 counties and combines geographic mapping, blockwise regression, six secondary CAHPS outcomes, and eight sensitivity analyses. Adding county context increased adjusted R-squared from 0.077 to 0.202. After full adjustment, for-profit hospices had overall caregiver ratings 3.832 percentage points lower than nonprofit hospices. This negative association appeared across all six secondary CAHPS domains and remained significant in every sensitivity specification. Higher county social vulnerability was also associated with poorer caregiver experience, although its magnitude depended partly on the specification of community health burden. These findings show that county context materially improves hospice performance assessment but does not eliminate the ownership difference. The framework supports context-aware monitoring, peer comparison, and targeted quality improvement.

stat.AP↗

Feeling Left Behind? Territorial Disadvantage, Well-Being and Cohesion across European Regions

Left-behind places are usually identified using economic, demographic and accessibility indicators, but these may not align with well-being and social cohesion. We combine ten years of territorial indicators with subjective measures derived from georeferenced social-media data for over 1,200 NUTS-3 regions in 28 European countries (2013-2023). Regional profiles and within-between panel models reveal uneven relationships: disadvantaged regions can display contrasting levels of well-being and cohesion, national context alters cross-country associations, and between-region differences diverge from within-region change. The findings show that aggregate measures conceal distinct forms and trajectories of territorial disadvantage.

stat.AP↗