SearcharxivSearch

arXiv subjects

Harrison Katz

Publications and source records attributed to Harrison Katz.

16 recordsLinked to original sources

When Should Forecasting Models Be Re-Specified? A Cost-Sensitive Trigger for Adaptive Model-Form Updating

Routine refresh bundles two operations that need not travel together: estimating parameters and selecting the model form. The second is often unnecessary. Under a reduced-update policy the maintenance question is a stopping problem: when has enough evidence accumulated against the deployed form to justify re-specifying it? We define specification debt as that evidence, derive a cost-sensitive rule that re-specifies when expected avoidable loss exceeds re-specification cost, and recover fixed update frequencies as the constant-accumulation case. The rule then raises its own question: how much machinery the monitor needs. We take this to all 47,982 monthly M4 series over an exponential smoothing grid and field three rules ordered by machinery: a fixed cadence with no monitoring, a one-parameter tracking signal after Trigg and Brown on the deployed model's one-step errors, and an evidence-gated trigger that fits a twelve-candidate grid for a validation score gap. Model form matters little on average here, so a cheap cadence matches full updating. The ordering within that band is informative. The tracking signal matches full updating at the three-step benchmark, but that edge dies under rate matching: against a cadence making the same nine searches, the gap is null at every horizon. What survives is the detector, once separated from the selector. The score-gap signal loses to its cap-matched cadence when a fire deploys the validation winner, but beats it at all five horizons when a fire re-selects by AICc, never worse at any lead; the validation-winner selector trades the first lead for the largest long-horizon gains, beating the cadence at horizon 18 by 0.4 percent (t of -9.0). Whether a gate repays its compute reduces to a break-even weight the evaluation reports. Everything rests on realized out-of-sample loss; the in-sample analogue does not predict degradation.

stat.AP

Cost-sensitive retraining via posterior learning debt

Deployed prediction systems are often retrained on fixed calendars, even when model staleness and retraining burden vary over time. This short communication formulates retraining for Bayesian prediction systems as a cost-sensitive predictive-regret decision. The central monitoring state is posterior learning debt, defined as the Kullback--Leibler divergence from a reference shadow posterior to the deployed frozen posterior. In the decision layer, a retraining cost is compared with the expected one-period predictive regret of waiting. A continuous-severity version retrains when calibrated expected regret exceeds the retraining cost, while the familiar two-state excess-loss rule is a special case. The empirical study is an exact-state proof-of-concept in a synthetic conjugate simulation with warm-started deployed and shadow normal-inverse-gamma posteriors, separate update, monitoring, and evaluation batches, lagged deployment actions, expanded baseline grids, and score-unit sensitivity. Under the primary 75th-percentile score-unit scaling, an age-adjusted debt-threshold policy improves on tuned calendar retraining in all 72 non-stable scenario cells and on tuned CUSUM in 58 of 72 cells, with mean relative objectives 0.677 and 0.975, respectively. Debt-utility and hybrid-utility policies also improve strongly over tuned calendar retraining, but they do not dominate tuned CUSUM. Median and mean score-unit sensitivities show the same main calendar result, while the CUSUM comparison remains policy-dependent. The contribution is a transparent decision layer for deployed Bayesian prediction systems, not a universal replacement for drift detection.

stat.AP

Retraining as Approximate Bayesian Inference

Model retraining is usually treated as an ongoing maintenance task. But as Harrison Katz now argues, retraining can be better understood as approximate Bayesian inference under computational constraints. The gap between a continuously updated belief state and your frozen deployed model is "learning debt," and the retraining decision is a cost minimization problem with a threshold that falls out of your loss function. In this article Katz provides a decision-theoretic framework for retraining policies. The result is evidence-based triggers that replace calendar schedules and make governance auditable. For readers less familiar with the Bayesian and decision-theoretic language, key terms are defined in a glossary at the end of the article.

cs.AI

Coupled Supply and Demand Forecasting in Platform Accommodation Markets

Tourism demand forecasting is methodologically mature, but it typically treats accommodation supply as fixed or exogenous. In platform-mediated short-term rentals, supply is elastic, decision-driven, and co-evolves with demand through pricing, information design, and interventions. I reframe the core issue as endogenous stock-out censoring: realized booked nights satisfy B_{k,t} <= min(D_{k,t}, S_{k,t}), so booking models that ignore supply learn a regime-specific ceiling and become fragile under policy changes and supply shocks. This narrated review synthesizes work from tourism forecasting, revenue management, two-sided market economics, and Bayesian time-series methods; develops a three-part coupling framework (behavioral, informational, intervention); and illustrates the identification failure with a toy simulation. I conclude with a focused research agenda for jointly forecasting supply, demand, and their compositions.

stat.AP

Forecasting the Evolving Composition of Inbound Tourism Demand: A Bayesian Compositional Time Series Approach Using Platform Booking Data

Understanding how the composition of guest origin markets evolves over time is critical for destination marketing organizations, hospitality businesses, and tourism planners. We develop and apply Bayesian Dirichlet autoregressive moving average (BDARMA) models to forecast the compositional dynamics of guest origin market shares using proprietary Airbnb booking data spanning 2017--2025 across four major destination regions. Our analysis reveals substantial pandemic-induced structural breaks in origin composition, with heterogeneous recovery patterns across markets. In our analysis, the BDARMA framework achieves the lowest forecast error for EMEA and competitive performance across destination regions, outperforming standard benchmarks including na\"ive forecasts, exponential smoothing, and SARIMA on log-ratio transformed data in compositionally complex markets. For EMEA destinations, BDARMA achieves 27% lower forecast error than na\"ive methods ($p < 0.001$), with the greatest gains where multiple origin markets compete in the 5-25% share range. By modeling compositions directly on the simplex with a Dirichlet likelihood and incorporating seasonal variation in both mean and precision parameters, our approach produces coherent forecasts that respect the unit-sum constraint while capturing complex temporal dependencies. The methodology provides destination stakeholders with probabilistic forecasts of source market shares, enabling more informed strategic planning for marketing resource allocation, infrastructure investment, and crisis response.

stat.AP

Directional-Shift Dirichlet ARMA Models for Compositional Time Series with Structural Break Intervention

Compositional time series frequently exhibit structural breaks due to external shocks, policy changes, or market disruptions. Standard methods either ignore such breaks or handle them through fixed effects that cannot extrapolate beyond the sample, or step-function dummies that impose instantaneous adjustment. We develop a Bayesian Dirichlet ARMA model augmented with a directional-shift intervention mechanism that captures structural breaks through three interpretable parameters: a direction vector specifying which components gain or lose share, an amplitude controlling redistribution magnitude, and a logistic gate governing transition timing and speed. The model preserves compositional constraints by construction, maintains DARMA dynamics for short-run dependence, and produces coherent probabilistic forecasts through and after structural breaks. The intervention trajectory corresponds to geodesic motion on the simplex and is invariant to the choice of ILR basis. A simulation study with 400 fits across 8 scenarios shows near-zero amplitude bias and nominal 80\% credible interval coverage when the shift direction is correctly identified (77.5\% of cases); supplementary studies confirm robustness across extreme transition speeds and non-monotone DGPs. Two empirical applications to COVID-era Airbnb data characterize performance relative to simpler alternatives. Where the break is monotone and ongoing, the intervention model achieves near-nominal calibration (79.6\%) while the fixed effect substantially under-covers (66.1\%). Where post-break dynamics are non-monotone, both models are acceptably calibrated and the fixed effect outperforms on point accuracy. The intervention model's advantages are thus specific to settings with roughly monotone structural transitions.

stat.ME

Impact by design: translating Lead times in flux into an R handbook with code

This commentary translates the central ideas in Lead times in flux into a practice ready handbook in R. The original article measures change in the full distribution of booking lead times with a normalized L1 distance and tracks that divergence across months relative to year over year and to a fixed 2018 reference. It also provides a bound that links divergence and remaining horizon to the relative error of pickup forecasts. We implement these ideas end to end in R, using a minimal data schema and providing runnable scripts, simulated examples, and a prespecified evaluation plan. All results use synthetic data so the exposition is fully reproducible without reference to proprietary sources.

q-fin.ST

Centered-Innovation MA for Bayesian Dirichlet ARMA: Theoretical Equivalence and an Application to Bank-Asset Shares

We study a minimal change to an observation-driven Bayesian Dirichlet ARMA (B--DARMA) for compositional time series: replace the raw additive log-ratio (ALR) residual in the moving-average block with a centered innovation that subtracts the Dirichlet conditional ALR mean, available in closed form via digamma identities. We prove a recursion-level first-order equivalence (in $1/\phi$) between the centered specification and a digamma-link DARMA at fixed parameters, under explicit interior and lag-stability conditions. The result clarifies why the two specifications should be predictively indistinguishable in the high-precision regime but does not by itself govern the geometry of the Bayesian posteriors that re-estimation produces. On weekly Federal Reserve H.8 bank-asset shares (October~2015 through October~2025, $T=522$ weeks), predictive performance is statistically indistinguishable across $104$ rolling weekly origins on every accuracy metric examined, while Hamiltonian Monte Carlo divergent transitions are approximately an order of magnitude more frequent under the raw specification, driven by isolated rolling fits at which the raw posterior exhibits localized pathologies. A four-reference sensitivity analysis confirms that predictive equivalence is reference-invariant and that the geometric advantage of centering is preserved across references but varies with the prevalence of pathological raw fits, from a substantial reduction at the loans reference to parity at the cash reference. The practical implication is operational rather than predictive: centering avoids the catastrophic raw-MA divergence spikes that occur at isolated rolling origins, which matters for production workflows in which posterior simulation feeds downstream stress tests. The adjustment is analytic and plug-in, and requires only a local change to the MA innovation calculation.

stat.ME

Slomads Rising: Stay Length Shifts in Digital Nomad Travel, United States 2019-2024

Using all U.S. Airbnb reservations created in 2019-2024 (booking-count weighted), we quantify pandemic-era shifts in nights per booking (NPB) and the mechanism behind them. The mean rose from 3.68 pre-COVID to 4.36 during restrictions and stabilized near 4.07 post-2021 (about 10% above 2019); the booking-weighted median moved from 2 to 3 nights. A two-parameter log-normal fits best by wide AIC/BIC margins, indicating heavy tails. A negative-binomial model with month effects implies post-vaccine bookings are 6.5% shorter than restriction-era bookings, while pre-COVID bookings are 16% shorter. In a two-part model at 28 nights, the booking share of month-plus stays rose from 1.43% (pre) to 2.72% (restriction) and settled at 2.04% (post); conditional means among long stays were about 55-60 nights. Thus the higher average reflects more long stays rather than longer long stays. A SARIMA(0,1,1)(0,1,1)12 with pandemic-phase dummies improves fit (LR=8.39, df=2, p=0.015), consistent with a structural level shift.

q-fin.ST

A Bayesian Dirichlet Auto-Regressive Conditional Heteroskedasticity Model for Forecasting Currency Shares

We analyze daily Airbnb service-fee shares across eleven settlement currencies, a compositional series that shows bursts of volatility after shocks such as the COVID-19 pandemic. Standard Dirichlet time series models assume constant precision and therefore miss these episodes. We introduce B-DARMA-DARCH, a Bayesian Dirichlet autoregressive moving average model with a Dirichlet ARCH component, which lets the precision parameter follow an ARMA recursion. The specification preserves the Dirichlet likelihood so forecasts remain valid compositions while capturing clustered volatility. Simulations and out-of-sample tests show that B-DARMA-DARCH lowers forecast error and improves interval calibration relative to Dirichlet ARMA and log-ratio VARMA benchmarks, providing a concise framework for settings where both the level and the volatility of proportions matter.

stat.ME

Forecasting the U.S. Renewable-Energy Mix with an ALR-BDARMA Compositional Time-Series Framework

Accurate forecasts of the US renewable-generation mix are critical for planning transmission upgrades, sizing storage, and setting balancing-market rules. We present a Bayesian Dirichlet ARMA (BDARMA) model for monthly shares of hydro, geothermal, solar, wind, wood, municipal waste, and biofuels from January 2010 to January 2025. The mean vector follows a parsimonious VAR(2) in additive-log-ratio space, while the Dirichlet concentration parameter combines an intercept with ten Fourier harmonics, letting predictive dispersion expand or contract with the seasons. A 61-split rolling-origin study generates twelve-month density forecasts from January 2019 to January 2024. Relative to three benchmarks, a Gaussian VAR(2) in transform space, a seasonal naive copy of last year's proportions, and a drift-free additive-log-ratio random walk, BDARMA lowers the mean continuous ranked probability score by fifteen to sixty percent, achieves component-wise ninety percent interval coverage close to nominal, and matches Gaussian VAR point accuracy through eight months with a maximum loss of 0.02 Aitchison units thereafter. BDARMA therefore delivers sharp, well-calibrated probabilistic forecasts of multivariate renewable-energy shares without sacrificing point precision.

stat.AP

Sensitivity Analysis of Priors in the Bayesian Dirichlet Auto-Regressive Moving Average Model

Prior choice can strongly influence Bayesian Dirichlet ARMA (B-DARMA) inference for compositional time-series. Using simulations with (i) correct lag order, (ii) overfitting, and (iii) underfitting, we assess five priors: weakly-informative, horseshoe, Laplace, mixture-of-normals, and hierarchical. With the true lag order, all priors achieve comparable RMSE, though horseshoe and hierarchical slightly reduce bias. Under overfitting, aggressive shrinkage-especially the horseshoe-suppresses noise and improves forecasts, yet no prior rescues a model that omits essential VAR or VMA terms. We then fit B-DARMA to daily SP 500 sector weights using an intentionally large lag structure. Shrinkage priors curb spurious dynamics, whereas weakly-informative priors magnify errors in volatile sectors. Two lessons emerge: (1) match shrinkage strength to the degree of overparameterization, and (2) prioritize correct lag selection, because no prior repairs structural misspecification. These insights guide prior selection and model complexity management in high-dimensional compositional time-series applications.

stat.ME

Two-Part Forecasting for Time-Shifted Metrics

Katz, Savage, and Brusch propose a two-part forecasting method for sectors where event timing differs from recording time. They treat forecasting as a time-shift operation, using univariate time series for total bookings and a Bayesian Dirichlet Auto-Regressive Moving Average (B-DARMA) model to allocate bookings across trip dates based on lead time. Analysis of Airbnb data shows that this approach is interpretable, flexible, and potentially more accurate for forecasting demand across multiple time axes.

stat.AP

Bayesian Shrinkage in High-Dimensional VAR Models: A Comparative Study

High-dimensional vector autoregressive (VAR) models offer a versatile framework for multivariate time series analysis, yet face critical challenges from over-parameterization and uncertain lag order. In this paper, we systematically compare three Bayesian shrinkage priors (horseshoe, lasso, and normal) and two frequentist regularization approaches (ridge and nonparametric shrinkage) under three carefully crafted simulation scenarios. These scenarios encompass (i) overfitting in a low-dimensional setting, (ii) sparse high-dimensional processes, and (iii) a combined scenario where both large dimension and overfitting complicate inference. We evaluate each method in quality of parameter estimation (root mean squared error, coverage, and interval length) and out-of-sample forecasting (one-step-ahead forecast RMSE). Our findings show that local-global Bayesian methods, particularly the horseshoe, dominate in maintaining accurate coverage and minimizing parameter error, even when the model is heavily over-parameterized. Frequentist ridge often yields competitive point forecasts but underestimates uncertainty, leading to sub-nominal coverage. A real-data application using macroeconomic variables from Canada illustrates how these methods perform in practice, reinforcing the advantages of local-global priors in stabilizing inference when dimension or lag order is inflated.

stat.ME

Lead Times in Flux: Analyzing Airbnb Booking Dynamics During Global Upheavals (2018-2022)

Short-term shifts in booking behaviors can disrupt forecasting in the travel and hospitality industry, especially during global crises. Traditional metrics like average or median lead times often overlook important distribution changes. This study introduces a normalized L1 (Manhattan) distance to assess Airbnb booking lead time divergences from 2018 to 2022, focusing on the COVID-19 pandemic across four major U.S. cities. We identify a two-phase disruption: an abrupt change at the pandemic's onset followed by partial recovery with persistent deviations from pre-2018 patterns. Our method reveals changes in travelers' planning horizons that standard statistics miss, highlighting the need to analyze the entire lead-time distribution for more accurate demand forecasting and pricing strategies. The normalized L1 metric provides valuable insights for tourism stakeholders navigating ongoing market volatility.

stat.AP

A Bayesian Dirichlet Auto-Regressive Moving Average Model for Forecasting Lead Times

Lead time data is compositional data found frequently in the hospitality industry. Hospitality businesses earn fees each day, however these fees cannot be recognized until later. For business purposes, it is important to understand and forecast the distribution of future fees for the allocation of resources, for business planning, and for staffing. Motivated by 5 years of daily fees data, we propose a new class of Bayesian time series models, a Bayesian Dirichlet Auto-Regressive Moving Average (B-DARMA) model for compositional time series, modeling the proportion of future fees that will be recognized in 11 consecutive 30 day windows and 1 last consecutive 35 day window. Each day's compositional datum is modeled as Dirichlet distributed given the mean and a scale parameter. The mean is modeled with a Vector Autoregressive Moving Average process after transforming with an additive log ratio link function and depends on previous compositional data, previous compositional parameters and daily covariates. The B-DARMA model offers solutions to data analyses of large compositional vectors and short or long time series, offers efficiency gains through choice of priors, provides interpretable parameters for inference, and makes reasonable forecasts.

stat.ME