SearcharxivSearch

arXiv subjects

Robert E. Weiss

Publications and source records attributed to Robert E. Weiss.

13 recordsLinked to original sources

A Bayesian Dirichlet Auto-Regressive Conditional Heteroskedasticity Model for Forecasting Currency Shares

We analyze daily Airbnb service-fee shares across eleven settlement currencies, a compositional series that shows bursts of volatility after shocks such as the COVID-19 pandemic. Standard Dirichlet time series models assume constant precision and therefore miss these episodes. We introduce B-DARMA-DARCH, a Bayesian Dirichlet autoregressive moving average model with a Dirichlet ARCH component, which lets the precision parameter follow an ARMA recursion. The specification preserves the Dirichlet likelihood so forecasts remain valid compositions while capturing clustered volatility. Simulations and out-of-sample tests show that B-DARMA-DARCH lowers forecast error and improves interval calibration relative to Dirichlet ARMA and log-ratio VARMA benchmarks, providing a concise framework for settings where both the level and the volatility of proportions matter.

stat.ME

Bayesian Shrinkage in High-Dimensional VAR Models: A Comparative Study

High-dimensional vector autoregressive (VAR) models offer a versatile framework for multivariate time series analysis, yet face critical challenges from over-parameterization and uncertain lag order. In this paper, we systematically compare three Bayesian shrinkage priors (horseshoe, lasso, and normal) and two frequentist regularization approaches (ridge and nonparametric shrinkage) under three carefully crafted simulation scenarios. These scenarios encompass (i) overfitting in a low-dimensional setting, (ii) sparse high-dimensional processes, and (iii) a combined scenario where both large dimension and overfitting complicate inference. We evaluate each method in quality of parameter estimation (root mean squared error, coverage, and interval length) and out-of-sample forecasting (one-step-ahead forecast RMSE). Our findings show that local-global Bayesian methods, particularly the horseshoe, dominate in maintaining accurate coverage and minimizing parameter error, even when the model is heavily over-parameterized. Frequentist ridge often yields competitive point forecasts but underestimates uncertainty, leading to sub-nominal coverage. A real-data application using macroeconomic variables from Canada illustrates how these methods perform in practice, reinforcing the advantages of local-global priors in stabilizing inference when dimension or lag order is inflated.

stat.ME

Sensitivity Analysis of Priors in the Bayesian Dirichlet Auto-Regressive Moving Average Model

Prior choice can strongly influence Bayesian Dirichlet ARMA (B-DARMA) inference for compositional time-series. Using simulations with (i) correct lag order, (ii) overfitting, and (iii) underfitting, we assess five priors: weakly-informative, horseshoe, Laplace, mixture-of-normals, and hierarchical. With the true lag order, all priors achieve comparable RMSE, though horseshoe and hierarchical slightly reduce bias. Under overfitting, aggressive shrinkage-especially the horseshoe-suppresses noise and improves forecasts, yet no prior rescues a model that omits essential VAR or VMA terms. We then fit B-DARMA to daily SP 500 sector weights using an intentionally large lag structure. Shrinkage priors curb spurious dynamics, whereas weakly-informative priors magnify errors in volatile sectors. Two lessons emerge: (1) match shrinkage strength to the degree of overparameterization, and (2) prioritize correct lag selection, because no prior repairs structural misspecification. These insights guide prior selection and model complexity management in high-dimensional compositional time-series applications.

stat.ME

A Bayesian Dirichlet Auto-Regressive Moving Average Model for Forecasting Lead Times

Lead time data is compositional data found frequently in the hospitality industry. Hospitality businesses earn fees each day, however these fees cannot be recognized until later. For business purposes, it is important to understand and forecast the distribution of future fees for the allocation of resources, for business planning, and for staffing. Motivated by 5 years of daily fees data, we propose a new class of Bayesian time series models, a Bayesian Dirichlet Auto-Regressive Moving Average (B-DARMA) model for compositional time series, modeling the proportion of future fees that will be recognized in 11 consecutive 30 day windows and 1 last consecutive 35 day window. Each day's compositional datum is modeled as Dirichlet distributed given the mean and a scale parameter. The mean is modeled with a Vector Autoregressive Moving Average process after transforming with an additive log ratio link function and depends on previous compositional data, previous compositional parameters and daily covariates. The B-DARMA model offers solutions to data analyses of large compositional vectors and short or long time series, offers efficiency gains through choice of priors, provides interpretable parameters for inference, and makes reasonable forecasts.

stat.ME

A Spatially Varying Hierarchical Random Effects Model for Longitudinal Macular Structural Data in Glaucoma Patients

We model longitudinal macular thickness measurements to monitor the course of glaucoma and prevent vision loss due to disease progression. The macular thickness varies over a 6$\times$6 grid of locations on the retina with additional variability arising from the imaging process at each visit. Currently, ophthalmologists estimate slopes using repeated simple linear regression for each subject and location. To estimate slopes more precisely, we develop a novel Bayesian hierarchical model for multiple subjects with spatially varying population-level and subject-level coefficients, borrowing information over subjects and measurement locations. We augment the model with visit effects to account for observed spatially correlated visit-specific errors. We model spatially varying (a) intercepts, (b) slopes, and (c) log residual standard deviations (SD) with multivariate Gaussian process priors with Matérn cross-covariance functions. Each marginal process assumes an exponential kernel with its own SD and spatial correlation matrix. We develop our models for and apply them to data from the Advanced Glaucoma Progression Study. We show that including visit effects in the model reduces error in predicting future thickness measurements and greatly improves model fit.

stat.AP

Hierarchical Topic Presence Models

Topic models analyze text from a set of documents. Documents are modeled as a mixture of topics, with topics defined as probability distributions on words. Inferences of interest include the most probable topics and characterization of a topic by inspecting the topic's highest probability words. Motivated by a data set of web pages (documents) nested in web sites, we extend the Poisson factor analysis topic model to hierarchical topic presence models for analyzing text from documents nested in known groups. We incorporate an unknown binary topic presence parameter for each topic at the web site and/or the web page level to allow web sites and/or web pages to be sparse mixtures of topics and we propose logistic regression modeling of topic presence conditional on web site covariates. We introduce local topics into the Poisson factor analysis framework, where each web site has a local topic not found in other web sites. Two data augmentation methods, the Chinese table distribution and Pólya-Gamma augmentation, aid in constructing our sampler. We analyze text from web pages nested in United States local public health department web sites to abstract topical information and understand national patterns in topic presence.

cs.IR

Local and Global Topics in Text Modeling of Web Pages Nested in Web Sites

Topic models are popular models for analyzing a collection of text documents. The models assert that documents are distributions over latent topics and latent topics are distributions over words. A nested document collection is where documents are nested inside a higher order structure such as stories in a book, articles in a journal, or web pages in a web site. In a single collection of documents, topics are global, or shared across all documents. For web pages nested in web sites, topic frequencies likely vary between web sites. Within a web site, topic frequencies almost certainly vary between web pages. A hierarchical prior for topic frequencies models this hierarchical structure and specifies a global topic distribution. Web site topic distributions vary around the global topic distribution and web page topic distributions vary around the web site topic distribution. In a nested collection of web pages, some topics are likely unique to a single web site. Local topics in a nested collection of web pages are topics unique to one web site. For US local health department web sites, brief inspection of the text shows local geographic and news topics specific to each department that are not present in others. Topic models that ignore the nesting may identify local topics, but do not label topics as local nor do they explicitly identify the web site owner of the local topic. For web pages nested inside web sites, local topic models explicitly label local topics and identifies the owning web site. This identification can be used to adjust inferences about global topics. In the US public health web site data, topic coverage is defined at the web site level after removing local topic words from pages. Hierarchical local topic models can be used to identify local topics, adjust inferences about if web sites cover particular health topics, and study how well health topics are covered.

cs.IR

Explaining the Decline of Child Mortality in 44 Developing Countries: A Bayesian Extension of Oaxaca Decomposition Methods

We investigate the decline of infant mortality in 42 low and middle income countries (LMIC) using detailed micro data from 84 Demographic and Health Surveys. We estimate infant mortality risk for each infant in our data and develop a novel extension of Oaxaca decomposition to understand the sources of these changes. We find that the decline in infant mortality is due to a declining propensity for parents with given characteristics to experience the death of an infant rather than due to changes in the distributions of these characteristics over time. Our results suggest that technical progress and policy health interventions in the form of public goods are the main drivers of the the recent decline in infant mortality in LMIC.

stat.AP

Measuring Within and Between Group Inequality in Early-Life Mortality Over Time: A Bayesian Approach with Application to India

Most studies on inequality in infant and child mortality compare average mortality rates between large groups of births, for example, comparing births from different countries, income groups, ethnicities, or different times. These studies do not measure within-group disparities. The few studies that have measured within-group variability in infant and child mortality have used tools from the income inequality literature, such as Gini indices. We show that the latter are inappropriate for infant and child mortality. We develop novel tools that are appropriate for analyzing infant and child mortality inequality, including inequality measures, covariate adjustments, and ANOVA methods. We illustrate how to handle uncertainty about complex inference targets, including ensembles of probabilities and kernel density estimates. We illustrate our methodology using a large data set from India, where we estimate infant and child mortality risk for over 400,000 births using a Bayesian hierarchical model. We show that most of the variance in mortality risk exists within groups of births, not between them, and thus that within-group mortality needs to be taken into account when assessing inequality in infant and child mortality. Our approach has broad applicability to many health indicators.

stat.AP

Sex, lies and self-reported counts: Bayesian mixture models for heaping in longitudinal count data via birth-death processes

Surveys often ask respondents to report nonnegative counts, but respondents may misremember or round to a nearby multiple of 5 or 10. This phenomenon is called heaping, and the error inherent in heaped self-reported numbers can bias estimation. Heaped data may be collected cross-sectionally or longitudinally and there may be covariates that complicate the inferential task. Heaping is a well-known issue in many survey settings, and inference for heaped data is an important statistical problem. We propose a novel reporting distribution whose underlying parameters are readily interpretable as rates of misremembering and rounding. The process accommodates a variety of heaping grids and allows for quasi-heaping to values nearly but not equal to heaping multiples. We present a Bayesian hierarchical model for longitudinal samples with covariates to infer both the unobserved true distribution of counts and the parameters that control the heaping process. Finally, we apply our methods to longitudinal self-reported counts of sex partners in a study of high-risk behavior in HIV-positive youth.

stat.AP

Functional Time Series Models for Ultrafine Particle Distributions

We propose Bayesian random effect functional time series models to model the impact of engine idling on ultrafine particle (UFP) counts inside school buses. UFPs are toxic to humans with health effects strongly linked to particle size. School engines emit particles primarily in the UFP size range and as school buses idle at bus stops, UFPs penetrate into cabins through cracks, doors, and windows. How UFP counts inside buses vary by particle size over time and under different idling conditions is not yet well understood. We model UFP counts at a given time with a cubic B-spline basis as a function of size and allow counts to increase over time at a size dependent rate once the engine turns on. We explore alternate parametric models for the engine-on increase which also vary smoothly over size. The log residual variance over size is modeled using a quadratic B-spline basis to account for heterogeneity and an autoregressive model is used for the residual. Model predictions are communicated graphically. These methods provide information needed for regulating vehicle emissions to minimize UFP exposure in the future.

stat.AP

Using a Birth-Death Process to Account for Reporting Errors in Longitudinal Self-reported Counts of Behavior

We analyze longitudinal self-reported counts of sexual partners from youth living with HIV. In self-reported survey data, subjects recall counts of events or behaviors such as the number of sexual partners or the number of drug uses in the past three months. Subjects with small counts may report the exact number, whereas subjects with large counts may have difficulty recalling the exact number. Thus, self-reported counts are noisy, and mis-reporting induces errors in the count variable. As a naive method for analyzing self-reported counts, the Poisson random effects model treats the observed counts as true counts and reporting errors in the outcome variable are ignored. Inferences are therefore based on incorrect information and may lead to conclusions unsupported by the data. We describe a Bayesian model for analyzing longitudinal self-reported count data that formally accounts for reporting error. We model reported counts conditional on underlying true counts using a linear birth-death process and use a Poisson random effects model to model the underlying true counts. A regression version of our model can identify characteristics of subjects with greater or lesser reporting error. We demonstrate several approaches to prior specification.

stat.ME

Meta-Analysis of Odds Ratios With Incomplete Extracted Data

A typical random effects meta-analysis of odds-ratios assumes binomially distributed numbers of events in a treatment and control group and requires the proportion of deaths to be extracted from published papers. This data is often not available in the publications due to loss to follow-up. When the Kaplan Meier survival plot is available, it is common practice to manually measure the needed information from the plot and infer the probability of survival and then to infer a best-guess of the number of deaths. Uncertainty introduced from theses guesses is not accounted for in current models. This naive approach leads to over-certain results and potentially inaccurate conclusions. We propose the Uncertain Reading-Estimated Events model to construct each study's contribution to the meta-analysis separately using the data available for extraction in the publications. We use real and simulated data to illustrate our methods. Meta-analysis based on the observed number of deaths lead to biased estimates while our proposed model does not. Our results show increases in the standard deviation of the log-odds as compared to a naive meta-analysis that assumes ideal extracted data, equivalent to a reduction of the overall sample size of 43% in our example.

stat.ME