SearcharxivSearch

arXiv subjects

Hedibert F. Lopes

Publications and source records attributed to Hedibert F. Lopes.

14 recordsLinked to original sources

A Conversation with Mike West

Mike West is currently the Arts & Sciences Distinguished Professor Emeritus of Statistics and Decision Sciences at Duke University. Mike's research in Bayesian analysis spans multiple interlinked areas: theory and methods of dynamic models in time series analysis, foundations of inference and decision analysis, multivariate and latent structure analysis, stochastic computation and optimisation, among others. Inter-disciplinary R&D has ranged across applications in commercial forecasting, dynamic networks, finance, econometrics, signal processing, climatology, systems biology, genomics and neuroscience, among other areas. Among Mike's currently active research areas are forecasting, causal prediction and decision analysis in business, economic policy and finance, as well as in personal decision making. Mike led the development of academic statistics at Duke University from 1990-2002, and has been broadly engaged in professional leadership elsewhere. He is past president of the International Society for Bayesian Analysis (ISBA), and has served in founding roles and as board member for several professional societies, national and international centres and institutes. Recipient of numerous awards, Mike has been active in research with various companies, banks, government agencies and academic centres, co-founder of a successful biotechnology company, and board member for several financial and IT companies. He has published 4 books, several edited volumes and over 200 papers. Mike has worked with many undergraduate and Master's research students, and as of 2025 has mentored around 65 primary PhD students and postdoctoral associates who moved to academic, industrial or governmental positions involving advanced statistical and data science research.

stat.OT

Lower-dimensional posterior density and cluster summaries for overparameterized Bayesian models

The usefulness of Bayesian models for density and cluster estimation is well established across multiple literatures. However, there is still a known tension between the use of simpler, more interpretable models and more flexible, complex ones. In this paper, we propose a novel method that integrates these two approaches by projecting the fit of a flexible, overparameterized model onto a lower-dimensional parametric surrogate, which serves as a summary. This process increases interpretability while preserving most of the fit of the original model. Our approach involves three main steps. First, we fit the data using nonparametric or overparameterized models. Second, we project the posterior predictive distribution of the original model onto a sequence of parametric summary point estimates with varying dimensions using a decision-theoretic approach. Finally, given the parametric summary estimate, obtained in the second step, that best approximates the original model, we construct uncertainty quantification for this summary by projecting the original posterior distribution. We demonstrate the effectiveness of our method for generating summaries for both nonparametric and overparameterized models, delivering both point estimates and uncertainty quantification for density and cluster summaries across synthetic and real datasets.

stat.ME

Minnesota BART

Vector autoregression (VAR) models are widely used for forecasting and macroeconomic analysis, yet they remain limited by their reliance on a linear parameterization. Recent research has introduced nonparametric alternatives, such as Bayesian additive regression trees (BART), which provide flexibility without strong parametric assumptions. However, existing BART-based frameworks do not account for time dependency or allow for sparse estimation in the construction of regression tree priors, leading to noisy and inefficient high-dimensional representations. This paper introduces a sparsity-inducing Dirichlet hyperprior on the regression tree's splitting probabilities, allowing for automatic variable selection and high-dimensional VARs. Additionally, we propose a structured shrinkage prior that decreases the probability of splitting on higher-order lags, aligning with the Minnesota prior's principles. Empirical results demonstrate that our approach improves predictive accuracy over the baseline BART prior and Bayesian VAR (BVAR), particularly in capturing time-dependent relationships and enhancing density forecasts. These findings highlight the potential of developing domain-specific nonparametric methods in macroeconomic forecasting.

stat.ME

A Bayesian Additive Regression Tree Model for Learning Conditional Average Treatment Effects in Regression Discontinuity Designs

This paper develops an effective Bayesian approach to conditional average treatment effect (CATE) estimation in regression discontinuity designs (RDD), an increasingly prevalent form of quasi-experiment that facilitates causal inference. Earlier Bayesian approaches do not easily accommodate CATE estimation while recent frequentist approaches to this problem assume a known basis expansion, a steep model specification requirement that our approach avoids. The new model is a variant of a Bayesian additive regression tree (BART) model with linear leaf-level regressions on the running variable and a treatment dummy (and their interaction). The model adaptively partitions covariate space into regions where the slope on the running variable appreciably differs, providing interpretable Bayesian inference on conditional average treatment effects near the cutoff.

stat.ME

Dynamic Portfolio Allocation in High Dimensions using Sparse Risk Factors

We propose a fast and flexible method to scale multivariate return volatility predictions up to high-dimensions using a dynamic risk factor model. Our approach increases parsimony via time-varying sparsity on factor loadings and is able to sequentially learn the use of constant or time-varying parameters and volatilities. We show in a dynamic portfolio allocation problem with 452 stocks from the S&P 500 index that our dynamic risk factor model is able to produce more stable and sparse predictions, achieving not just considerable portfolio performance improvements but also higher utility gains for the mean-variance investor compared to the traditional Wishart benchmark and the passive investment on the market index.

q-fin.ST

Dynamic Ordering Learning in Multivariate Forecasting

In many fields where the main goal is to produce sequential forecasts for decision making problems, the good understanding of the contemporaneous relations among different series is crucial for the estimation of the covariance matrix. In recent years, the modified Cholesky decomposition appeared as a popular approach to covariance matrix estimation. However, its main drawback relies on the imposition of the series ordering structure. In this work, we propose a highly flexible and fast method to deal with the problem of ordering uncertainty in a dynamic fashion with the use of Dynamic Order Probabilities. We apply the proposed method in two different forecasting contexts. The first is a dynamic portfolio allocation problem, where the investor is able to learn the contemporaneous relationships among different currencies improving final decisions and economic performance. The second is a macroeconomic application, where the econometrician can adapt sequentially to new economic environments, switching the contemporaneous relations among macroeconomic variables over time.

econ.EM

Trend-Following Strategies via Dynamic Momentum Learning

Time series momentum strategies are widely applied in the quantitative financial industry and its academic research has grown rapidly since the work of Moskowitz, Ooi and Pedersen (2012). However, trading signals are usually obtained via simple observation of past return measurements. In this article we study the benefits of incorporating dynamic econometric models to sequentially learn the time-varying importance of different look-back periods for individual assets. By the use of a dynamic binary classifier model, the investor is able to switch between time-varying or constant relations between past momentum and future returns, dynamically combining or selecting different momentum speeds during turning points, improving trading signals accuracy and portfolio performance. Using data from 56 future contracts we show that a mean-variance investor will be willing to pay a considerable management fee to switch from the traditional naive time series momentum strategy to the dynamic classifier approach.

q-fin.ST

Decoupling Shrinkage and Selection in Gaussian Linear Factor Analysis

Factor Analysis is a popular method for modeling dependence in multivariate data. However, determining the number of factors and obtaining a sparse orientation of the loadings are still major challenges. In this paper, we propose a decision-theoretic approach that brings to light the relation between a sparse representation of the loadings and factor dimension. This relation is done through a summary from information contained in the multivariate posterior. To construct such summary, we introduce a three-step approach. In the first step, the model is fitted with a conservative factor dimension. In the second step, a series of sparse point-estimates, with a decreasing number of factors, is obtained by minimizing an expected predictive loss function. In step three, the degradation in utility in relation to the sparse loadings and factor dimensions is displayed in the posterior summary. The findings are illustrated with applications in classical data from the Factor Analysis literature. We used different prior choices and factor dimensions to demonstrate the flexibility of the proposed method.

stat.ME

Dynamic sparsity on dynamic regression models

In the present work, we consider variable selection and shrinkage for the Gaussian dynamic linear regression within a Bayesian framework. In particular, we propose a novel method that allows for time-varying sparsity, based on an extension of spike-and-slab priors for dynamic models. This is done by assigning appropriate Markov switching priors for the time-varying coefficients' variances, extending the previous work of Ishwaran and Rao (2005). Furthermore, we investigate different priors, including the common Inverted gamma prior for the process variances, and other mixture prior distributions such as Gamma priors for both the spike and the slab, which leads to a mixture of Normal-Gammas priors (Griffin ad Brown, 2010) for the coefficients. In this sense, our prior can be view as a dynamic variable selection prior which induces either smoothness (through the slab) or shrinkage towards zero (through the spike) at each time point. The MCMC method used for posterior computation uses Markov latent variables that can assume binary regimes at each time point to generate the coefficients' variances. In that way, our model is a dynamic mixture model, thus, we could use the algorithm of Gerlach et al (2000) to generate the latent processes without conditioning on the states. Finally, our approach is exemplified through simulated examples and a real data application.

stat.ME

The Illusion of the Illusion of Sparsity: An exercise in prior sensitivity

The emergence of Big Data raises the question of how to model economic relations when there is a large number of possible explanatory variables. We revisit the issue by comparing the possibility of using dense or sparse models in a Bayesian approach, allowing for variable selection and shrinkage. More specifically, we discuss the results reached by Giannone, Lenza, and Primiceri (2020) through a "Spike-and-Slab" prior, which suggest an "illusion of sparsity" in economic data, as no clear patterns of sparsity could be detected. We make a further revision of the posterior distributions of the model, and propose three experiments to evaluate the robustness of the adopted prior distribution. We find that the pattern of sparsity is sensitive to the prior distribution of the regression coefficients, and present evidence that the model indirectly induces variable selection and shrinkage, which suggests that the "illusion of sparsity" could be, itself, an illusion. Code is available on github.com/bfava/IllusionOfIllusion.

stat.ME

Learning a latent pattern of heterogeneity in the innovation rates of a time series of counts

We develop a Bayesian hierarchical semiparametric model for phenomena related to time series of counts. The main feature of the model is its capability to learn a latent pattern of heterogeneity in the distribution of the process innovation rates, which are softly clustered through time with the help of a Dirichlet process placed at the top of the model hierarchy. The probabilistic forecasting capabilities of the model are put to test in the analysis of crime data in Pittsburgh, with favorable results.

stat.ME

Bayesian Hypothesis Testing: Redux

Bayesian hypothesis testing is re-examined from the perspective of an a priori assessment of the test statistic distribution under the alternative. By assessing the distribution of an observable test statistic, rather than prior parameter values, we provide a practical default Bayes factor which is straightforward to interpret. To illustrate our methodology, we provide examples where evidence for a Bayesian strikingly supports the null, but leads to rejection under a classical test. Finally, we conclude with directions for future research.

math.ST

Measuring the vulnerability of the Uruguayan population to vector-borne diseases via spatially hierarchical factor models

We propose a model-based vulnerability index of the population from Uruguay to vector-borne diseases. We have available measurements of a set of variables in the census tract level of the 19 Departmental capitals of Uruguay. In particular, we propose an index that combines different sources of information via a set of micro-environmental indicators and geographical location in the country. Our index is based on a new class of spatially hierarchical factor models that explicitly account for the different levels of hierarchy in the country, such as census tracts within the city level, and cities in the country level. We compare our approach with that obtained when data are aggregated in the city level. We show that our proposal outperforms current and standard approaches, which fail to properly account for discrepancies in the region sizes, for example, number of census tracts. We also show that data aggregation can seriously affect the estimation of the cities vulnerability rankings under benchmark models.

stat.AP

Particle Learning and Smoothing

Particle learning (PL) provides state filtering, sequential parameter learning and smoothing in a general class of state space models. Our approach extends existing particle methods by incorporating the estimation of static parameters via a fully-adapted filter that utilizes conditional sufficient statistics for parameters and/or states as particles. State smoothing in the presence of parameter uncertainty is also solved as a by-product of PL. In a number of examples, we show that PL outperforms existing particle filtering alternatives and proves to be a competitor to MCMC.

stat.ME