SearcharxivSearch

arXiv subjects

Stephan Smeekes

Publications and source records attributed to Stephan Smeekes.

At least 19 recordsLinked to original sources

Drivers of Success: A Bayesian State-Space Model to Disentangling Latent Driver and Constructor Abilities in Formula One

Formula One outcomes reflect the joint contributions of drivers and constructors, but these contributions are unobserved and vary over time. We propose a Bayesian state-space model that disentangles dynamic driver and constructor abilities using two observed outcomes: fastest qualifying lap times and race rankings. Both outcomes depend jointly on latent driver and constructor states that evolve at the Grand Prix level, while the race equation additionally accounts for starting-grid position. The decomposition is supported by constraints that center the driver and constructor abilities at zero, together with variation in driver-constructor assignments over time. Bayesian inference is performed using the No-U-Turn sampler under weakly informative priors that treat driver and constructor abilities symmetrically. Applying the model to the Formula One hybrid era from 2014 to 2021, we find substantial heterogeneity in both driver and constructor abilities. Driver abilities are generally more stable over time, whereas constructor abilities exhibit greater variation and, for many driver--constructor combinations, contribute more strongly to observed performance.

econ.EM

Sparse Tree-Based Aggregation for Time Series Regressions

High-dimensional time series regressions are often regularized to produce sparse coefficients. We show that temporal aggregation provides a powerful alternative to reduce dimensionality in high-order autoregressions and mixed-frequency regressions. To this end, we propose StarTime (Sparse Tree-based Aggregation for Time Series), a convex penalization method that uses a temporal tree to arrange lags hierarchically from high to low frequency. StarTime then flexibly selects coefficients to be aggregated at possibly varying frequencies, sparse or a combination thereof. We provide new error bounds for StarTime, demonstrate improved estimation accuracy and recovery of aggregation and sparsity in simulations relative to benchmarks, and illustrate StarTime's relevance for financial and macroeconomic applications.

econ.EM

Autotune: fast, accurate, and automatic tuning parameter selection for Lasso

Least absolute shrinkage and selection operator (Lasso), the popular variable selection engine for high-dimensional regression, is commonly tuned using cross-validation (CV). This is known to be slow and loses accuracy in high-dimension, low signal-to-noise ratio (SNR) settings. These issues are exacerbated for high-dimensional time series models such as the vector autoregression (VAR), where time series cross-validaiton (TSCV) is used for tuning. We propose $\texttt{autotune}$, a strategy for the Lasso to tune itself automatically by exploiting information contained in partial residuals computed during a single Lasso fit. The strategy can also be viewed as alternately estimating noise standard deviation and column space of relevant predictors. Numerical experiments on regression and VAR models show that $\texttt{autotune}$ is faster than existing alternatives, and more accurate when the SNR is low. It also provides an accurate estimator of noise scale and diagnostic plots akin to screeplots for checking model sparsity. We demonstrate the benefit of $\texttt{autotune}$ on several real data sets and develop an R package available on CRAN.

stat.ME

Estimation of Latent Group Structures in Time-Varying Panel Data Models

We consider panel data models where coefficients change smoothly over time and follow a latent group structure, being homogeneous within but heterogeneous across groups. To jointly estimate the group membership and group-specific coefficient trajectories, we propose FUSE-TIME, a pairwise adaptive group fused-Lasso estimator combined with polynomial spline sieves. We establish consistency, derive the asymptotic distributions of the penalized sieve estimator and its post-selection version, and show oracle efficiency. Monte Carlo experiments demonstrate strong finite-sample performance in terms of estimation accuracy and group identification. An application to the CO2 intensity of GDP highlights the relevance of addressing both cross-sectional heterogeneity and time-variance in empirical exercises.

econ.EM

High-Dimensional Granger Causality for Climatic Attribution

In this paper we test for Granger causality in high-dimensional vector autoregressive models (VARs) to disentangle and interpret the complex causal chains linking radiative forcings and global temperatures. By allowing for high dimensionality in the model, we can enrich the information set with relevant natural and anthropogenic forcing variables to obtain reliable causal relations. This provides a step forward from existing climatology literature, which has mostly treated these variables in isolation in small models. Additionally, our framework allows to disregard the order of integration of the variables by directly estimating the VAR in levels, thus avoiding accumulating biases coming from unit-root and cointegration tests. This is of particular appeal for climate time series which are well known to contain stochastic trends and long memory. We are thus able to establish causal networks linking radiative forcings to global temperatures and to connect radiative forcings among themselves, thereby allowing for tracing the path of dynamic causal effects through the system.

econ.EM

Transmission Channel Analysis in Dynamic Models

We propose a framework for analysing transmission channels in a large class of dynamic models. We formulate our approach both using graph theory and potential outcomes, which we show to be equivalent. Our method, labelled Transmission Channel Analysis (TCA), allows for the decomposition of total effects captured by impulse response functions into the effects flowing through transmission channels, thereby providing a quantitative assessment of the strength of various well-defined channels. We establish that this requires no additional identification assumptions beyond the identification of the structural shock whose effects the researcher wants to decompose. Additionally, we prove that impulse response functions are sufficient statistics for the computation of transmission effects. We demonstrate the empirical relevance of TCA for policy evaluation by decomposing the effects of policy shocks arising from a variety of popular macroeconomic models.

econ.EM

Local Projection Inference in High Dimensions

In this paper, we estimate impulse responses by local projections in high-dimensional settings. We use the desparsified (de-biased) lasso to estimate the high-dimensional local projections, while leaving the impulse response parameter of interest unpenalized. We establish the uniform asymptotic normality of the proposed estimator under general conditions. Finally, we demonstrate small sample performance through a simulation study and consider two canonical applications in macroeconomic research on monetary policy and government spending.

econ.EM

Inference in Non-stationary High-Dimensional VARs

In this paper we construct an inferential procedure for Granger causality in high-dimensional non-stationary vector autoregressive (VAR) models. Our method does not require knowledge of the order of integration of the time series under consideration. We augment the VAR with at least as many lags as the suspected maximum order of integration, an approach which has been proven to be robust against the presence of unit roots in low dimensions. We prove that we can restrict the augmentation to only the variables of interest for the testing, thereby making the approach suitable for high dimensions. We combine this lag augmentation with a post-double-selection procedure in which a set of initial penalized regressions is performed to select the relevant variables for both the Granger causing and caused variables. We then establish uniform asymptotic normality of a second-stage regression involving only the selected variables. Finite sample simulations show good performance, an application to investigate the (predictive) causes and effects of economic uncertainty illustrates the need to allow for unknown orders of integration.

econ.EM

A Residual Bootstrap for Conditional Value-at-Risk

A fixed-design residual bootstrap method is proposed for the two-step estimator of Francq and Zakoïan (2015) associated with the conditional Value-at-Risk. The bootstrap's consistency is proven for a general class of volatility models and intervals are constructed for the conditional Value-at-Risk. A simulation study reveals that the equal-tailed percentile bootstrap interval tends to fall short of its nominal value. In contrast, the reversed-tails bootstrap interval yields accurate coverage. We also compare the theoretically analyzed fixed-design bootstrap with the recursive-design bootstrap. It turns out that the fixed-design bootstrap performs equally well in terms of average coverage, yet leads on average to shorter intervals in smaller samples. An empirical application illustrates the interval estimation.

econ.EM

Sparse High-Dimensional Vector Autoregressive Bootstrap

We introduce a high-dimensional multiplier bootstrap for time series data based on capturing dependence through a sparsely estimated vector autoregressive model. We prove its consistency for inference on high-dimensional means under two different moment assumptions on the errors, namely sub-gaussian moments and a finite number of absolute moments. In establishing these results, we derive a Gaussian approximation for the maximum mean of a linear process, which may be of independent interest.

econ.EM

Lasso Inference for High-Dimensional Time Series

In this paper we develop valid inference for high-dimensional time series. We extend the desparsified lasso to a time series setting under Near-Epoch Dependence (NED) assumptions allowing for non-Gaussian, serially correlated and heteroskedastic processes, where the number of regressors can possibly grow faster than the time dimension. We first derive an error bound under weak sparsity, which, coupled with the NED assumption, means this inequality can also be applied to the (inherently misspecified) nodewise regressions performed in the desparsified lasso. This allows us to establish the uniform asymptotic normality of the desparsified lasso under general conditions, including for inference on parameters of increasing dimensions. Additionally, we show consistency of a long-run variance estimator, thus providing a complete set of tools for performing inference in high-dimensional linear time series models. Finally, we perform a simulation exercise to demonstrate the small sample properties of the desparsified lasso in common time series settings.

econ.EM

bootUR: An R Package for Bootstrap Unit Root Tests

Unit root tests form an essential part of any time series analysis. We provide practitioners with a single, unified framework for comprehensive and reliable unit root testing in the R package bootUR.The package's backbone is the popular augmented Dickey-Fuller test paired with a union of rejections principle, which can be performed directly on single time series or multiple (including panel) time series. Accurate inference is ensured through the use of bootstrap methods. The package addresses the needs of both novice users, by providing user-friendly and easy-to-implement functions with sensible default options, as well as expert users, by giving full user-control to adjust the tests to one's desired settings. Our parallelized C++ implementation ensures that all unit root tests are scalable to datasets containing many time series.

econ.EM

Min(d)ing the President: A text analytic approach to measuring tax news

Economic agents react to signals about future tax policy changes. Consequently, estimating their macroeconomic effects requires identification of such signals. We propose a novel text analytic approach for transforming textual information into an economically meaningful time series. Using this method, we create a tax news measure from all publicly available post-war communications of U.S. presidents. Our measure predicts the direction and size of future tax changes and contains signals not present in previously considered (narrative) measures of tax changes. We investigate the effects of tax news and find that, for long anticipation horizons, pre-implementation effects lead initially to contractions in output.

econ.EM

Granger Causality Testing in High-Dimensional VARs: a Post-Double-Selection Procedure

We develop an LM test for Granger causality in high-dimensional VAR models based on penalized least squares estimations. To obtain a test retaining the appropriate size after the variable selection done by the lasso, we propose a post-double-selection procedure to partial out effects of nuisance variables and establish its uniform asymptotic validity. We conduct an extensive set of Monte-Carlo simulations that show our tests perform well under different data generating processes, even without sparsity. We apply our testing procedure to find networks of volatility spillovers and we find evidence that causal relationships become clearer in high-dimensional compared to standard low-dimensional VARs.

econ.EM

An Automated Approach Towards Sparse Single-Equation Cointegration Modelling

In this paper we propose the Single-equation Penalized Error Correction Selector (SPECS) as an automated estimation procedure for dynamic single-equation models with a large number of potentially (co)integrated variables. By extending the classical single-equation error correction model, SPECS enables the researcher to model large cointegrated datasets without necessitating any form of pre-testing for the order of integration or cointegrating rank. Under an asymptotic regime in which both the number of parameters and time series observations jointly diverge to infinity, we show that SPECS is able to consistently estimate an appropriate linear combination of the cointegrating vectors that may occur in the underlying DGP. In addition, SPECS is shown to enable the correct recovery of sparsity patterns in the parameter space and to posses the same limiting distribution as the OLS oracle procedure. A simulation study shows strong selective capabilities, as well as superior predictive performance in the context of nowcasting compared to high-dimensional models that ignore cointegration. An empirical application to nowcasting Dutch unemployment rates using Google Trends confirms the strong practical performance of our procedure.

econ.EM

A statistical analysis of time trends in atmospheric ethane

Ethane is the most abundant non-methane hydrocarbon in the Earth's atmosphere and an important precursor of tropospheric ozone through various chemical pathways. Ethane is also an indirect greenhouse gas (global warming potential), influencing the atmospheric lifetime of methane through the consumption of the hydroxyl radical (OH). Understanding the development of trends and identifying trend reversals in atmospheric ethane is therefore crucial. Our dataset consists of four series of daily ethane columns obtained from ground-based FTIR measurements. As many other decadal time series, our data are characterized by autocorrelation, heteroskedasticity, and seasonal effects. Additionally, missing observations due to instrument failure or unfavorable measurement conditions are common in such series. The goal of this paper is therefore to analyze trends in atmospheric ethane with statistical tools that correctly address these data features. We present selected methods designed for the analysis of time trends and trend reversals. We consider bootstrap inference on broken linear trends and smoothly varying nonlinear trends. In particular, for the broken trend model, we propose a bootstrap method for inference on the break location and the corresponding changes in slope. For the smooth trend model we construct simultaneous confidence bands around the nonparametrically estimated trend. Our autoregressive wild bootstrap approach, combined with a seasonal filter, is able to handle all issues mentioned above.

stat.AP

A dynamic factor model approach to incorporate Big Data in state space models for official statistics

In this paper we consider estimation of unobserved components in state space models using a dynamic factor approach to incorporate auxiliary information from high-dimensional data sources. We apply the methodology to unemployment estimation as done by Statistics Netherlands, who uses a multivariate state space model to produce monthly figures for the unemployment using series observed with the labour force survey (LFS). We extend the model by including auxiliary series of Google Trends about job-search and economic uncertainty, and claimant counts, partially observed at higher frequencies. Our factor model allows for nowcasting the variable of interest, providing reliable unemployment estimates in real-time before LFS data become available.

econ.EM

Autoregressive Wild Bootstrap Inference for Nonparametric Trends

In this paper we propose an autoregressive wild bootstrap method to construct confidence bands around a smooth deterministic trend. The bootstrap method is easy to implement and does not require any adjustments in the presence of missing data, which makes it particularly suitable for climatological applications. We establish the asymptotic validity of the bootstrap method for both pointwise and simultaneous confidence bands under general conditions, allowing for general patterns of missing data, serial dependence and heteroskedasticity. The finite sample properties of the method are studied in a simulation study. We use the method to study the evolution of trends in daily measurements of atmospheric ethane obtained from a weather station in the Swiss Alps, where the method can easily deal with the many missing observations due to adverse weather conditions.

stat.ME