SearcharxivSearch

arXiv subjects

Angelos Alexopoulos

Publications and source records attributed to Angelos Alexopoulos.

15 recordsLinked to original sources

Bias-robust causal inference for panel data

We develop a bias-robust causal inference method for observational panel data settings. Such methods typically impute untreated outcomes, so counterfactual error passes straight into the estimated treatment effect while conventional standard errors ignore it. We adapt bias-aware minimax methods, developed for estimating regression coefficients in factor-model panels, to a causal target: the average effect on the treated, which has to be imputed and may vary across units and periods. The estimator corrects the imputed counterfactual with weighted untreated residuals and reports intervals with an explicit allowance for the error that remains. In simulations the proposed method holds nominal coverage where alternatives such as the generalized synthetic control have almost none, especially when the factor rank is underfitted, at the cost of wider intervals. By applying the developed methodology to real data the estimated effect remains significant for counterfactual errors nearly twice the size that the design's placebos typically exhibit.

econ.EM

Exchangeable Gaussian Processes for Staggered-Adoption Policy Evaluation

We study the use of exchangeable multi-task Gaussian processes (GPs) for causal inference in panel data, applying the framework to two settings: one with a single treated unit subject to a once-and-for-all treatment and another with multiple treated units and staggered treatment adoption. Our approach models the joint evolution of outcomes for treated and control units through a GP prior that ensures exchangeability across units while allowing for flexible nonlinear trends over time. The resulting posterior predictive distribution for the untreated potential outcomes of the treated unit provides a counterfactual path, from which we derive pointwise and cumulative treatment effects, along with credible intervals to quantify uncertainty. We implement several variations of the exchangeable GP model using different kernel functions. To assess prediction accuracy, we conduct a placebo-style validation within the pre-intervention window by selecting a ``fake'' intervention date. Ultimately, this study illustrates how exchangeable GPs serve as a flexible tool for policy evaluation in panel data settings and proposes a novel approach to staggered-adoption designs with a large number of treated and control units.

stat.ME

On robust Bayesian causal inference

This paper develops a Bayesian framework for robust causal inference from longitudinal observational data. Many contemporary methods rely on structural assumptions, such as factor models, to adjust for unobserved confounding, but they can lead to biased causal estimands when mis-specified. We focus on directly estimating time--unit--specific causal effects and use generalised Bayesian inference to quantify model mis-specification and adjust for it, while retaining interpretable posterior inference. We select the learning rate~$\omega$ based on a proper scoring rule that jointly evaluates point and interval accuracy of the causal estimand, thus providing a coherent, decision-theoretic foundation for tuning~$\omega$. Simulation studies and applications to real data demonstrate improved calibration, sharpness, and robustness in estimating causal effects.

stat.ME

A Model-Based Approach to Automated Digital Twin Generation in Manufacturing

Modern manufacturing demands high flexibility and reconfigurability to adapt to dynamic production needs. Model-based Engineering (MBE) supports rapid production line design, but final reconfiguration requires simulations and validation. Digital Twins (DTs) streamline this process by enabling real-time monitoring, simulation, and reconfiguration. This paper presents a novel platform that automates DT generation and deployment using AutomationML-based factory plans. The platform closes the loop with a GAI-powered simulation scenario generator and automatic physical line reconfiguration, enhancing efficiency and adaptability in manufacturing.

eess.SY

Digital Twin based Automatic Reconfiguration of Robotic Systems in Smart Environments

Robotic systems have become integral to smart environments, enabling applications ranging from urban surveillance and automated agriculture to industrial automation. However, their effective operation in dynamic settings - such as smart cities and precision farming - is challenged by continuously evolving topographies and environmental conditions. Traditional control systems often struggle to adapt quickly, leading to inefficiencies or operational failures. To address this limitation, we propose a novel framework for autonomous and dynamic reconfiguration of robotic controllers using Digital Twin technology. Our approach leverages a virtual replica of the robot's operational environment to simulate and optimize movement trajectories in response to real-world changes. By recalculating paths and control parameters in the Digital Twin and deploying the updated code to the physical robot, our method ensures rapid and reliable adaptation without manual intervention. This work advances the integration of Digital Twins in robotics, offering a scalable solution for enhancing autonomy in smart, dynamic environments.

cs.RO

Gaussian Invariant Markov Chain Monte Carlo

We develop sampling methods, which consist of Gaussian invariant versions of random walk Metropolis (RWM), Metropolis adjusted Langevin algorithm (MALA) and second order Hessian or Manifold MALA. Unlike standard RWM and MALA, we show that Gaussian invariant sampling can lead to ergodic estimators with improved statistical efficiency. This is due to a remarkable property of Gaussian invariance that allows us to obtain exact analytical solutions to the Poisson equation for Gaussian targets. These solutions can be used to construct efficient and easy to use control variates for variance reduction of estimators under any intractable target. We demonstrate the new samplers and estimators in several examples, including high dimensional targets in latent Gaussian models where we compare against several advanced methods and obtain state-of-the-art results. We also provide theoretical results regarding geometric ergodicity, and an optimal scaling analysis that shows the dependence of the optimal acceptance rate on the Gaussianity of the target.

stat.ML

A computationally efficient framework for realistic epidemic modelling through Gaussian Markov random fields

We tackle limitations of ordinary differential equation-driven Susceptible-Infections-Removed (SIR) models and their extensions that have recently be employed for epidemic nowcasting and forecasting. In particular, we deal with challenges related to the extension of SIR-type models to account for the so-called \textit{environmental stochasticity}, i.e., external factors, such as seasonal forcing, social cycles and vaccinations that can dramatically affect outbreaks of infectious diseases. Typically, in SIR-type models environmental stochasticity is modelled through stochastic processes. However, this stochastic extension of epidemic models leads to models with large dimension that increases over time. Here we propose a Bayesian approach to build an efficient modelling and inferential framework for epidemic nowcasting and forecasting by using Gaussian Markov random fields to model the evolution of these stochastic processes over time and across population strata. Importantly, we also develop a bespoke and computationally efficient Markov chain Monte Carlo algorithm to estimate the large number of parameters and latent states of the proposed model. We test our approach on simulated data and we apply it to real data from the Covid-19 pandemic in the United Kingdom.

stat.CO

The heterogeneous causal effects of the EU's Cohesion Fund

This paper estimates the causal effect of EU cohesion policy on regional output and investment, focusing on the Cohesion Fund (CF), a comparatively understudied instrument. Departing from standard approaches such as regression discontinuity (RDD) and instrumental variables (IV), we use a recently developed causal inference method based on matrix completion within a factor model framework. This yields a new framework to evaluate the CF and to characterize the time-varying distribution of its causal effects across EU regions, along with distributional metrics relevant for policy assessment. Our results show that average treatment effects conceal substantial heterogeneity and may lead to misleading conclusions about policy effectiveness. The CF's impact is front-loaded, peaking within the first seven years after a region's initial inclusion. During this first seven-year funding cycle, the distribution of effects is right-skewed with relatively thick tails, indicating generally positive but uneven gains across regions. Effects are larger for regions that are relatively poorer at baseline, and we find a non-linear, diminishing-returns relationship: beyond a threshold, the impact declines as the ratio of CF receipts to regional gross value added (GVA) increases.

econ.GN

Real-time modelling of the SARS-CoV-2 pandemic in England 2020-2023: a challenging data integration

A central pillar of the UK's response to the SARS-CoV-2 pandemic was the provision of up-to-the moment nowcasts and short term projections to monitor current trends in transmission and associated healthcare burden. Here we present a detailed deconstruction of one of the 'real-time' models that was key contributor to this response, focussing on the model adaptations required over three pandemic years characterised by the imposition of lockdowns, mass vaccination campaigns and the emergence of new pandemic strains. The Bayesian model integrates an array of surveillance and other data sources including a novel approach to incorporating prevalence estimates from an unprecedented large-scale household survey. We present a full range of estimates of the epidemic history and the changing severity of the infection, quantify the impact of the vaccination programme and deconstruct contributing factors to the reproduction number. We further investigate the sensitivity of model-derived insights to the availability and timeliness of prevalence data, identifying its importance to the production of robust estimates.

stat.AP

Variance Reduction for Metropolis-Hastings Samplers

We introduce a general framework that constructs estimators with reduced variance for random walk Metropolis and Metropolis-adjusted Langevin algorithms. The resulting estimators require negligible computational cost and are derived in a post-process manner utilising all proposal values of the Metropolis algorithms. Variance reduction is achieved by producing control variates through the approximate solution of the Poisson equation associated with the target density of the Markov chain. The proposed method is based on approximating the target density with a Gaussian and then utilising accurate solutions of the Poisson equation for the Gaussian case. This leads to an estimator that uses two key elements: (i) a control variate from the Poisson equation that contains an intractable expectation under the proposal distribution, (ii) a second control variate to reduce the variance of a Monte Carlo estimate of this latter intractable expectation. Simulated data examples are used to illustrate the impressive variance reduction achieved in the Gaussian target case and the corresponding effect when target Gaussianity assumption is violated. Real data examples on Bayesian logistic regression and stochastic volatility models verify that considerable variance reduction is achieved with negligible extra computational cost.

stat.ME

Evaluating the impact of local tracing partnerships on the performance of contact tracing for COVID-19 in England

Assessing the impact of an intervention using time-series observational data on multiple units and outcomes is a frequent problem in many fields of scientific research. In this paper, we present a novel method to estimate intervention effects in such a setting by generalising existing approaches based on the factor analysis model and developing a Bayesian algorithm for inference. Our method is one of the few that can simultaneously: deal with outcomes of mixed type (continuous, binomial, count); increase efficiency in the estimates of the causal effects by jointly modelling multiple outcomes affected by the intervention; easily provide uncertainty quantification for all causal estimands of interest. We use the proposed approach to evaluate the impact that local tracing partnerships (LTP) had on the effectiveness of England's Test and Trace (TT) programme for COVID-19. Our analyses suggest that, overall, LTPs had a small positive impact on TT. However, there is considerable heterogeneity in the estimates of the causal effects over units and time.

stat.AP

A network and machine learning approach to detect Value Added Tax fraud

Value Added Tax (VAT) fraud erodes public revenue and puts legitimate businesses at a disadvantaged position thereby impacting inequality. Identifying and combating VAT fraud before it occurs is therefore important for welfare. This paper proposes flexible machine learning algorithms which detect fraudulent transactions, utilising the information provided by the complex VAT network structure of a large dimension. VAT fraud detection is implemented through a combination of a suitably constructed Laplacian matrix with classification algorithms that rely on scalable machine learning techniques. The method is implemented on the universe of Bulgarian VAT data and detects around 50 percent of the VAT fraud, outperforming well-known techniques that ignore the information provided by the network of VAT transactions. Importantly, the proposed methods are automated, and can be implemented following the taxpayers submission of their VAT returns. This allows tax revenue authorities to prevent large losses of tax revenues through performing early identification of fraud between business-to-business transactions within the VAT system.

physics.soc-ph

Bayesian Variable Selection for Gaussian copula regression models

We develop a novel Bayesian method to select important predictors in regression models with multiple responses of diverse types. A sparse Gaussian copula regression model is used to account for the multivariate dependencies between any combination of discrete and/or continuous responses and their association with a set of predictors. We utilize the parameter expansion for data augmentation strategy to construct a Markov chain Monte Carlo algorithm for the estimation of the parameters and the latent variables of the model. Based on a centered parametrization of the Gaussian latent variables, we design a fixed-dimensional proposal distribution to update jointly the latent binary vectors of important predictors and the corresponding non-zero regression coefficients. For Gaussian responses and for outcomes that can be modeled as a dependent version of a Gaussian response, this proposal leads to a Metropolis-Hastings step that allows an efficient exploration of the predictors' model space. The proposed strategy is tested on simulated data and applied to real data sets in which the responses consist of low-intensity counts, binary, ordinal and continuous variables.

stat.ME

Bayesian prediction of jumps in large panels of time series data

We take a new look at the problem of disentangling the volatility and jumps processes of daily stock returns. We first provide a computational framework for the univariate stochastic volatility model with Poisson-driven jumps that offers a competitive inference alternative to the existing tools. This methodology is then extended to a large set of stocks for which we assume that their unobserved jump intensities co-evolve in time through a dynamic factor model. To evaluate the proposed modelling approach we conduct out-of-sample forecasts and we compare the posterior predictive distributions obtained from the different models. We provide evidence that joint modelling of jumps improves the predictive ability of the stochastic volatility models.

q-fin.ST

Bayesian forecasting of mortality rates using latent Gaussian models

We provide forecasts for mortality rates by using two different approaches. First we employ dynamic non-linear logistic models based on Heligman-Pollard formula. Second, we assume that the dynamics of the mortality rates can be modelled through a Gaussian Markov random field. We use efficient Bayesian methods to estimate the parameters and the latent states of the proposed models. Both methodologies are tested with past data and are used to forecast mortality rates both for large (UK and Wales) and small (New Zealand) populations up to 21 years ahead. We demonstrate that predictions for individual survivor functions and other posterior summaries of demographic and actuarial interest are readily obtained. Our results are compared with other competing forecasting methods.

stat.AP