SearcharxivSearch

arXiv subjects

Konstantinos Kalogeropoulos

Publications and source records attributed to Konstantinos Kalogeropoulos.

At least 19 recordsLinked to original sources

Evaluating the Impact of COVID-19 Vaccination in the United Kingdom: A Gaussian Process Approach

The rapid rollout of COVID-19 vaccines in the United Kingdom in early 2021 differed markedly from that of many other European countries, providing a natural setting to assess the impact of vaccination speed on public health outcomes. We evaluate the impact of the accelerated UK vaccination rollout and associated policy transition on COVID-19 mortality and transmission dynamics by constructing a probabilistic reference trajectory for the UK under a slower vaccination and reopening trajectory. The proposed framework combines ideas from interrupted time series analysis and synthetic control methods with flexible probabilistic modelling based on multi-output Gaussian processes. These models capture non-linear and heterogeneous dependence structures across countries and over time, while providing uncertainty quantification through predictive distributions. A central feature of the methodology is a design-consistent validation strategy based on predictive performance in held-out pre-intervention periods, which is used both to guide model specification and to assess the plausibility of the reconstructed reference trajectory. The empirical results indicate a substantial reduction in COVID-19 mortality associated with the accelerated vaccination-policy transition, with little evidence of an effect on transmission rates. Generally, the framework illustrates how flexible probabilistic models and predictive validation can support causal and policy evaluation in complex time series settings.

stat.ME

Bayesian Benefit-Risk Assessment with Dependent Outcomes via Latent Factor Models

Approving and assessing new drugs is complex because multiple criteria must be considered simultaneously. A common approach is benefit-risk analysis, often conducted within a Bayesian framework to account for uncertainty and combine data with expert judgement, typically through multi-criteria decision analysis (MCDA) scores. This requires models that accommodate mixed and potentially correlated outcomes; latent factor models provide a natural framework. We develop a coherent Bayesian framework for benefit-risk analysis that addresses these challenges and supports sequential decision-making. We extend structured factor models to mixed outcomes and introduce a principled approach for selecting among competing specifications that combines model fit with out-of-sample predictive performance. We then develop a sequential estimation framework that updates MCDA scores as new data become available, allowing treatment comparisons to evolve over time. This supports early stopping when conclusions are clear and permits dynamic treatment allocation aligned with study objectives. To make this feasible, we develop tailored sequential Monte Carlo methods adapted to the model structure. The methodology is illustrated using data on patients with type II diabetes treated with Metformin, Rosiglitazone, and their combination.

stat.ME

Dynamic Inference in Term Structure Models with Unspanned Latent Risks

We propose a parsimonious class of arbitrage-free, yields-only dynamic term structure models (DTSMs) with unspanned latent risks. To enable sequential estimation and forecasting, we develop a Sequential Monte Carlo framework that combines particle learning for static parameters with Kalman filter updates for latent states, yielding joint posterior inference and predictive distributions that account for both parameter and state uncertainty. We use this framework to assess the out-of-sample statistical and economic value of bond return predictability from the perspective of a Bayesian investor. Empirically, we find that unspanned latent factors contain predictive information beyond that embedded in the yield curve, improving out-of-sample forecasting performance relative to standard benchmark models. These gains translate into economically meaningful utility improvements across a range of portfolio settings. Finally, we show that the hidden component of the slope-related risk factor is countercyclical and associated with real economic activity, suggesting that the latent factors capture economically relevant variation not directly reflected in yields.

stat.AP

Semi-Markov Models with Particle-Based Bayesian Inference for Epidemics

The COVID-19 pandemic has been characterised by multiple waves of transmission driven by interventions and emerging variants, challenging epidemic models that assume gradually evolving transmission dynamics. We propose a class of state-space models in which the transmission rate evolves through persistent regimes of random duration, governed by a semi-Markov process. This formulation yields an interpretable representation of sustained transmission phases and retains a parsimonious parameterisation. Particle-based Bayesian methods are well established for standard state-space models, but their use in semi-Markov settings has received comparatively limited attention. In epidemic applications, inference is further complicated by differential equation-driven latent dynamics and observation models defined through functionals of the latent process. We develop an inferential framework that accommodates these features, combining particle-based state updates with gradient-based parameter updates and enabling batch and sequential inference via particle and sequential Monte Carlo. We apply the proposed methodology to COVID-19 data from the United Kingdom and show that combining reported cases and deaths leads to more precise and stable inference compared to using deaths alone. These results illustrate the practical value of semi-Markov transmission models for epidemic analysis under complex observation schemes.

stat.AP

Exchangeable Gaussian Processes for Staggered-Adoption Policy Evaluation

We study the use of exchangeable multi-task Gaussian processes (GPs) for causal inference in panel data, applying the framework to two settings: one with a single treated unit subject to a once-and-for-all treatment and another with multiple treated units and staggered treatment adoption. Our approach models the joint evolution of outcomes for treated and control units through a GP prior that ensures exchangeability across units while allowing for flexible nonlinear trends over time. The resulting posterior predictive distribution for the untreated potential outcomes of the treated unit provides a counterfactual path, from which we derive pointwise and cumulative treatment effects, along with credible intervals to quantify uncertainty. We implement several variations of the exchangeable GP model using different kernel functions. To assess prediction accuracy, we conduct a placebo-style validation within the pre-intervention window by selecting a ``fake'' intervention date. Ultimately, this study illustrates how exchangeable GPs serve as a flexible tool for policy evaluation in panel data settings and proposes a novel approach to staggered-adoption designs with a large number of treated and control units.

stat.ME

Exchangeable Gaussian Processes with application to epidemics

We develop a Bayesian non-parametric framework based on multi-task Gaussian processes, appropriate for temporal shrinkage. We focus on a particular class of dynamic hierarchical models to obtain evidence-based knowledge of infectious disease burden. These models induce a parsimonious way to capture cross-dependence between groups while retaining a natural interpretation based on an underlying mean process, itself expressed as a Gaussian process. We analyse distinct types of outbreak data from recent epidemics and find that the proposed models result in improved predictive ability against competing alternatives.

stat.ME

De novo peptide sequencing rescoring and FDR estimation with Winnow

Machine learning has markedly advanced de novo peptide sequencing (DNS) for mass spectrometry-based proteomics. DNS tools offer a reliable way to identify peptides without relying on reference databases, extending proteomic analysis and unlocking applications into less-charted regions of the proteome. However, they still face a key limitation. DNS tools lack principled methods for estimating false discovery rates (FDR) and instead rely on model-specific confidence scores that are often miscalibrated. This limits trust in results, hinders cross-model comparisons and reduces validation success. Here we present Winnow, a model-agnostic framework for estimating FDR from calibrated DNS outputs. Winnow maps raw model scores to calibrated confidences using a neural network trained on peptide-spectrum match (PSM)-derived features. From these calibrated scores, Winnow computes PSM-specific error metrics and an experiment-wide FDR estimate using a novel decoy-free FDR estimator. It supports both zero-shot and dataset-specific calibration, enabling flexible application via direct inference, fine-tuning, or training a custom model. We demonstrate that, when applied to InstaNovo predictions, Winnow's calibrator improves recall at fixed FDR thresholds, and its FDR estimator tracks true error rates when benchmarked against reference proteomes and database search. Winnow ensures accurate FDR control across datasets, helping unlock the full potential of DNS.

q-bio.QM

Bayesian analysis of diffusion-driven multi-type epidemic models with application to COVID-19

We consider a flexible Bayesian evidence synthesis approach to model the age-specific transmission dynamics of COVID-19 based on daily mortality counts. The temporal evolution of transmission rates in populations containing multiple types of individuals is reconstructed via an appropriate dimension-reduction formulation driven by independent diffusion processes. A suitably tailored compartmental model is used to learn the latent counts of infection, accounting for fluctuations in transmission influenced by public health interventions and changes in human behaviour. The model is fitted to freely available COVID-19 data sources from the UK, Greece, and Austria and validated using a large-scale prevalence survey in England. In particular, we demonstrate how model expansion can facilitate evidence reconciliation at a latent level. The code implementing this work is made freely available via the Bernadette R package.

stat.CO

A modelling framework for the analysis of the SARS-CoV2 transmission dynamics

Despite the progress in medical data collection the actual burden of SARS-CoV-2 remains unknown due to under-ascertainment of cases. This was apparent in the acute phase of the pandemic and the use of reported deaths has been pointed out as a more reliable source of information, likely less prone to under-reporting. Since daily deaths occur from past infections weighted by their probability of death, one may infer the total number of infections accounting for their age distribution, using the data on reported deaths. We adopt this framework and assume that the dynamics generating the total number of infections can be described by a continuous time transmission model expressed through a system of non-linear ordinary differential equations where the transmission rate is modelled as a diffusion process allowing to reveal both the effect of control strategies and the changes in individuals behavior. We develop this flexible Bayesian tool in Stan and study 3 pairs of European countries, estimating the time-varying reproduction number($R_t$) as well as the true cumulative number of infected individuals. As we estimate the true number of infections we offer a more accurate estimate of $R_t$. We also provide an estimate of the daily reporting ratio and discuss the effects of changes in mobility and testing on the inferred quantities.

stat.AP

Dynamic Term Structure Models with Nonlinearities using Gaussian Processes

The importance of unspanned macroeconomic variables for Dynamic Term Structure Models has been intensively discussed in the literature. To our best knowledge the earlier studies considered only linear interactions between the economy and the real-world dynamics of interest rates in DTSMs. We propose a generalized modelling setup for Gaussian DTSMs which allows for unspanned nonlinear associations between the two and we exploit it in forecasting. Specifically, we construct a custom sequential Monte Carlo estimation and forecasting scheme where we introduce Gaussian Process priors to model nonlinearities. Sequential scheme we propose can also be used with dynamic portfolio optimization to assess the potential of generated economic value to investors. The methodology is presented using US Treasury data and selected macroeconomic indices. Namely, we look at core inflation and real economic activity. We contrast the results obtained from the nonlinear model with those stemming from an application of a linear model. Unlike for real economic activity, in case of core inflation we find that, compared to linear models, application of nonlinear models leads to statistically significant gains in economic value across considered maturities.

stat.AP

Sequential Bayesian Learning for Hidden Semi-Markov Models

In this paper, we explore the class of the Hidden Semi-Markov Model (HSMM), a flexible extension of the popular Hidden Markov Model (HMM) that allows the underlying stochastic process to be a semi-Markov chain. HSMMs are typically used less frequently than their basic HMM counterpart due to the increased computational challenges when evaluating the likelihood function. Moreover, while both models are sequential in nature, parameter estimation is mainly conducted via batch estimation methods. Thus, a major motivation of this paper is to provide methods to estimate HSMMs (1) in a computationally feasible time, (2) in an exact manner, i.e. only subject to Monte Carlo error, and (3) in a sequential setting. We provide and verify an efficient computational scheme for Bayesian parameter estimation on HSMMs. Additionally, we explore the performance of HSMMs on the VIX time series using Autoregressive (AR) models with hidden semi-Markov states and demonstrate how this algorithm can be used for regime switching, model selection and clustering purposes.

stat.AP

Sequential Learning and Economic Benefits from Dynamic Term Structure Models

We explore the statistical and economic importance of restrictions on the dynamics of risk compensation from the perspective of a real-time Bayesian learner who predicts bond excess returns using dynamic term structure models (DTSMs). The question on whether potential statistical predictability offered by such models can generate economically significant portfolio benefits out-of-sample, is revisited while imposing restrictions on their risk premia parameters. To address this question, we propose a methodological framework that successfully handles sequential model search and parameter estimation over the restriction space in real time, allowing investors to revise their beliefs when new information arrives, thus informing their asset allocation and maximising their expected utility. Empirical results reinforce the argument of sparsity in the market price of risk specification since we find strong evidence of out-of-sample predictability only for those models that allow for level risk to be priced and, additionally, only one or two of these risk premia parameters to be different than zero. Most importantly, such statistical evidence is turned into economically significant utility gains, across prediction horizons, different time periods and portfolio specifications. In addition to identifying successful DTSMs, the sequential version of the stochastic search variable selection (SSVS) scheme developed can be applied on its own and also offer useful diagnostics monitoring key quantities over time. Connections with predictive regressions are also provided.

stat.AP

Model Assessment for a Generalised Bayesian Structural Equation Model

The paper proposes a novel model assessment paradigm aiming to address shortcoming of posterior predictive $p-$values, which provide the default metric of fit for Bayesian structural equation modelling (BSEM). The model framework of the paper focuses on the approximate zero approach, according to which parameters that would before set to zero (e.g. factor loadings) are now formulated to be approximate zero via informative priors (Muthen and Asparouhov, 2012). The introduced model assessment procedure monitors the out-of-sample predictive performance of the fitted model, and together with a list of guidelines we provide, one can investigate whether the hypothesised model is supported by the data. We incorporate scoring rules and cross-validation to supplement existing model assessment metrics for Bayesian SEM. The proposed tools can be applied to models for both categorical and continuous data. The modelling of categorical and non-normally distributed continuous data is facilitated with the introduction of an item-individual random effect that can also be used for outlier detection. We study the performance of the proposed methodology via simulations. The factor model for continuous and binary data is fitted to data on the `Big-5' personality scale and the Fagerstrom test for nicotine dependence respectively.

stat.ME

Sequential Bayesian Inference for Factor Analysis

We develop an efficient Bayesian sequential inference framework for factor analysis models observed via various data types, such as continuous, binary and ordinal data. In the continuous data case, where it is possible to marginalise over the latent factors, the proposed methodology tailors the Iterated Batch Importance Sampling (IBIS) of Chopin (2002) to handle such models and we incorporate Hamiltonian Markov Chain Monte Carlo. For binary and ordinal data, we develop an efficient IBIS scheme to handle the parameter and latent factors, combining with Laplace or Variational Bayes approximations. The methodology can be used in the context of sequential hypothesis testing via Bayes factors, which are known to have advantages over traditional null hypothesis testing. Moreover, the developed sequential framework offers multiple benefits even in non-sequential cases, by providing posterior distribution, model evidence and scoring rules (under the prequential framework) in one go, and by offering a more robust alternative computational scheme to Markov Chain Monte Carlo that can be useful in problematic target distributions.

stat.ME

Bayesian Inference for partially observed SDEs Driven by Fractional Brownian Motion

We consider continuous-time diffusion models driven by fractional Brownian motion. Observations are assumed to possess a non-trivial likelihood given the latent path. Due to the non-Markovianity and high-dimensionality of the latent paths, estimating posterior expectations is a computationally challenging undertaking. We present a reparameterization framework based on the Davies and Harte method for sampling stationary Gaussian processes and use this framework to construct a Markov chain Monte Carlo algorithm that allows computationally efficient Bayesian inference. The Markov chain Monte Carlo algorithm is based on a version of hybrid Monte Carlo that delivers increased efficiency when applied on the high-dimensional latent variables arising in this context. We specify the methodology on a stochastic volatility model allowing for memory in the volatility increments through a fractional specification. The methodology is illustrated on simulated data and on the S&P500/VIX time series and is shown to be effective. Contrary to a long range dependence attribute of such models often assumed in the literature, with Hurst parameter larger than 1/2, the posterior distribution favours values smaller than 1/2, pointing towards medium range dependence.

stat.ME

A Bayesian approach to estimate changes in condom use from limited HIV prevalence data

Evaluation of HIV large scale interventions programme is becoming increasingly important, but impact estimates frequently hinge on knowledge of changes in behaviour such as the frequency of condom use (CU) over time, or other self-reported behaviour changes, for which we generally have limited or potentially biased data. We employ a Bayesian inference methodology that incorporates a dynamic HIV transmission dynamics model to estimate CU time trends from HIV prevalence data. Estimation is implemented via particle Markov Chain Monte Carlo methods, applied for the first time in this context. The preliminary choice of the formulation for the time varying parameter reflecting the proportion of CU is critical in the context studied, due to the very limited amount of CU and HIV data available We consider various novel formulations to explore the trajectory of CU in time, based on diffusion-driven trajectories and smooth sigmoid curves. Extensive series of numerical simulations indicate that informative results can be obtained regarding the amplitude of the increase in CU during an intervention, with good levels of sensitivity and specificity performance in effectively detecting changes. The application of this method to a real life problem illustrates how it can help evaluate HIV intervention from few observational studies and suggests that these methods can potentially be applied in many different contexts.

stat.AP

Advanced MCMC Methods for Sampling on Diffusion Pathspace

The need to calibrate increasingly complex statistical models requires a persistent effort for further advances on available, computationally intensive Monte Carlo methods. We study here an advanced version of familiar Markov Chain Monte Carlo (MCMC) algorithms that sample from target distributions defined as change of measures from Gaussian laws on general Hilbert spaces. Such a model structure arises in several contexts: we focus here at the important class of statistical models driven by diffusion paths whence the Wiener process constitutes the reference Gaussian law. Particular emphasis is given on advanced Hybrid Monte-Carlo (HMC) which makes large, derivative-driven steps in the state space (in contrast with local-move Random-walk-type algorithms) with analytical and experimental results. We illustrate it's computational advantages in various diffusion processes and observation regimes; examples include stochastic volatility and latent survival models. In contrast with their standard MCMC counterparts, the advanced versions have mesh-free mixing times, as these will not deteriorate upon refinement of the approximation of the inherently infinite-dimensional diffusion paths by finite-dimensional ones used in practice when applying the algorithms on a computer.

stat.ME

Capturing the time-varying drivers of an epidemic using stochastic dynamical systems

Epidemics are often modelled using non-linear dynamical systems observed through partial and noisy data. In this paper, we consider stochastic extensions in order to capture unknown influences (changing behaviors, public interventions, seasonal effects etc). These models assign diffusion processes to the time-varying parameters, and our inferential procedure is based on a suitably adjusted adaptive particle MCMC algorithm. The performance of the proposed computational methods is validated on simulated data and the adopted model is applied to the 2009 H1N1 pandemic in England. In addition to estimating the effective contact rate trajectories, the methodology is applied in real time to provide evidence in related public health decisions. Diffusion driven SEIR-type models with age structure are also introduced.

stat.AP