SearcharxivSearch

arXiv subjects

Luiz Max Carvalho

Publications and source records attributed to Luiz Max Carvalho.

15 recordsLinked to original sources

From priors to performance: enhancing statistical efficiency with Bayesian dynamic borrowing

One of the main advantages of the Bayesian approach to statistical inference is the flexibility in incorporating information from various sources, from expert opinion to historical data. Whereas the literature on Bayesian dynamic borrowing is rich, practical guidance for how to implement such methods is comparatively sparse. We thoroughly review the most common dynamic borrowing approaches for external data, including how to incorporate multiple historical data sets. We then show how these techniques can be used in practice, using publicly available software, to perform important statistical tasks such as model selection, average treatment effect estimation and prior sensitivity analysis. The aim is to make analysts more confident in their use of Bayesian methods to incorporate external data. All code is made publicly available in a companion GitHub repository (https://github.com/EzequielEBS/hdbayes-tutorial).

stat.ME

An adaptive time-tree transition kernel for Bayesian phylogenetic inference

Bayesian phylogenetic and phylodynamic analyses can be very time-consuming, owing to the combination of complex models that are used to estimate key parameters from increasingly large genomic data sets and their associated metadata. The use of high-performance computer hardware can -- to a certain extent -- alleviate the computational burden and markedly decrease the time to results. Still, even converging to the posterior can be a lengthy endeavour, with the burn-in aspect of such analyses potentially taking days or even weeks for large data sets. One of the key aspects that hampers performance in Bayesian phylogenetic inference is the efficiency with which tree topology proposals explore tree space. We here propose a novel adaptive tree transition kernel, which we call `subTreeLeap' (STL), which involves modifying the phylogeny by walking along patristic distance paths in the tree according to an adaptable radius parameter. STL is a general proposal, which can be used with contemporaneous or time-calibrated sequence data, being particularly suited to the latter due to respecting temporal precedence constraints. We carefully assess its impact on convergence and statistical mixing of the exploration of posterior tree space, by comparison to replicate ``golden runs'' obtained from lengthy analyses of empirical data under standard tree transition kernels. We find that STL successfully explores the same posterior tree space as standard kernels, but often does so in a more efficient manner. We discuss limitations as well as future potential improvements to STL that could substantially increase the speed at which Bayesian phylogenetic inferences are obtained.

q-bio.PE

Mapping the landscape of mathematical models for antimicrobial resistance: a scoping review

Background: Antimicrobial resistance (AMR) is a major global public health problem, contributing to an estimated 4.95 million deaths in 2019 and projected to cause up to 10 million deaths annually and 100 trillion dollars in cumulative economic losses by 2050. Its emergence and spread result from complex biological, ecological, and socioeconomic interactions. Mathematical modelling is a key tool to study AMR dynamics, yet the literature remains fragmented and methodologically limited. This review synthesizes recent mathematical modelling studies to identify trends, biases, and research gaps. Methods: A scoping review following PRISMA-ScR guidelines was conducted. PubMed, Web of Science, and Scopus were searched for studies published between 2019 and 2024 that developed mathematical models of AMR. After screening and duplicate removal, 36 studies were included. Data were extracted using a structured framework covering model context, construction and parameters, and outputs and validation. Results: Most studies relied on deterministic ordinary differential equation (ODE) models and focused on bacterial resistance in human hosts, with only one adopting a One Health perspective. Conjugation and mutation were the most commonly modelled resistance mechanisms, whereas transduction, transformation, host immunity, spatial heterogeneity, and environmental components were rarely included. Few studies incorporated economic impacts, and a strong geographic bias toward high-income countries was observed. Conclusion: Mathematical modelling of AMR is an active field but characterized by limited methodological diversity. While deterministic ODE models have advanced understanding of AMR dynamics, future work should integrate stochasticity, spatial structure, ecological interactions, and One Health perspectives, as well as economic and social variables, to inform global strategies to mitigate AMR.

q-bio.PE

Adaptive truncation of infinite sums: applications to Statistics

It is often the case in Statistics that one needs to compute sums of infinite series, especially in marginalising over discrete latent variables. This has become more relevant with the popularization of gradient-based techniques (e.g. Hamiltonian Monte Carlo) in the Bayesian inference context, for which discrete latent variables are hard or impossible to deal with. For many commonly used infinite series, custom algorithms have been developed which exploit specific features of each problem. General techniques, suitable for a large class of problems with limited input from the user are less established. We employ basic results from the theory of infinite series to investigate general, problem-agnostic algorithms to truncate infinite sums within an arbitrary tolerance $\varepsilon > 0$ and provide robust computational implementations with provable guarantees. We compare three tentative solutions to estimating the infinite sum of interest: (i) a "naive" approach that sums terms until the terms are below the threshold $\varepsilon$; (ii) a `bounding pair' strategy based on trapping the true value between two partial sums; and (iii) a `batch' strategy that computes the partial sums in regular intervals and stops when their difference is less than $\varepsilon$. We show under which conditions each strategy guarantees the truncated sum is within the required tolerance and compare the error achieved by each approach, as well as the number of function evaluations necessary for each one. A detailed discussion of numerical issues in practical implementations is also provided. The paper provides some theoretical discussion of a variety of statistical applications, including raw and factorial moments and count models with observation error. Finally, detailed illustrations in the form noisy MCMC for Bayesian inference and maximum marginal likelihood estimation are presented.

stat.ME

Long-term predictive models for mosquito borne diseases: a narrative review

In face of climate change and increasing urbanization, the predictive mosquito-borne diseases (MBD) transmission models require constant updates. Thus, is urgent to comprehend the driving forces of this non stationary behavior, observed through spatial and incidence expansion. We observed that temperature is a critical driver in predictive models for MBD transmission, also being consistently used in multiple reviewed papers with considerable incidence predictive capacity. Rainfall, however, have more subtle importance as moderate precipitation creates breeding sites for mosquitoes, but excessive rainfall can reduce larvae populations. We highlight the frequent use of mechanistic models, particularly those that integrate temperature-dependent biological parameters of disease transmission in incidence proxies as the Vectorial Capacity (VC) and temperature-based basic reproduction number $R_0(t)$, for example. These models show the importance of climate variables, but the socio-demographic factors are often not considered. This gap is a significant opportunity for future research to incorporate socio-demographic data into long-term predictive models for more comprehensive and reliable forecasts. With this survey, we outline the most promising paths to be followed by long-term MBD transmission research and highlighting the potential facing challenges. Thus, we offer a valuable foundation for enhancing disease forecasting models and supporting more effective public health interventions, specially in the long term.

q-bio.QM

Mosqlimate: a platform to providing automatable access to data and forecasting models for arbovirus disease

Dengue is a climate-sensitive mosquito-borne disease with a complex transmission dynamic. Data related to climate, environmental and sociodemographic characteristics of the target population are important for project scenarios. Different datasets and methodologies have been applied to build complex models for dengue forecast, stressing the need to evaluate these models and their relative accuracy grounded on a reproducible methodology. The goal of this work is to describe and present Mosqlimate, a web-based platform composed by a dashboard, a data store, model and rediction registries and support for a community of practice in arbovirus forecasting. Multiple API endpoints give access to data for development, open registration of predictive models from different approaches and sharing of predictive models for arboviruses incidence, facilitating interaction between modellers and allowing for proper comparison of the performance of different registered models, by means of probabilistic scores. Epidemiological, entomological, climatic and sociodemographic datasets related to arboviruses in Brazil, are freely available for download, alongside full documentation.

stat.AP

On the importance of assessing topological convergence in Bayesian phylogenetic inference

Modern phylogenetics research is often performed within a Bayesian framework, using sampling algorithms such as Markov chain Monte Carlo (MCMC) to approximate the posterior distribution. These algorithms require careful evaluation of the quality of the generated samples. Within the field of phylogenetics, one frequently adopted diagnostic approach is to evaluate the effective sample size (ESS) and to investigate trace graphs of the sampled parameters. A major limitation of these approaches is that they are developed for continuous parameters and therefore incompatible with a crucial parameter in these inferences: the tree topology. Several recent advancements have aimed at extending these diagnostics to topological space. In this reflection paper, we present two case studies - one on Ebola virus and one on HIV - illustrating how these topological diagnostics can contain information not found in standard diagnostics, and how decisions regarding which of these diagnostics to compute can impact inferences regarding MCMC convergence and mixing. Our results show the importance of running multiple replicate analyses and of carefully assessing topological convergence using the output of these replicate analyses. To this end, we illustrate different ways of assessing and visualizing the topological convergence of these replicates. Given the major importance of detecting convergence and mixing issues in Bayesian phylogenetic analyses, the lack of a unified approach to this problem warrants further action, especially now that additional tools are becoming available to researchers.

q-bio.PE

Embarrassingly Parallel GFlowNets

GFlowNets are a promising alternative to MCMC sampling for discrete compositional random variables. Training GFlowNets requires repeated evaluations of the unnormalized target distribution or reward function. However, for large-scale posterior sampling, this may be prohibitive since it incurs traversing the data several times. Moreover, if the data are distributed across clients, employing standard GFlowNets leads to intensive client-server communication. To alleviate both these issues, we propose embarrassingly parallel GFlowNet (EP-GFlowNet). EP-GFlowNet is a provably correct divide-and-conquer method to sample from product distributions of the form $R(\cdot) \propto R_1(\cdot) ... R_N(\cdot)$ -- e.g., in parallel or federated Bayes, where each $R_n$ is a local posterior defined on a data partition. First, in parallel, we train a local GFlowNet targeting each $R_n$ and send the resulting models to the server. Then, the server learns a global GFlowNet by enforcing our newly proposed \emph{aggregating balance} condition, requiring a single communication step. Importantly, EP-GFlowNets can also be applied to multi-objective optimization and model reuse. Our experiments illustrate the EP-GFlowNets's effectiveness on many tasks, including parallel Bayesian phylogenetics, multi-objective multiset, sequence generation, and federated Bayesian structure learning.

cs.LG

Locking and Quacking: Stacking Bayesian model predictions by log-pooling and superposition

Combining predictions from different models is a central problem in Bayesian inference and machine learning more broadly. Currently, these predictive distributions are almost exclusively combined using linear mixtures such as Bayesian model averaging, Bayesian stacking, and mixture of experts. Such linear mixtures impose idiosyncrasies that might be undesirable for some applications, such as multi-modality. While there exist alternative strategies (e.g. geometric bridge or superposition), optimising their parameters usually involves computing an intractable normalising constant repeatedly. We present two novel Bayesian model combination tools. These are generalisations of model stacking, but combine posterior densities by log-linear pooling (locking) and quantum superposition (quacking). To optimise model weights while avoiding the burden of normalising constants, we investigate the Hyvarinen score of the combined posterior predictions. We demonstrate locking with an illustrative example and discuss its practical application with importance sampling.

stat.ML

Bivariate beta distribution: parameter inference and diagnostics

Correlated proportions appear in many real-world applications and present a unique challenge in terms of finding an appropriate probabilistic model due to their constrained nature. The bivariate beta is a natural extension of the well-known beta distribution to the space of correlated quantities on $[0, 1]^2$. Its construction is not unique, however. Over the years, many bivariate beta distributions have been proposed, ranging from three to eight or more parameters, and for which the joint density and distribution moments vary in terms of mathematical tractability. In this paper, we investigate the construction proposed by Olkin & Trikalinos (2015), which strikes a balance between parameter-richness and tractability. We provide classical (frequentist) and Bayesian approaches to estimation in the form of method-of-moments and latent variable/data augmentation coupled with Hamiltonian Monte Carlo, respectively. The elicitation of bivariate beta as a prior distribution is also discussed. The development of diagnostics for checking model fit and adequacy is explored in depth with the aid of Monte Carlo experiments under both well-specified and misspecified data-generating settings. Keywords: Bayesian estimation; bivariate beta; correlated proportions; diagnostics; method of moments.

stat.ME

Beyond the shortest path: the path length index as a distribution

The traditional complex network approach considers only the shortest paths from one node to another, not taking into account several other possible paths. This limitation is significant, for example, in urban mobility studies. In this short report, as the first steps, we present an exhaustive approach to address that problem and show we can go beyond the shortest path, but we do not need to go so far: we present an interactive procedure and an early stop possibility. After presenting some fundamental concepts in graph theory, we presented an analytical solution for the problem of counting the number of possible paths between two nodes in complete graphs, and a depth-limited approach to get all possible paths between each pair of nodes in a general graph (an NP-hard problem). We do not collapse the distribution of path lengths between a pair of nodes into a scalar number, we look at the distribution itself - taking all paths up to a pre-defined path length (considering a truncated distribution), and show the impact of that approach on the most straightforward distance-based graph index: the walk/path length.

cs.DM

On the normalized power prior

The power prior is a popular tool for constructing informative prior distributions based on historical data. The method consists of raising the likelihood to a discounting factor in order to control the amount of information borrowed from the historical data. It is customary to perform a sensitivity analysis reporting results for a range of values of the discounting factor. However, one often wishes to assign it a prior distribution and estimate it jointly with the parameters, which in turn necessitates the computation of a normalising constant. In this paper we are concerned with how to recycle computations from a sensitivity analysis in order to approximately sample from joint posterior of the parameters and the discounting factor. We first show a few important properties of the normalising constant and then use these results to motivate a bisection-type algorithm for computing it on a fixed budget of evaluations. We give a large array of illustrations and discuss cases where the normalising constant is known in closed-form and where it is not. We show that the proposed method produces approximate posteriors that are very close to the exact distributions when those are available and also produces posteriors that cover the data-generating parameters with higher probability in the intractable case. Our results show that proper inclusion the normalising constant is crucial to the correct quantification of uncertainty and that the proposed method is an accurate and easy to implement technique to include this normalisation, being applicable to a large class of models. Key-words: Doubly-intractable; elicitation; historical data; normalisation; power prior; sensitivity analysis.

stat.AP

Spatio-temporal Dynamics of Foot-and-Mouth Disease Virus in South America

Although foot-and-mouth disease virus (FMDV) incidence has decreased in South America over the last years, the pathogen still circulates in the region and the risk of re-emergence in previously FMDV-free areas is a veterinary public health concern. In this paper we merge environmental, epidemiological and genetic data to reconstruct spatiotemporal patterns and determinants of FMDV serotypes A and O dispersal in South America. Our dating analysis suggests that serotype A emerged in South America around 1930, while serotype O emerged around 1990. The rate of evolution for serotype A was significantly higher compared to serotype O. Phylogeographic inference identified two well-connected sub networks of viral flow, one including Venezuela, Colombia and Ecuador; another including Brazil, Uruguay and Argentina. The spread of serotype A was best described by geographic distances, while trade of live cattle was the predictor that best explained serotype O spread. Our findings show that the two serotypes have different underlying evolutionary and spatial dynamics and may pose different threats to control programmes. Key-words: Phylogeography, foot-and-mouth disease virus, South America, animal trade.

q-bio.PE

Estimating the Attack Ratio of Dengue Epidemics under Time-varying Force of Infection using Aggregated Notification Data

Quantifying the attack ratio of disease is key to epidemiological inference and Public Health planning. For multi-serotype pathogens, however, different levels of serotype-specific immunity make it difficult to assess the population at risk. In this paper we propose a Bayesian method for estimation of the attack ratio of an epidemic and the initial fraction of susceptibles using aggregated incidence data. We derive the probability distribution of the effective reproductive number, R t , and use MCMC to obtain posterior distributions of the parameters of a single-strain SIR transmission model with time-varying force of infection. Our method is showcased in a data set consisting of 18 years of dengue incidence in the city of Rio de Janeiro, Brazil. We demonstrate that it is possible to learn about the initial fraction of susceptibles and the attack ratio even in the absence of serotype specific data. On the other hand, the information provided by this approach is limited, stressing the need for detailed serological surveys to characterise the distribution of serotype-specific immunity in the population.

q-bio.PE

piBUSS: a parallel BEAST/BEAGLE utility for sequence simulation under complex evolutionary scenarios

Background: Simulated nucleotide or amino acid sequences are frequently used to assess the performance of phylogenetic reconstruction methods. BEAST, a Bayesian statistical framework that focuses on reconstructing time-calibrated molecular evolutionary processes, supports a wide array of evolutionary models, but lacked matching machinery for simulation of character evolution along phylogenies. Results: We present a flexible Monte Carlo simulation tool, called piBUSS, that employs the BEAGLE high performance library for phylogenetic computations within BEAST to rapidly generate large sequence alignments under complex evolutionary models. piBUSS sports a user-friendly graphical user interface (GUI) that allows combining a rich array of models across an arbitrary number of partitions. A command-line interface mirrors the options available through the GUI and facilitates scripting in large-scale simulation studies. Analogous to BEAST model and analysis setup, more advanced simulation options are supported through an extensible markup language (XML) specification, which in addition to generating sequence output, also allows users to combine simulation and analysis in a single BEAST run. Conclusions: piBUSS offers a unique combination of flexibility and ease-of-use for sequence simulation under realistic evolutionary scenarios. Through different interfaces, piBUSS supports simulation studies ranging from modest endeavors for illustrative purposes to complex and large-scale assessments of evolutionary inference procedures. The software aims at implementing new models and data types that are continuously being developed as part of BEAST/BEAGLE.

q-bio.PE