SearcharxivSearch

arXiv subjects

Joe Meagher

Publications and source records attributed to Joe Meagher.

3 recordsLinked to original sources

Modelling superspreading dynamics and circadian rhythms in online discussion boards using Hawkes processes

Online boards offer a platform for sharing and discussing content, where discussion emerges as a cascade of comments in response to a post. Branching point process models offer a practical approach to modelling these cascades; however, existing models do not account for apparent features of empirical data. We address this gap by illustrating the flexibility of Hawkes processes to model data arising from this context as well as outlining the computational tools needed to service this class of models. For example, the distribution of replies within discussions tends to have a heavy tail. As such, a small number of posts and comments may generate many replies, while most generate few or none, similar to `superspreading' in epidemics. Here, we propose a novel model for online discussion, motivated by a dataset arising from discussions on the r/ireland subreddit, that accommodates such phenomena and develop a framework for Bayesian inference that considers in- and out-of-sample tests for goodness-of-fit. This analysis shows that discussions within this community follow a circadian rhythm and are subject to moderate superspreading dynamics. For example, we estimate that the expected discussion size is approximately four for initial posts between 04:00 and 12:00 but approximately 2.5 from 15:00 to 02:00. We also estimate that 58% to 62% of posts fail to generate any discussion, with 95% posterior probability. Thus, we demonstrate that our framework offers a general approach to modelling discussion on online boards.

stat.AP

Classification of cow diet based on milk mid infrared spectra: a data analysis competition at the "International workshop of spectroscopy and chemometrics 2022"

In April 2022, the Vistamilk SFI Research Centre organized the second edition of the "International Workshop on Spectroscopy and Chemometrics - Applications in Food and Agriculture". Within this event, a data challenge was organized among participants of the workshop. Such data competition aimed at developing a prediction model to discriminate dairy cows' diet based on milk spectral information collected in the mid-infrared region. In fact, the development of an accurate and reliable discriminant model for dairy cows' diet can provide important authentication tools for dairy processors to guarantee product origin for dairy food manufacturers from grass-fed animals. Different statistical and machine learning modelling approaches have been employed during the workshop, with different pre-processing steps involved and different degree of complexity. The present paper aims to describe the statistical methods adopted by participants to develop such classification model.

q-bio.QM

Assessing epidemic curves for evidence of superspreading

The expected number of secondary infections arising from each index case, referred to as the reproduction or $R$ number, is a vital summary statistic for understanding and managing epidemic diseases. There are many methods for estimating $R$; however, few explicitly model heterogeneous disease reproduction, which gives rise to superspreading within the population. We propose a parsimonious discrete-time branching process model for epidemic curves that incorporates heterogeneous individual reproduction numbers. Our Bayesian approach to inference illustrates that this heterogeneity results in less certainty on estimates of the time-varying cohort reproduction number $R_t$. We apply these methods to a COVID-19 epidemic curve for the Republic of Ireland and find support for heterogeneous disease reproduction. Our analysis allows us to estimate the expected proportion of secondary infections attributable to the most infectious proportion of the population. For example, we estimate that the 20% most infectious index cases account for approximately 75-98% of the expected secondary infections with 95% posterior probability. In addition, we highlight that heterogeneity is a vital consideration when estimating $R_t$.

stat.AP