SearcharxivSearch

arXiv subjects

Clement Lee

Publications and source records attributed to Clement Lee.

11 recordsLinked to original sources

Evidencing preferential attachment in dependency network evolution

Preferential attachment is often suggested to be the underlying mechanism of the growth of a network, largely due to that many real networks are, to a certain extent, scale-free. However, such attribution is usually made under debatable practices of determining scale-freeness and when only snapshots of the degree distribution are observed. In the presence of the evolution history of the network, modelling the increments of the evolution allows us to measure preferential attachment directly. Therefore, we propose a generalised linear model for such purpose, where the in-degrees and their increments are the covariate and response, respectively. Not only are the parameters that describe the preferential attachment directly incorporated, they also ensure that the tail heaviness of the asymptotic degree distribution is realistic. The Bayesian approach to inference enables the hierarchical version of the model to be implemented naturally. The application to the dependency network of R packages reveals subtly different behaviours between new dependencies by new and existing packages, and between addition and removal of dependencies.

stat.AP

A spliced preferential attachment model for degree distributions in networks

Identifying the generating mechanism of a network is challenging as, more often than not, only snapshots are available, but not the full evolution. One candidate for the generating mechanism is the general preferential attachment (GPA), in which existing nodes gain new connections at a rate governed by a preference function of their current degree, which, in its simplest form, results in a degree distribution that follows the power law. However, the ubiquity of the power law in real-life networks has been challenged on two fronts: alternative distributions often fit comparably well, and recent works using extreme value methods have shown that the tail of the degree distribution, while still regularly varying, tends to be lighter than the body implies. In this paper, we propose a GPA model with a flexible preference function. Using methods for discrete extremes, we characterise the tail behaviour of the limiting degree distribution directly by the preference function. This direct connection facilitates the inference of the model parameters using snapshot data alone, and sidesteps the need of traditional threshold-based extreme value methods, which lack interpretability and suffer from identifiability issues. Comprehensive simulation studies show that our model recovers the parameters well, while applications to real-life networks demonstrate comparable performance to established alternatives and provide insights into the growth dynamics of the networks.

stat.ME

Conditional Extremes with Graphical Models

Multivariate extreme value analysis quantifies the probability and magnitude of joint extreme events. Classical multivariate models, such as max-stable or multivariate generalised Pareto distributions, generally have a high computational cost of fitting, which limits their application. To overcome this, models based on the asymptotically dependent multivariate Pareto distribution have recently incorporated graphical models to induce sparsity and reduce the dimension of the parameter space. While this approach is computationally efficient, the assumption of asymptotic dependence is inappropriate for many applications. The conditional multivariate extreme value model (CMEVM) is a popular model for which the asymptotic dependence assumption is not required. Unfortunately, inference for this model is semi-parametric, and consequently, it has poor predictive performance in high dimensions. An extension of the CMEVM that allows both the incorporation and selection of sparse dependence structures, and fully parametric prediction is proposed. The approach fills a current gap in statistical methodology by extending graphical models to asymptotically independent multivariate extreme value models. To support inference in high dimensions, a stepwise inference procedure that is computationally efficient and loses no information or predictive power is proposed. Simulation studies show the model is highly flexible, and an application to discharges in the upper Danube River basin provides promising results.

stat.ME

A Bayesian Nonparametric Stochastic Block Model for Directed Acyclic Graphs

Random graphs have been widely used in statistics, for example in network analysis and graphical models. In some applications, the data may contain an inherent hierarchical ordering among its vertices, which prevents directed edges between pairs of vertices that do not respect this order. For example, in bibliometrics, older papers cannot cite newer ones. In such situations, the resulting graph forms a Directed Acyclic Graph. In this article, we extend the Stochastic Block Model (SBM) to account for the presence of such ordering in the data, ignoring which can lead to biased estimates of the number of blocks. The proposed approach includes in the model likelihood a topological ordering, which is treated as an unknown parameter and endowed with a prior distribution. We describe how to formalise the model and perform posterior inference for a Bayesian nonparametric version of the SBM in which both the hierarchical ordering and the number of latent blocks are learnt from the data. Finally, an illustration with real-world datasets from bibliometrics is presented. Additional supplementary materials are available online.

stat.ME

Degree distributions in networks: beyond the power law

The power law is useful in describing count phenomena such as network degrees and word frequencies. With a single parameter, it captures the main feature that the frequencies are linear on the log-log scale. Nevertheless, there have been criticisms of the power law, for example that a threshold needs to be pre-selected without its uncertainty quantified, that the power law is simply inadequate, and that subsequent hypothesis tests are required to determine whether the data could have come from the power law. We propose a modelling framework that combines two different generalisations of the power law, namely the generalised Pareto distribution and the Zipf-polylog distribution, to resolve these issues. The proposed mixture distributions are shown to fit the data well and quantify the threshold uncertainty in a natural way. A model selection step embedded in the Bayesian inference algorithm further answers the question whether the power law is adequate.

stat.AP

A Review of Stochastic Block Models and Extensions for Graph Clustering

There have been rapid developments in model-based clustering of graphs, also known as block modelling, over the last ten years or so. We review different approaches and extensions proposed for different aspects in this area, such as the type of the graph, the clustering approach, the inference approach, and whether the number of groups is selected or estimated. We also review models that combine block modelling with topic modelling and/or longitudinal modelling, regarding how these models deal with multiple types of data. How different approaches cope with various issues will be summarised and compared, to facilitate the demand of practitioners for a concise overview of the current status of these areas of literature.

stat.ML

A hierarchical model of non-homogeneous Poisson processes for Twitter retweets

We present a hierarchical model of non-homogeneous Poisson processes (NHPP) for information diffusion on online social media, in particular Twitter retweets. The retweets of each original tweet are modelled by a NHPP, for which the intensity function is a product of time-decaying components and another component that depends on the follower count of the original tweet author. The latter allows us to explain or predict the ultimate retweet count by a network centrality-related covariate. The inference algorithm enables the Bayes factor to be computed, in order to facilitate model selection. Finally, the model is applied to the retweet data sets of two hashtags.

stat.AP

A Social Network Analysis of Articles on Social Network Analysis

A collection of articles on the statistical modelling and inference of social networks is analysed in a network fashion. The references of these articles are used to construct a citation network data set, which is almost a directed acyclic graph because only existing articles can be cited. A mixed membership stochastic block model is then applied to this data set to soft cluster the articles. The results obtained from a Gibbs sampler give us insights into the influence and the categorisation of these articles.

stat.AP

Performance and sensitivities of home detection from mobile phone data

Large-scale location based traces, such as mobile phone data, have been identified as a promising data source to complement or even enrich official statistics. In many cases, a prerequisite step to deploy the massively gathered data is the detection of home location from individual users. The problem is that little research exists on the validation (comparison with ground truth datasets) or the uncertainty estimation of home detection methods, not at individual user level, nor at nation-wide levels. In this paper, we present an extensive empirical analysis of home detection methods when performed on a nation-wide mobile phone dataset from France. We analyze the validity of 9 different Home Detection Algorithms (HDAs), and we assess different sources of uncertainty. Based on 225 different set-ups for the home detection of around 18 million users we discuss different measures for validation and investigate sensitivity to user choices such as HDA parameter choice and observation period restriction. Our findings show that nation-wide performance of home detection is moderate at best, with correlations to ground truth maximizing at 0.60 only. Additionally, we show that time and duration of observation have a clear effect on performance, and that the effect of HDA criteria and parameter choice are rather small compared to other uncertainties. Our findings and discussion offer welcoming insights to other practitioners who want to apply home detection on similar datasets, or who are in need of an assessment of the challenges and uncertainties related to mobilizing mobile phone data for official statistics.

cs.CY

A Network Epidemic Model for Online Community Commissioning Data

A statistical model assuming a preferential attachment network, which is generated by adding nodes sequentially according to a few simple rules, usually describes real-life networks better than a model assuming, for example, a Bernoulli random graph, in which any two nodes have the same probability of being connected, does. Therefore, to study the propogation of "infection" across a social network, we propose a network epidemic model by combining a stochastic epidemic model and a preferential attachment model. A simulation study based on the subsequent Markov Chain Monte Carlo algorithm reveals an identifiability issue with the model parameters. Finally, the network epidemic model is applied to a set of online commissioning data.

stat.CO

Optimal scaling of the independence sampler: Theory and Practice

The independence sampler is one of the most commonly used MCMC algorithms usually as a component of a Metropolis-within-Gibbs algorithm. The common focus for the independence sampler is on the choice of proposal distribution to obtain an as high as possible acceptance rate. In this paper we have a somewhat different focus concentrating on the use of the independence sampler for updating augmented data in a Bayesian framework where a natural proposal distribution for the independence sampler exists. Thus we concentrate on the proportion of the augmented data to update to optimise the independence sampler. Generic guidelines for optimising the independence sampler are obtained for independent and identically distributed product densities mirroring findings for the random walk Metropolis algorithm. The generic guidelines are shown to be informative beyond the narrow confines of idealised product densities in two epidemic examples.

stat.CO