SearcharxivSearch

arXiv subjects

Frank Takes

Publications and source records attributed to Frank Takes.

2 recordsLinked to original sources

MARS: A framework for modelling register-based social networks

Register-based social networks have become of increasing interest in countries where formal government-curated microdata is available. Due to the non-trivial generative process of register-based networks, existing random graph models fail to facilitate effective structural analysis, hindering the discovery of meaningful insights in the underlying social system. In this paper we introduce the Multiplex Affiliation-based Random Spatially-embedded (MARS) graph framework, which replicates the construction method of register-based social networks. We derive fundamental statistical properties of MARS ensembles in general and special cases. To demonstrate the applicability of the framework, we implement a simple model under the MARS framework and show that it recovers similar properties to those exhibited by the population-scale register-based social network of the Netherlands. Furthermore, we analyse the effect of spatial tie strength on closure in the network and compare our results with existing empirical findings, showing that increased spatial freedom is correlated with decreased social cohesion.

cs.SI

Improving the output quality of official statistics based on machine learning algorithms

National statistical institutes currently investigate how to improve the output quality of official statistics based on machine learning algorithms. A key obstacle is concept drift, i.e., when the joint distribution of independent variables and a dependent (categorical) variable changes over time. Under concept drift, a statistical model requires regular updating to prevent it from becoming biased. However, updating a model asks for additional data, which are not always available. In the literature, we find a variety of bias correction methods as a promising solution. In the paper, we will compare two popular correction methods: the misclassification estimator and the calibration estimator. For prior probability shift (a specific type of concept drift), we investigate the two correction methods theoretically as well as experimentally. Our theoretical results are expressions for the bias and variance of both methods. As experimental result, we present a decision boundary (as a function of (a) model accuracy, (b) class distribution and (c) test set size) for the relative performance of the two methods. Close inspection of the results will provide a deep insight into the effect of prior probability shift on output quality, leading to practical recommendations on the use of machine learning algorithms in official statistics.

stat.ME