Searcharxiv⌕ Search

arXiv subjects

Peter-Paul de Wolf

Publications and source records attributed to Peter-Paul de Wolf.

3 recordsLinked to original sources

Disclosure risk in a geo-spatial setting

Using thematic maps to publish statistical information has become a popular visualization. As is the case with all statistical publications, thematic maps also have to deal with the balance between disclosure risk and utility. However, most risk and utility measures do not take into account the spatial character of a map. Some of the proposed spatial risk measures suffer from the Modifiable Areal Unit Problem (MAUP): slightly changing regional classifications may influence the risk. Indeed, even a small translation of for example a grid may influence that risk. We propose a new risk measure that does not suffer from MAUP. Moreover, our risk is directly related to the local density of the (target) population and takes into account that often multiple units may be connected to a single location. We show the behavior of our risk measure using an example dataset of fake but realistic locations of enterprises. Our risk measure can be adapted to take into account the effect on the (perceived) risk of zooming in or out and the effect of the used resolution.

stat.ME↗

A density ratio framework for evaluating the utility of synthetic data

Synthetic data generation is a promising technique to facilitate the use of sensitive data while mitigating the risk of privacy breaches. However, for synthetic data to be useful in downstream analysis tasks, it needs to be of sufficient quality. Various methods have been proposed to measure the utility of synthetic data, but their results are often incomplete or even misleading. In this paper, we propose using density ratio estimation to improve quality evaluation for synthetic data, and thereby the quality of synthesized datasets. We show how this framework relates to and builds on existing measures, yielding global and local utility measures that are informative and easy to interpret. We develop an estimator which requires little to no manual tuning due to automatic selection of a nonparametric density ratio model. Through simulations, we find that density ratio estimation yields more accurate estimates of global utility than established procedures. A real-world data application demonstrates how the density ratio can guide refinements of synthesis models and can be used to improve downstream analyses. We conclude that density ratio estimation is a valuable tool in synthetic data generation workflows and provide these methods in the accessible open source R-package densityratio.

stat.ML↗

When Machine Learning Models Leak: An Exploration of Synthetic Training Data

We investigate an attack on a machine learning model that predicts whether a person or household will relocate in the next two years, i.e., a propensity-to-move classifier. The attack assumes that the attacker can query the model to obtain predictions and that the marginal distribution of the data on which the model was trained is publicly available. The attack also assumes that the attacker has obtained the values of non-sensitive attributes for a certain number of target individuals. The objective of the attack is to infer the values of sensitive attributes for these target individuals. We explore how replacing the original data with synthetic data when training the model impacts how successfully the attacker can infer sensitive attributes.

cs.LG↗