arXiv · 0810.5582
Anonymizing Unstructured Data
Abstract
In this paper we consider the problem of anonymizing datasets in which each individual is associated with a set of items that constitute private information about the individual. Illustrative datasets include market-basket datasets and search engine query logs. We formalize the notion of k-anonymity for set-valued data as a variant of the k-anonymity model for traditional relational datasets. We define an optimization problem that arises from this definition of anonymity and provide O(klogk) and O(1)-approximation algorithms for the same. We demonstrate applicability of our algorithms to the America Online query log dataset.
Explore related subjects
Keep this discovery
Rajeev Motwani, Shubha U. Nabar. 2008-11-03. Anonymizing Unstructured Data. https://arxiv.org/abs/0810.5582
Cite the original work for its findings. Save a collection to share your selection of sources.