Contrastive Privacy: A Semantic Approach to Measuring Privacy of AI-based Sanitization
AI-based sanitization can remove concepts from images and text, but privacy evaluation remains largely ad hoc. We propose contrastive privacy, a formal definition that yields a quantitative test with a semantic interpretation. Under formal assumptions, we derive a conditional sufficiency result for a class of sanitized renderings (i.e., media files). We operationalize the definition using imperfect semantic-distance models such as CLIP. The test compares sanitized renderings under audit with both the original and sanitized versions of reference renderings known to contain privacy-relevant properties; if the rendering under audit is semantically closer to the unsanitized reference, then the former might leak private information even after sanitization. Importantly, the test is able to conditionally audit an abstract privacy target without enumerating its constituent properties or requiring per-item ground-truth labels; results remain relative to the chosen models and data. We evaluate 34 image-sanitization configurations, 15 social-media text models and one entity recognizer, and four off-the-shelf PII sanitizers on Enron emails. The tests detect residual semantic associations in every setting, including all four PII tools even when they sanitize every candidate their detectors return. Two further studies examine sensitivity to the semantic-distance model. For a synthetic identity, fine-tuning reveals target associations missed by the base model; after broader sanitization, the adapted test detects none. Across 94 matched image collections sanitized to conceal Leonardo DiCaprio's identity, three models yield broadly correlated assessments but sometimes disagree on which sanitizations appear most private. These findings support using multiple models and show how contrastive privacy can reveal retained identifiers and identity-revealing context across sanitization methods.