arXiv · 2605.12264
Reconstruction of Personally Identifiable Information from Proprietary Data in Supervised Fine-Tuned Models
Abstract
Supervised Finetuning (SFT) has become one of the primary methods for adapting a large language model (LLM) with extensive pre-trained knowledge to domain-specific, instruction-following tasks. SFT datasets, composed of instruction-response pairs, often include user-provided information that may contain sensitive data such as personally identifiable information (PII), raising privacy concerns. This paper studies the problem of targeted PII reconstruction from models fine- tuned on proprietary SFT data, in which an adversary attempts to recover PII associated with a specific identity. We construct multi-turn, user-centric Q&A datasets in sensitive domains, specifically medical and legal settings, that incorporate PII to enable realistic evaluation of leakage. We then propose COVA, a coverage-aware decoding algorithm for targeted PII reconstruction under prefix-based attacks. Using COVA, we study how the amount of information available about a target user affects the recovery of their PII from SFT models. Across datasets, COVA consistently improves PII reconstruction over baseline decoding methods, and even limited contextual knowledge can substantially increase an adversary's success. Our findings demonstrate that small, proprietary SFT datasets can induce meaningful privacy leakage through the reconstruction of target-PII associations learned during fine-tuning.
Explore related subjects
Keep this discovery
Sae Furukawa, Alina Oprea. 2026-05-12. Reconstruction of Personally Identifiable Information from Proprietary Data in Supervised Fine-Tuned Models. https://arxiv.org/abs/2605.12264
Cite the original work for its findings. Save a collection to share your selection of sources.