SearcharxivSearch

arXiv subjects

Sae Furukawa

Publications and source records attributed to Sae Furukawa.

2 recordsLinked to original sources

Reconstruction of Personally Identifiable Information from Proprietary Data in Supervised Fine-Tuned Models

Supervised Finetuning (SFT) has become one of the primary methods for adapting a large language model (LLM) with extensive pre-trained knowledge to domain-specific, instruction-following tasks. SFT datasets, composed of instruction-response pairs, often include user-provided information that may contain sensitive data such as personally identifiable information (PII), raising privacy concerns. This paper studies the problem of targeted PII reconstruction from models fine- tuned on proprietary SFT data, in which an adversary attempts to recover PII associated with a specific identity. We construct multi-turn, user-centric Q&A datasets in sensitive domains, specifically medical and legal settings, that incorporate PII to enable realistic evaluation of leakage. We then propose COVA, a coverage-aware decoding algorithm for targeted PII reconstruction under prefix-based attacks. Using COVA, we study how the amount of information available about a target user affects the recovery of their PII from SFT models. Across datasets, COVA consistently improves PII reconstruction over baseline decoding methods, and even limited contextual knowledge can substantially increase an adversary's success. Our findings demonstrate that small, proprietary SFT datasets can induce meaningful privacy leakage through the reconstruction of target-PII associations learned during fine-tuning.

cs.CR

Towards the Development of a Real-Time Deepfake Audio Detection System in Communication Platforms

Deepfake audio poses a rising threat in communication platforms, necessitating real-time detection for audio stream integrity. Unlike traditional non-real-time approaches, this study assesses the viability of employing static deepfake audio detection models in real-time communication platforms. An executable software is developed for cross-platform compatibility, enabling real-time execution. Two deepfake audio detection models based on Resnet and LCNN architectures are implemented using the ASVspoof 2019 dataset, achieving benchmark performances compared to ASVspoof 2019 challenge baselines. The study proposes strategies and frameworks for enhancing these models, paving the way for real-time deepfake audio detection in communication platforms. This work contributes to the advancement of audio stream security, ensuring robust detection capabilities in dynamic, real-time communication scenarios.

cs.SD