SearcharxivSearch

arXiv subjects

Michele Miranda

Publications and source records attributed to Michele Miranda.

3 recordsLinked to original sources

Differentially Private De-identification of Dutch Clinical Notes: A Comparative Evaluation

Protecting patient privacy in clinical narratives is essential for enabling secondary use of healthcare data under regulations such as GDPR and HIPAA. While manual de-identification remains the gold standard, it is costly and slow, motivating the need for automated methods that combine privacy guarantees with high utility. Most automated text de-identification pipelines employed named entity recognition (NER) to identify protected entities for redaction. Although methods based on differential privacy (DP) provide formal privacy guarantees, more recently also large language models (LLMs) are increasingly used for text de-identification in the clinical domain. In this work, we present the first comparative study of DP, NER, and LLMs for Dutch clinical text de-identification. We investigate these methods separately as well as hybrid strategies that apply NER or LLM preprocessing prior to DP, and assess performance in terms of privacy leakage and extrinsic evaluation (entity and relation classification). We show that DP mechanisms alone degrade utility substantially, but combining them with linguistic preprocessing, especially LLM-based redaction, significantly improves the privacy-utility trade-off.

cs.CR

Preserving Privacy in Large Language Models: A Survey on Current Threats and Solutions

Large Language Models (LLMs) represent a significant advancement in artificial intelligence, finding applications across various domains. However, their reliance on massive internet-sourced datasets for training brings notable privacy issues, which are exacerbated in critical domains (e.g., healthcare). Moreover, certain application-specific scenarios may require fine-tuning these models on private data. This survey critically examines the privacy threats associated with LLMs, emphasizing the potential for these models to memorize and inadvertently reveal sensitive information. We explore current threats by reviewing privacy attacks on LLMs and propose comprehensive solutions for integrating privacy mechanisms throughout the entire learning pipeline. These solutions range from anonymizing training datasets to implementing differential privacy during training or inference and machine unlearning after training. Our comprehensive review of existing literature highlights ongoing challenges, available tools, and future directions for preserving privacy in LLMs. This work aims to guide the development of more secure and trustworthy AI systems by providing a thorough understanding of privacy preservation methods and their effectiveness in mitigating risks.

cs.CR

On integration by parts formula on open convex sets in Wiener spaces

In Euclidean space, it is well known that any integration by parts formula for a set of finite perimeter $Ω$ is expressed by the integration with respect to a measure $P(Ω,\cdot)$ which is equivalent to the one-codimensional Hausdorff measure restricted to the reduced boundary of $Ω$. The same result has been proved in an abstract Wiener space, typically an infinite dimensional space, where the surface measure considered is the one-codimensional spherical Hausdorff-Gauss measure $\mathscr S^{\infty-1}$ restricted to the measure-theoretic boundary of $Ω$. In this paper we consider an open convex set $Ω$ and we provide an explicit formula for the density of $P(Ω,\cdot)$ with respect to $\mathscr S^{\infty-1}$. In particular, the density can be written in terms of the Minkowski functional $\p$ of $Ω$ with respect to an inner point of $Ω$. As a consequence, we obtain an integration by parts formula for open convex sets in Wiener spaces.

math.FA