SearcharxivSearch

arXiv subjects

Jan Greve

Publications and source records attributed to Jan Greve.

3 recordsLinked to original sources

Named Entity Swapping for Metadata Anonymization in a Text Corpus

This work introduces an anonymization scheme for a corpus of texts to safeguard metadata from disclosure. It specifically aims to prevent large language models from identifying metadata associated with texts, thereby avoiding their influence on query responses. The core mechanism is called named entity swapping, a technique inspired by data swapping in statistical disclosure control. Our method randomly selects pairs of semantically similar substrings from different texts based on the similarity of their embedding vectors and interchanges some named entities between them. This prevents certain combinations of named entities from being uniquely associated with the metadata of individual texts. Our approach offers two key advantages. First, it enables users to determine the optimal level of anonymization that balances data utility and data risk through a calibration of several key decision variables. Second, it leverages text embeddings both to compute swapping weights and to assess data utility, enabling a high degree of flexibility and customization in the overall workflow. The effectiveness of the proposed method is demonstrated with an application that prevents the disclosure of company names in a cross-sectional dataset of earnings call transcripts.

stat.AP

A New Representation of Ewens-Pitman's Partition Structure and Its Characterization via Riordan Array Sums

Ewens-Pitman's partition structure arises as a system of sampling consistent probability distributions on set partitions induced by the Pitman-Yor process. It is widely used in statistical applications, particularly in species sampling models in Bayesian nonparametrics. Drawing references from the area of representation theory of the infinite symmetric group, we view Ewens-Pitman's partition structure as an example of a non-extreme harmonic function on a branching graph, specifically, the Kingman graph. Taking this perspective enables us to obtain combinatorial and algebraic constructions of this distribution using the interpolation polynomial approach proposed by Borodin and Olshanski (The Electronic Journal of Combinatorics, 7, 2000). We provide a new explicit representation of Ewens-Pitman's partition structure using modern umbral interpolation based on Sheffer polynomial sequences. In addition, we show that a certain type of marginals of this distribution can be computed using weighted row sums of a Riordan array. In this way, we show that some summary statistics and estimators derived from Ewens-Pitman's partition structure can be obtained using methods of generating functions. This approach simplifies otherwise cumbersome calculations of these quantities often involving various special combinatorial functions. In addition, it has the added benefit of being amenable to symbolic computation.

stat.ME

Spying on the prior of the number of data clusters and the partition distribution in Bayesian cluster analysis

Cluster analysis aims at partitioning data into groups or clusters. In applications, it is common to deal with problems where the number of clusters is unknown. Bayesian mixture models employed in such applications usually specify a flexible prior that takes into account the uncertainty with respect to the number of clusters. However, a major empirical challenge involving the use of these models is in the characterisation of the induced prior on the partitions. This work introduces an approach to compute descriptive statistics of the prior on the partitions for three selected Bayesian mixture models developed in the areas of Bayesian finite mixtures and Bayesian nonparametrics. The proposed methodology involves computationally efficient enumeration of the prior on the number of clusters in-sample (termed as ``data clusters'') and determining the first two prior moments of symmetric additive statistics characterising the partitions. The accompanying reference implementation is made available in the R package 'fipp'. Finally, we illustrate the proposed methodology through comparisons and also discuss the implications for prior elicitation in applications.

stat.ME