arXiv · 2609.33280
The Text Beside the Image: Detection, Utility and Leakage for Trustworthy Multimodal Medical Data and Beyond
Abstract
Medical images are released with the reports that describe them, and protecting the image does not protect the report. This paper measures the text component of such releases. We measure identifier detection, downstream utility and residual identity leakage on the same documents, with the pseudonymisation policy as the variable under test: 15 detectors, three release conditions and four corpora of medical reports, legal judgments, news and other genres, and e-mail, in German, English, Chinese and Arabic. A fixed 13-detector union reaches a person sensitivity of 0.9998 at specificity 0.8686 on the medical reports, 0.9958 at 0.8504 on the legal judgments, 0.9352 at 0.9318 on news and other genres, and 0.9906 at 0.6235 on e-mail. With this ensemble, frequency matching with a public name list recovers zero identities by alignment across the four corpora; the names it got right were ones the detector missed, left in clear text. Cross-document linkage ranks the correct person first for 0.93% of e-mail queries without training and 3.94% with it, against 1/3697 chance and 71.98% on unmodified text. On the medical reports it recovers nothing without training and 0.71% of 138 queries with it, against 1/207 chance and a 2.73% ceiling on unmodified text.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Andreas Maier, Monica Hinrichs-Mayer, Franziska Weber, Niklas Lackner, Matthias May, Bernhard Kainz, Siming Bayer. 2026-09-27. The Text Beside the Image: Detection, Utility and Leakage for Trustworthy Multimodal Medical Data and Beyond. https://arxiv.org/abs/2609.33280
Cite the original work for its findings. Save a collection to share your selection of sources.