SearcharxivSearch

arXiv subjects

Joshua Swanson

Publications and source records attributed to Joshua Swanson.

4 recordsLinked to original sources

Large-scale online deanonymization with LLMs

We show that large language models can be used to perform at-scale deanonymization. With full Internet access, our agent can re-identify Hacker News users and Anthropic Interviewer participants at high precision, given pseudonymous online profiles and conversations alone, matching what would take hours for a dedicated human investigator. We then design attacks for the closed-world setting. Given two databases of pseudonymous individuals, each containing unstructured text written by or about that individual, we implement a scalable attack pipeline that uses LLMs to: (1) extract identity-relevant features, (2) search for candidate matches via semantic embeddings, and (3) reason over top candidates to verify matches and reduce false positives. Compared to classical deanonymization work (e.g., on the Netflix prize) that required structured data, our approach works directly on raw user content across arbitrary platforms. We construct three datasets with known ground-truth data to evaluate our attacks. The first links Hacker News to LinkedIn profiles, using cross-platform references that appear in the profiles. Our second dataset matches users across Reddit movie discussion communities; and the third splits a single user's Reddit history in time to create two pseudonymous profiles to be matched. In each setting, LLM-based methods substantially outperform classical baselines, achieving up to 68% recall at 90% precision compared to near 0% for the best non-LLM method. Our results show that the practical obscurity protecting pseudonymous users online no longer holds and that threat models for online privacy need to be reconsidered.

cs.CR

Modal Aphasia: Can Unified Multimodal Models Describe Images From Memory?

We present modal aphasia, a systematic dissociation in which current unified multimodal models accurately memorize concepts visually but fail to articulate them in writing, despite being trained on images and text simultaneously. For one, we show that leading frontier models can generate near-perfect reproductions of iconic movie artwork, but confuse crucial details when asked for textual descriptions. We corroborate those findings through controlled experiments on synthetic datasets in multiple architectures. Our experiments confirm that modal aphasia reliably emerges as a fundamental property of current unified multimodal models, not just as a training artifact. In practice, modal aphasia can introduce vulnerabilities in AI safety frameworks, as safeguards applied to one modality may leave harmful concepts accessible in other modalities. We demonstrate this risk by showing how a model aligned solely on text remains capable of generating unsafe images.

cs.CV

Stirling numbers for complex reflection groups

In an earlier paper, we defined and studied q-analogues of the Stirling numbers of both types for the Coxeter group of type B. In the present work, we show how this approach can be extended to all irreducible complex reflection groups G. The Stirling numbers of the first and second kind are defined via the Whitney numbers of the first and second kind, respectively, of the intersection lattice of G. For the groups G(m,p,n), these numbers and polynomials can be given combinatorial interpretations in terms of various statistics. The ordered version of ths q-Stirling numbers of the second kind also show up in conjectured Hilbert series for certain super coinvariant algebras.

math.CO

Refined Cyclic Sieving on Words for the Major Index Statistic

Reiner-Stanton-White defined the cyclic sieving phenomenon (CSP) associated to a finite cyclic group action and a polynomial. A key example arises from the length generating function for minimal length coset representatives of a parabolic quotient of a finite Coxeter group. In type A, this result can be phrased in terms of the natural cyclic action on words of fixed content. There is a natural notion of refinement for many CSP's. We formulate and prove a refinement, with respect to the major index statistic, of this CSP on words of fixed content by also fixing the cyclic descent type. The argument presented is completely different from Reiner-Stanton-White's representation-theoretic approach. It is combinatorial and largely, though not entirely, bijective in a sense we make precise with a "universal" sieving statistic on words, "flex". A building block of our argument involves cyclic sieving for shifted subset sums, which also appeared in Reiner-Stanton-White. We give an alternate, largely bijective proof of a refinement of this result by extending some ideas of Wagon-Wilf.

math.CO