SearcharxivSearch

arXiv subjects

Mojtaba Abdolmaleki

Publications and source records attributed to Mojtaba Abdolmaleki.

2 recordsLinked to original sources

A Training-free Method for LLM Text Attribution

Verifying the provenance of content is crucial to the functioning of many organizations, e.g., educational institutions, social media platforms, and firms. This problem is becoming increasingly challenging as text generated by Large Language Models (LLMs) becomes almost indistinguishable from human-generated content. In addition, many institutions use in-house LLMs and want to ensure that external, non-sanctioned LLMs do not produce content within their institutions. In this paper, we answer the following question: Given a piece of text, can we identify whether it was produced by a particular LLM, while ensuring a guaranteed low false positive rate? We model LLM text as a sequential stochastic process with complete dependence on history. We then design zero-shot statistical tests to (i) distinguish between text generated by two different known sets of LLMs $A$ (non-sanctioned) and $B$ (in-house), and (ii) identify whether text was generated by a known LLM or by any unknown model. We prove that the Type I and Type II errors of our test decrease exponentially with the length of the text. We also extend our theory to black-box access via sampling and characterize the required sample size to obtain essentially the same Type I and Type II error upper bounds as in the white-box setting (i.e., with access to $A$). We show the tightness of our upper bounds by providing an information-theoretic lower bound. We next present numerical experiments to validate our theoretical results and assess their robustness in settings with adversarial post-editing. Our work has a host of practical applications in which determining the origin of a text is important and can also be useful for combating misinformation and ensuring compliance with emerging AI regulations. See https://github.com/TaraRadvand74/llm-text-detection for code, data, and an online demo of the project.

stat.ML

Minimum Weight Pairwise Distance Preservers

In this paper, we study the Minimum Weight Pairwise Distance Preservers (MWPDP) problem. Consider a positively weighted undirected/directed connected graph $G = (V, E, c)$ and a subset $P$ of pairs of vertices, also called demand pairs. A subgraph $G'$ is a distance preserver with respect to $P$ if and only if every pair $(u, w) \in P$ satisfies $dist_{G'} (u, w) = dist_{G}(u, w)$. In MWPDP problem, we aim to find the minimum-weight subgraph $G^*$ that is a distance preserver with respect to $P$. Taking a shortest path between each pair in $P$ gives us a trivial solution with the weight of at most $U=\sum_{(u,v) \in P} dist_{G} (u, w)$. Subsequently, we ask how much improvement we can make upon $U$. In other words, we opt to find a distance preserver $G^*$ that maximizes $U-c(G^*)$. Denote this problem as Cost Sharing Pairwise Distance Preservers (CSPDP), which has several applications in the planning and operations of transportation systems. The only known work that can provide a nontrivial solution for CSPDP is that of Chlamtáč et al. (SODA, 2017). This algorithm works for unweighted graphs and guarantees a non-zero objective only if the optimal solution is extremely sparse with respect to the trivial solution. We address this issue by proposing an $O(|E|^{1/2+ε})$-approximation algorithm for CSPDP in weighted graphs that runs in $O((|P||E|)^{2.38} (1/ε))$ time. Moreover, we prove CSPDP is at least as hard as $\text{LABEL-COVER}_{\max}$. This implies that CSPDP cannot be approximated within $O(|E|^{1/6-ε})$ factor in polynomial time, unless there is an improvement in the notoriously difficult $\text{LABEL-COVER}_{\max}$.

cs.DS