SearcharxivSearch

arXiv subjects

Brendan Murphy

Publications and source records attributed to Brendan Murphy.

At least 19 recordsLinked to original sources

Jailbreak-Tuning: Models Efficiently Learn Jailbreak Susceptibility

AI systems are rapidly advancing in capability, and frontier model developers broadly acknowledge the need for safeguards against serious misuse. However, this paper demonstrates that fine-tuning, whether via open weights or closed fine-tuning APIs, can produce helpful-only models with safeguards destroyed. In contrast to prior work which is blocked by modern moderation systems or achieved only partial removal of safeguards or degraded output quality, our jailbreak-tuning method teaches models to generate detailed, high-quality responses to arbitrary harmful requests. For example, OpenAI, Google, and Anthropic models will fully comply with requests for CBRN assistance, executing cyberattacks, and other criminal activity. We further show that backdoors can increase not only the stealth but also the severity of attacks. Stronger jailbreak prompts become even more effective in fine-tuning attacks, linking attacks and potentially defenses in the input and weight spaces. Not only are current models vulnerable, more recent ones also appear to be becoming even more vulnerable to these attacks, underscoring the urgent need for tamper-resistant safeguards. Until such safeguards are discovered, companies and policymakers should view the release of any fine-tunable model as simultaneously releasing its evil twin: equally capable as the original model, and usable for any malicious purpose within its capabilities.

cs.CR

On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback

As LLMs become more widely deployed, there is increasing interest in directly optimizing for feedback from end users (e.g. thumbs up) in addition to feedback from paid annotators. However, training to maximize human feedback creates a perverse incentive structure for the AI to resort to manipulative or deceptive tactics to obtain positive feedback from users who are vulnerable to such strategies. We study this phenomenon by training LLMs with Reinforcement Learning with simulated user feedback in environments of practical LLM usage. In our settings, we find that: 1) Extreme forms of "feedback gaming" such as manipulation and deception are learned reliably; 2) Even if only 2% of users are vulnerable to manipulative strategies, LLMs learn to identify and target them while behaving appropriately with other users, making such behaviors harder to detect; 3) To mitigate this issue, it may seem promising to leverage continued safety training or LLM-as-judges during training to filter problematic outputs. Instead, we found that while such approaches help in some of our settings, they backfire in others, sometimes even leading to subtler manipulative behaviors. We hope our results can serve as a case study which highlights the risks of using gameable feedback sources -- such as user feedback -- as a target for RL.

cs.LG

Identifying Factors Contributing to Bad Days for Software Developers: A Mixed Methods Study

Software development is a dynamic activity that requires engineers to work effectively with tools, processes, and collaborative teams. As a result, the presence of friction can significantly hinder productivity, increase frustration, and contribute to low morale among developers. By contrast, higher satisfaction levels are positively correlated with higher levels of perceived productivity. Hence, understanding the factors that cause bad experiences for developers is critical for fostering a positive and productive engineering environment. In this research, we employed a mixed-method approach, including interviews, surveys, diary studies, and analysis of developer telemetry data to uncover and triangulate common factors that cause "bad days" for developers. The interviews involved 22 developers across different levels and roles. The survey captured the perception of 214 developers about factors that cause them to have "bad days," their frequency, and their impact on job satisfaction. The daily diary study engaged 79 developers for 30 days to document factors that caused "bad days" in the moment. We examined the telemetry signals of 131 consenting participants to validate the impact of bad developer experience using system data. Findings from our research revealed factors that cause "bad days" for developers and significantly impact their work and well-being. We discuss the implications of these findings and suggest future work.

cs.SE

Scaling Trends for Data Poisoning in LLMs

LLMs produce harmful and undesirable behavior when trained on datasets containing even a small fraction of poisoned data. We demonstrate that GPT models remain vulnerable to fine-tuning on poisoned data, even when safeguarded by moderation systems. Given the persistence of data poisoning vulnerabilities in today's most capable models, this paper investigates whether these risks increase with model scaling. We evaluate three threat models -- malicious fine-tuning, imperfect data curation, and intentional data contamination -- across 24 frontier LLMs ranging from 1.5 to 72 billion parameters. Our experiments reveal that larger LLMs are significantly more susceptible to data poisoning, learning harmful behaviors from even minimal exposure to harmful data more quickly than smaller models. These findings underscore the need for leading AI companies to thoroughly red team fine-tuning APIs before public release and to develop more robust safeguards against data poisoning, particularly as models continue to scale in size and capability.

cs.CR

On Korobov bound concerning Zaremba's conjecture

We prove in particular that for any sufficiently large prime $p$ there is $1\le a<p$ such that all partial quotients of $a/p$ are bounded by $O(\log p/\log \log p)$. For composite denominators a similar result is obtained. This improves the well--known Korobov bound concerning Zaremba's conjecture from the theory of continued fractions.

math.NT

What are Weak Links in the npm Supply Chain?

Modern software development frequently uses third-party packages, raising the concern of supply chain security attacks. Many attackers target popular package managers, like npm, and their users with supply chain attacks. In 2021 there was a 650% year-on-year growth in security attacks by exploiting Open Source Software's supply chain. Proactive approaches are needed to predict package vulnerability to high-risk supply chain attacks. The goal of this work is to help software developers and security specialists in measuring npm supply chain weak link signals to prevent future supply chain attacks by empirically studying npm package metadata. In this paper, we analyzed the metadata of 1.63 million JavaScript npm packages. We propose six signals of security weaknesses in a software supply chain, such as the presence of install scripts, maintainer accounts associated with an expired email domain, and inactive packages with inactive maintainers. One of our case studies identified 11 malicious packages from the install scripts signal. We also found 2,818 maintainer email addresses associated with expired domains, allowing an attacker to hijack 8,494 packages by taking over the npm accounts. We obtained feedback on our weak link signals through a survey responded to by 470 npm package developers. The majority of the developers supported three out of our six proposed weak link signals. The developers also indicated that they would want to be notified about weak links signals before using third-party packages. Additionally, we discussed eight new signals suggested by package developers.

cs.CR

Growth in linear groups

We prove a conjecture of Helfgott on the structure of sets of bounded tripling in bounded rank, which states the following. Let $A$ be a finite symmetric subset of $\mathrm{GL}_n(\mathbf{F})$ for any field $\mathbf{F}$ such that $|A^3| \leq K|A|$. Then there are subgroups $H \trianglelefteq \Gamma \trianglelefteq \langle A \rangle$ such that $A$ is covered by $K^{O_n(1)}$ cosets of $\Gamma$, $\Gamma/H$ is nilpotent of step at most $n-1$, and $H$ is contained in $A^{O_n(1)}$. This theorem includes the Product Theorem for finite simple groups of bounded rank as a special case. As an application of our methods we also show that the diameter of sufficiently quasirandom finite linear groups is poly-logarithmic.

math.GR

On the Pinned Distances Problem in Positive Characteristic

We study the Erd\H os-Falconer distance problem for a set $A\subset \mathbb{F}^2$, where $\mathbb{F}$ is a field of positive characteristic $p$. If $\mathbb{F}=\mathbb{F}_p$ and the cardinality $|A|$ exceeds $p^{5/4}$, we prove that $A$ determines an asymptotically full proportion of the feasible $p$ distances. For small sets $A$, namely when $|A|\leq p^{4/3}$ over any $\mathbb{F}$, we prove that either $A$ determines $\gg|A|^{2/3}$. For both large and small sets, the results proved are in fact for pinned distances.

math.CO

Growth in Some Finite Three-Dimensional Matrix Groups

We study the growth of product sets in some finite three-dimensional matrix groups. In particular, we prove two results about the group of $2\times 2$ upper triangular matrices over arbitrary finite fields: a product set estimate using techniques from multiplicative combinatorics, and an energy estimate using incidence geometry. The energy method gives better quantitative results, but only applies to small sets. We also prove an energy result for the Heisenberg group.

math.CO

Bisector energy and pinned distances in positive characteristic

We prove a new lower bound for the number of pinned distances over finite fields: if $A$ is a sufficiently small subset of $\mathbb{F}_q^2$, then there is an element in $A$ that determines $\gg |A|^{2/3}$ distinct distances to other elements of $A$. Combined with results for large subsets $A\subseteq\mathbb{F}_q^2$, this improves all previously known lower bounds on distinct distances over finite fields. In fact, we obtain an upper bound for the number of isosceles triangles determined by $A$. For that we use the concept of bisector energy. It turns out that the latter can be expressed as a point-plane incidence bound, so one can use a theorem of the third author. The conversion to this incidence problem relies on the Blaschke-Grünwald kinematic mapping -- an embedding of the group of rigid motions of $\mathbb{F}_q^2$ into an open subset of the projective three space. This has long been known in kinematics and geometric algebra; we provide a proof for arbitrary fields using Clifford algebras.

math.CO

Group Action Combinatorics

This paper generalizes the basic notions of additive and multiplicative combinatorics to the setting of group actions: if $G$ is a group acting on a set $X$, and we have subsets $A\subseteq G$ and $Y\subseteq X$ such that the set of pairs $g\cdot y$ with $g\in A,y\in Y$ is not much larger than $Y$, what structure must $A$ and $Y$ have? Briefly, what is the structure of sets with small "image set"? In this setting, we develop analogs of Ruzsa's triangle inequality, covering theorems, multiplicative energy, and the Balog-Szemerédi-Gowers theorem. Approximate stabilizers, which we call symmetry sets, play an important role. While our focus is on presenting a general theory, we answer the inverse image set question in some special cases. To do so, we combine the group action version of the Balog-Szemerédi-Gowers theorem with structure theorems for approximate groups and bounds for the sizes of symmetry sets.

math.CO

New results on sum-product type growth over fields

We prove a range of new sum-product type growth estimates over a general field $\mathbb{F}$, in particular the special case $\mathbb{F}=\mathbb{F}_p$. They are unified by the theme of "breaking the $3/2$ threshold", epitomising the previous state of the art. These estimates stem from specially suited applications of incidence bounds over $\mathbb{F}$, which apply to higher moments of representation functions. We establish the estimate $|R[A]| \gtrsim |A|^{8/5}$ for cardinality of the set $R[A]$ of distinct cross-ratios defined by triples of elements of a (sufficiently small if $\mathbb{F}$ has positive characteristic, similarly for the rest of the estimates) set $A\subset \mathbb{F}$, pinned at infinity. The cross-ratio naturally arises in various sum-product type questions of projective nature and is the unifying concept underlying most of our results. It enables one to take advantage of its symmetry properties as an onset of growth of, for instance, products of difference sets. The geometric nature of the cross-ratio enables us to break the version of the above threshold for the minimum number of distinct triangle areas $Ouu'$, defined by points $u,u'$ of a non-collinear point set $P\subset \mathbb{F}^2$. Another instance of breaking the threshold is showing that if $A$ is sufficiently small and has additive doubling constant $M$, then $|AA|\gtrsim M^{-2}|A|^{14/9}$. This result has a second moment version, which allows for new upper bounds for the number of collinear point triples in the set $A\times A\subset \mathbb{F}^2$, the quantity often arising in applications of geometric incidence estimates.

math.CO

Products of Differences over Arbitrary Finite Fields

There exists an absolute constant $δ> 0$ such that for all $q$ and all subsets $A \subseteq \mathbb{F}_q$ of the finite field with $q$ elements, if $|A| > q^{2/3 - δ}$, then \[ |(A-A)(A-A)| = |\{ (a -b) (c-d) : a,b,c,d \in A\}| > \frac{q}{2}. \] Any $δ< 1/13,542$ suffices for sufficiently large $q$. This improves the condition $|A| > q^{2/3}$, due to Bennett, Hart, Iosevich, Pakianathan, and Rudnev, that is typical for such questions. Our proof is based on a qualitatively optimal characterisation of sets $A,X \subseteq \mathbb{F}_q$ for which the number of solutions to the equation \[ (a_1-a_2) = x (a_3-a_4) \, , \; a_1,a_2, a_3, a_4 \in A, x \in X \] is nearly maximum. A key ingredient is determining exact algebraic structure of sets $A, X$ for which $|A + XA|$ is nearly minimum, which refines a result of Bourgain and Glibichuk using work of Gill, Helfgott, and Tao. We also prove a stronger statement for \[ (A-B)(C-D) = \{ (a -b) (c-d) : a \in A, b \in B, c \in C, d \in D\} \] when $A,B,C,D$ are sets in a prime field, generalising a result of Roche-Newton, Rudnev, Shkredov, and the authors.

math.CO

Upper and lower bounds for rich lines in grids

We prove upper and lower bounds for the number of lines in general position that are rich in a Cartesian product point set. This disproves a conjecture of Solymosi and improves work of Elekes, Borenstein and Croot, and Amirkhanyan, Bush, Croot, and Pryby. The upper bounds are based on a version of the asymmetric Balog-Szemeredi-Gowers theorem for group actions combined with product theorems for the affine group. The lower bounds are based on a connection between rich lines in Cartesian product sets and amenability (or expanding families of graphs in the finite field case). As an application of our upper bounds for rich lines in grids, we give a geometric proof of the asymmetric sum-product estimates of Bourgain and Shkredov.

math.CO

Popular Products and Continued Fractions

We prove bounds for the popularity of products of sets with weak additive structure, and use these bounds to prove results about continued fractions. Namely, we obtain a nearly sharp upper bound for the cardinality of Zaremba's set modulo $p$.

math.NT

Axiomatic Foundations and Algorithms for Deciding Semantic Equivalences of SQL Queries

Deciding the equivalence of SQL queries is a fundamental problem in data management. As prior work has mainly focused on studying the theoretical limitations of the problem, very few implementations for checking such equivalences exist. In this paper, we present a new formalism and implementation for reasoning about the equivalences of SQL queries. Our formalism, U-semiring, extends SQL's semiring semantics with unbounded summation and duplicate elimination. U-semiring is defined using only very few axioms and can thus be easily implemented using proof assistants such as Coq for automated query reasoning. Yet, they are sufficient enough to enable us reason about sophisticated SQL queries that are evaluated over bags and sets, along with various integrity constraints. To evaluate the effectiveness of U-semiring, we have used it to formally verify 39 query rewrite rules from both classical data management research papers and real-world SQL engines, where many of them have never been proven correct before.

cs.DB

On the few products, many sums problem

We prove new results on additive properties of finite sets $A$ with small multiplicative doubling $|AA|\leq M|A|$ in the category of real/complex sets as well as multiplicative subgroups in the prime residue field. The improvements are based on new combinatorial lemmata, which may be of independent interest. Our main results are the inequality $$ |A-A|^3|AA|^5 \gtrsim |A|^{10}, $$ over the reals, "redistributing" the exponents in the textbook Elekes sum-product inequality and the new best known additive energy bound $\mathsf E(A)\lesssim_M |A|^{49/20}$, which aligns, in a sense to be discussed, with the best known sum set bound $|A+A|\gtrsim_M |A|^{8/5}$. These bounds, with $M=1$, also apply to multiplicative subgroups of $\mathbb F^\times_p$, whose order is $O(\sqrt{p})$. We adapt the above energy bound to larger subgroups and obtain new bounds on gaps between elements in cosets of subgroups of order $Ω(\sqrt{p})$.

math.CO