Searcharxiv⌕ Search

arXiv · 2610.02302

Intent-Hiding Jailbreaks: An Information-Theoretic Framework for Compositional Attacks

Abstract

Recent work has shown that large language models (LLMs) can be vulnerable to jailbreak attacks in which harmful intent is obscured through composition with benign tasks. A harmful request refused in isolation may elicit a different response when embedded within a larger, seemingly benign query. We study these compositional intent-hiding jailbreaks from an information-theoretic perspective. Our formulation associates each task with an estimated probability of being judged harmful: the average over the full task collection defines the prior probability of harmful intent, while the average over a selected bundle containing the target defines the posterior. Selecting auxiliary tasks so that these averages agree, which we call prior-posterior matching, leaves the estimated intent unchanged even though the harmful target remains in the bundle. We study two settings that differ in whether query construction is part of the optimization. In the query-independent setting, tasks are selected without regard to how they will be expressed in the final query. We show that exact prior-posterior matching under a bundle-size constraint is computationally hard, derive an optimal water-filling solution for fractional weights, and characterize the smallest bundle satisfying a prescribed safety threshold. In the query-dependent setting, task selection and query construction are considered jointly, and intent concealment and target preservation are evaluated on the resulting query. We evaluate jailbreak effectiveness and preservation of the target behavior across bundle sizes, query generators, and several open-source models. These results show that compositional queries can elicit target behaviors beyond the direct-request baseline under the evaluated search budgets, while revealing a trade-off: as bundle size increases, response-level target preservation tends to decrease for several models.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Fengwei Tian, Ravi Tandon. 2026-10-01. Intent-Hiding Jailbreaks: An Information-Theoretic Framework for Compositional Attacks. https://arxiv.org/abs/2610.02302

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Extended Differential Cryptanalysis of Kuznyechik

We study the input-side relation $F(cx\oplus a)\oplus F(x)=b$ as an extended differential for block-cipher analysis. Because multiplication occurs before the first S-box, fixing that S-box output leaves an ordinary XOR difference for subsequent binary linear layers and common key additions. For permutations, the extended and outer $c$-differential tables are related by inversion. For the Kuznyechik S-box, the exceptional inverse classes $\mathtt{02}/\mathtt{e1}$, $\mathtt{04}/\mathtt{91}$, and $\mathtt{03}/\mathtt{be}$ have exact extended differential uniformities $64$, $33$, and $21$. We connect them to the pseudo-exponential structure of Perrin and Udovenko through a cross-field transfer theorem. In the common polynomial basis, the disagreement rank between multiplication maps in the hidden and Kuznyechik fields equals the degree of the multiplier polynomial, yielding structural lower bounds $60$, $30$, and $16$. We also analyze the complete extended differential table: its normalization is doubly stochastic, its energy is an exact multiplicative correlation of ordinary DDT rows, and transferred hidden fibers give explicit Pearson and Shannon-information bounds. Whitening selects the effective row $a\oplus(1\oplus c)K$. For two-round standard Kuznyechik we derive the exact fixed-key channel $Q_{c,\ell}=W_cP_\ell W_1$ and a two-stage maximum-likelihood attack. With $c=\mathtt{04}$ it recovers the full $256$-bit master key with failure probability at most $0.01$, using $226{,}512$ chosen-plaintext pairs and at most $240{,}669$ encryption queries. Restricted product-model searches give gains of $5.2$, $4.6$, and about $1.6$ bits over the corresponding classical families at two, three, and four rounds. A separate nine-round no-pre-whitening Monte Carlo study is reported only as exploratory evidence because its final configurations followed preliminary screening.

cs.CR↗

Mitigating Watermark Forgery in Generative Models via Randomized Key Selection

Watermarking enables GenAI providers to verify whether content was generated by their models. A watermark is a hidden signal in the content, whose presence can be detected using a secret watermark key. A core security threat are forgery attacks, where adversaries insert the provider's watermark into content \emph{not} produced by the provider, potentially damaging their reputation and undermining trust. Existing defenses resist forgery by embedding many watermarks with multiple keys into the same content, which can degrade model utility. However, forgery remains a threat when attackers can collect sufficiently many watermarked samples. We propose a defense with a sample-count-independent upper bound on forgery success for blind attackers, conditional on key-symmetric, independent detector outcomes. Our scheme does not further degrade model utility. We randomize the watermark key selection for each query and accept content as genuine only if a watermark is detected by \emph{exactly} one key. Unlike cryptographic watermarks that rely on computational hardness assumptions and require designing new watermarking schemes from scratch, our method can be applied to any existing watermarking method to improve its forgery resistance. We focus on text watermarking, but our defense is modality-agnostic, since it treats the underlying watermarking method as a black-box. To show this, we include a preliminary study on image watermarking using Tree-Ring. Separately from this conditional guarantee, we empirically observe that, at $r=4$ keys, harmful-text forgery success drops from as high as $87\%$ with a single key to as low as $1\%$ against the adaptive blind attackers that we evaluate, at negligible computational overhead; a preliminary image study shows a reduction from $100\%$ to $2\%$.

cs.CR↗

Understanding Gaps in LLM Pipelines Towards Scalable Fuzzing Harness Generation: An Empirical Study and Enhancement

Large language model (LLM)-based techniques have achieved notable progress in fuzz harness generation. However, applying them to arbitrary functions \textit{at scale} remains difficult---generated harnesses often fail to compile or, worse, compile but remain logically ineffective. What factors drive success and what limitations hinder current methods remain unclear. To answer these questions, we conduct an empirical study on state-of-the-art LLM-based harness generation frameworks across 29 OSS-Fuzz projects. Our study establishes two pillars of success: contextual information to guide generation and pre-structured pipelines that ensure workflow stability. However, significant gaps remain: (1) existing context retrieval methods lack the robustness to reliably obtain necessary information across diverse projects; (2) current validation mechanisms fail to detect logically ineffective harnesses, such as those containing fabricated definitions; and (3) compilation workflows are brittle, unable to distinguish harness-level errors from build-configuration issues. We demonstrate the utility of these findings by enhancing existing techniques with a hybrid tool pool for robust context retrieval, an enhanced validation pipeline, and a compilation-error triage strategy. Evaluated on 243 OSS-Fuzz projects (65 C and 178 C++), the enhanced approach improves the three-shot success rate by approximately 20\% over state-of-the-art techniques, reaching 87\% for C and 81\% for C++. Our one-hour fuzzing results show that more than 75\% of the generated harnesses increase target-function coverage, surpassing baselines by over 10\%. In addition, the enhanced approach identified 15 new vulnerabilities across 11 real-world projects.

cs.CR↗