SearcharxivSearch

arXiv subjects

Guangjie Liu

Publications and source records attributed to Guangjie Liu.

7 recordsLinked to original sources

Inevitability of Encrypted Traffic Side-Channel Leakage in the Multi-Class Setting

The Side-Channel Existence Theorem proves $I(X;Y)>0$ in the binary, undefended setting, but is confined to pairwise arguments and ignores active defenses. We extend it to $k$ classes via the per-class decomposition $I(X;Y)=\sum_iπ_i D_{\mathrm{KL}}(P_{Y|i}\|P_Y)$, with defense cost modelled by per-class Wasserstein-1 constraints $\sup_x W_1(Q_x^D,P_x)\le B$. Three results follow: (1) a summation-form MI lower bound over all active classes; (2) a cascade critical cost theorem and a per-class budget corollary, nonzero where the uniform-budget bound vanishes; (3) an accuracy corollary $\mathrm{Acc}^*\ge 2^{I_0}/k>1/k$. On a 95-class website fingerprinting dataset the measured MI has a strictly positive $95\%$ confidence lower bound under every defense tested. Against the strongest pairwise baseline---a convex program over all $\binom{k}{2}$ triangle constraints, also $Θ(1)$ in $k$ under the same non-vanishing-gap conditions---the summation form is only $1.45\times$ stronger, so the case for the per-class decomposition is structural: only it gives each class a critical cost and a cascade. FRONT's apparent $122\times$ gap is inflated mainly by threshold exclusion rather than the inequality chain: on the active classes it is $21\times$, within $1.4\times$ of the $15\times$ measured undefended. Measuring the chain's two steps separately bounds the collapse onto one Lipschitz statistic below by $28\times$, against a divergence step measured at $1.5\times$. Undefended OVR distinguishability predicts post-defense per-class leakage at Spearman $ρ=0.62$--$0.77$, the transfer the certification procedure relies on. The framework carries over unchanged to a 100-class QUIC/TCP pair.

cs.CR

Rate-Distortion Function for Encrypted Traffic Side-Channel Defense

Parameter selection for encrypted traffic defense has long relied on empirical tuning, yet the fundamental question -- \emph{given a QoS cost budget $D$, how low can the leakage rate go under sustained observation?} -- lacks a provable, computable baseline. Taking the semantic label sequence $X^n$ as the source, the defended feature sequence $Y^n$ as the observation, and Wasserstein-1 distance as the defense cost, we define the \emph{side-channel rate-distortion function} $R^{\mathrm{sc}}(D)$ within the stationary memoryless defense class $Θ_{\mathrm{iid}}$ and provide its complete characterization. We prove that $R^{\mathrm{sc}}(D)$ is monotone decreasing, convex, and continuous, with exact endpoints; the optimal defense has an exponential-tilting (Boltzmann) structure governed by KKT conditions; and the curve constitutes the exact Pareto frontier within $Θ_{\mathrm{iid}}$. For binary equal-prior tasks, $D_{\max} = \tfrac{1}{2}W_1(P_0,P_1)$ via Kantorovich--Rubinstein duality. On real-world website-fingerprinting defenses, the framework locates Front ($Δ_{\mathrm{gap}}{=}0.028$\,bits), WTF-PAD ($0.034$\,bits), and TrafficSliver ($0.124$\,bits) above the theoretical curve, quantifying their suboptimality gaps.

cs.CR

A Queueing-Stability Criterion for Causal IPD-QIM Network Flow Watermarking

On multi-hop encrypted links such as Tor and cascaded VPNs, tunneling flattens packet lengths and protocol fields, leaving inter-packet delay (IPD) as the main carrier for active flow attribution. Causality lets the embedder delay packets but never advance them, so each quantization-index-modulation (QIM) alignment injects nonnegative dwell into a delay buffer; unbounded dwell breaks lattice alignment and delays the host connection unacceptably. Whether a causal QIM watermark embeds stably on bursty traffic has largely been left to empirical configuration rather than analysis. We model the embedder as a reflected dwell queue under the fixed dual-lattice, equiprobable-bit rule, where injection is state-dependent -- set by the current interval and bit -- rather than exogenous. The substitution $Y_i=δ_i-r_i$ gives only an algebraic Lindley-form identity; stability is governed by the busy-state drift at large dwell, where the effective interval collapses to zero and the mean injection becomes $Δ/4$. Away from the critical boundary, the buffer is stable iff $μ_d>Δ/4$ (i.e. $Δ<4μ_d$) for i.i.d. backgrounds, and, under stationary-ergodic and finite-state Markov-modulated traffic with instantaneous overload, iff the time-average intensity $\barρ<1$. With the exogenous decoding floor $Δ\ge cσ_ξ$ ($c=4Q^{-1}(ε/2)$), this yields the operating window $Δ\in[cσ_ξ,4\barμ_d)$. Simulations confirm a sharp transition at $ρ=1$ set only by the mean; on four real IPD traces, with each simulated chain confined to a single flow, the criterion gives the correct stability direction under flow-local correlation and burstiness, while pooled cross-flow means overestimate the margin. These results give a testable stable-embeddability criterion and a quantization-step configuration baseline for causal QIM network flow watermarking.

cs.CR

The Inevitability of Side-Channel Leakage in Encrypted Traffic

The widespread adoption of TLS 1.3 and QUIC has rendered payload content invisible, shifting traffic analysis toward side-channel features. However, rigorous justification for why side-channel leakage is inevitable in encrypted communications has been lacking. This paper establishes a strict foundation from information theory by constructing a formal model \(Σ=(Γ,Ω)\), where \(Γ=(A,Π,Φ,N)\) describes the causal chain of application generation, protocol encapsulation, encryption transformation, and network transmission, while \(Ω\) characterizes observation capabilities. Based on composite channel structure, data processing inequality, and Lipschitz statistics propagation, we propose and prove the Side-Channel Existence Theorem: for distinguishable semantic pairs, under conditions including mapping non-degeneracy (\(\mathbb{E}[d(z_P,z_N)\mid X]\le C\)), protocol-layer distinguishability (expectation difference \(\ge\barΔ\)), Lipschitz continuity, observation non-degeneracy (\(ρ>0\)), and propagation condition (\(C<\barΔ/2L_φ\)), the mutual information \(I(X;Y)\) is strictly positive with explicit lower bound. The corollary shows that in efficiency-prioritized systems, leakage is inevitable when at least one application pair is distinguishable. Three factors determine the boundary: non-degeneracy constant \(C\) constrained by efficiency, distinguishability \(\barΔ\) from application diversity, and \(ρ\) from analyst capabilities. This establishes the first rigorous information-theoretic foundation for encrypted traffic side channels, providing verifiable predictions for attack feasibility, quantifiable benchmarks for defenses, and mathematical basis for efficiency-privacy tradeoffs.

cs.CR

Guided Navigation in Knowledge-Dense Environments: Structured Semantic Exploration with Guidance Graphs

While Large Language Models (LLMs) exhibit strong linguistic capabilities, their reliance on static knowledge and opaque reasoning processes limits their performance in knowledge intensive tasks. Knowledge graphs (KGs) offer a promising solution, but current exploration methods face a fundamental trade off: question guided approaches incur redundant exploration due to granularity mismatches, while clue guided methods fail to effectively leverage contextual information for complex scenarios. To address these limitations, we propose Guidance Graph guided Knowledge Exploration (GG Explore), a novel framework that introduces an intermediate Guidance Graph to bridge unstructured queries and structured knowledge retrieval. The Guidance Graph defines the retrieval space by abstracting the target knowledge' s structure while preserving broader semantic context, enabling precise and efficient exploration. Building upon the Guidance Graph, we develop: (1) Structural Alignment that filters incompatible candidates without LLM overhead, and (2) Context Aware Pruning that enforces semantic consistency with graph constraints. Extensive experiments show our method achieves superior efficiency and outperforms SOTA, especially on complex tasks, while maintaining strong performance with smaller LLMs, demonstrating practical value.

cs.CL

IoT-AMLHP: Aligned Multimodal Learning of Header-Payload Representations for Resource-Efficient Malicious IoT Traffic Classification

Traffic classification is crucial for securing Internet of Things (IoT) networks. Deep learning-based methods can autonomously extract latent patterns from massive network traffic, demonstrating significant potential for IoT traffic classification tasks. However, the limited computational and spatial resources of IoT devices pose challenges for deploying more complex deep learning models. Existing methods rely heavily on either flow-level features or raw packet byte features. Flow-level features often require inspecting entire or most of the traffic flow, leading to excessive resource consumption, while raw packet byte features fail to distinguish between headers and payloads, overlooking semantic differences and introducing noise from feature misalignment. Therefore, this paper proposes IoT-AMLHP, an aligned multimodal learning framework for resource-efficient malicious IoT traffic classification. Firstly, the framework constructs a packet-wise header-payload representation by parsing packet headers and payload bytes, resulting in an aligned and standardized multimodal traffic representation that enhances the characterization of heterogeneous IoT traffic. Subsequently, the traffic representation is fed into a resource-efficient neural network comprising a multimodal feature extraction module and a multimodal fusion module. The extraction module employs efficient depthwise separable convolutions to capture multi-scale features from different modalities while maintaining a lightweight architecture. The fusion module adaptively captures complementary features from different modalities and effectively fuses multimodal features.

cs.NI

Covert Communication Gains from Adversary's Uncertainty of Phase Angles

This work investigates the phase gain of intelligent reflecting surface (IRS) covert communication over complex-valued additive white Gaussian noise (AWGN) channels. The transmitter Alice intends to transmit covert messages to the legitimate receiver Bob via reflecting the broadcast signals from a radio frequency (RF) source, while rendering the adversary Willie's detector arbitrarily close to ineffective. Our analyses show that, compared to the covert capacity for classical AWGN channels, we can achieve a covertness gain of value 2 by leveraging Willie's uncertainty of phase angles. This covertness gain is achieved when the number of possible phase angle pairs $N=2$. More interestingly, our results show that the covertness gain will not further increase with $N$ as long as $N \ge 2$, even if it approaches infinity.

cs.IT