Searcharxiv⌕ Search

arXiv subjects

Junlin Liao

Publications and source records attributed to Junlin Liao.

2 recordsLinked to original sources

TrackFlood: Relocating Latency Attacks from NMS-Free Detectors to Real-Time Trackers

We consider latency attacks on object detectors, where the attacker's goal is not to corrupt a prediction but to make the system fail to respond in time, targeting real-time applications such as autonomous driving. Modern object detectors eliminate Non-Maximum Suppression (NMS) through one-to-one assignment or set prediction, removing the classical detector-side latency bottleneck exploited by prior latency (``sponge'') attacks. We show that NMS-free does not mean latency-robust: this architectural change does not eliminate the attack surface but relocates it downstream to multi-object tracking, whose data-association cost remains content dependent. We present \emph{TrackFlood}, a unified white-box overload attack against NMS-free detect-then-track pipelines spanning both one-to-one detectors (YOLOv10 and YOLO26) and query-based detectors (RT-DETR). TrackFlood recovers differentiable confidence tensors and optimizes perturbations that flood the tracker with spatially distributed phantom detections while leaving detector inference unchanged. Evaluated entirely on an NVIDIA Jetson AGX Orin (TensorRT FP16), detector latency remains essentially constant, whereas tracker latency increases substantially. At a standard imperceptible budget ($L_\infty{=}8/255$), a universal perturbation produces clearly measurable tracker overload but only modest end-to-end slowdown, without deadline misses for the association-dominated trackers; a separate higher-budget stress test drives severe end-to-end slowdowns and sustained deadline misses. We further evaluate a lightweight, architecture-agnostic bounded-admission layer that caps the tracker workload and largely restores end-to-end latency, at a non-trivial cost in admitted clean detections. Our results demonstrate that evaluating NMS-free perception systems requires considering downstream tracking and end-to-end timing, not detector inference alone.

cs.CR↗

GateDrain: Availability Attacks and Admission-Side Defense for Confidence-Gated Edge-Cloud Inference

Confidence-gated edge--cloud inference accepts confident local predictions and offloads uncertain inputs to a stronger cloud model. We show that this routing decision creates an availability attack surface. We call this attack \emph{GateDrain}: bounded input perturbations lower calibrated confidence and redirect requests that would otherwise be answered locally into a shared cloud queue, without increasing the application request rate. Because escalated requests share a cloud service, an increase in per-request cloud demand can move a near-capacity deployment across a queueing knee, causing disproportionate tail-latency degradation for benign users. We evaluate white-box, transfer, decision-only, universal, and multi-gate attacks on the public EdgeBoost artifact. A fixed-application-volume comparison isolates the effect of confidence manipulation from added client traffic, while perturbation-budget and arrival-process sweeps show that the queueing transition persists across several workload models but its amplification depends on the operating point. Adaptive attacks also defeat the evaluated training-free preprocessing defenses. To contain the resulting cloud demand, we evaluate Bounded Escalation, which combines per-source admission budgets, protected capacity, and non-preemptive trusted-class priority; an optional global bucket adds an identity-independent bound on untrusted admissions. The evaluation makes the resulting policy trade-off explicit: authenticated clients receive latency isolation, whereas tighter aggregate containment can reject legitimate unauthenticated offloads and reduce overall expected accuracy through edge fallback.

cs.CR↗