Searcharxiv⌕ Search

arXiv subjects

Amin Saremi

Publications and source records attributed to Amin Saremi.

2 recordsLinked to original sources

Deep Learning Latency Attacks and Defenses: A Cross-Domain Survey of Availability Threats

Adversarial machine learning has focused mainly on integrity, but availability is an increasingly consequential complement. Latency attacks (also energy-latency attacks) increase inference-time work, energy, or response time, causing deadline misses, throughput collapse, or resource exhaustion in vehicle controllers, interactive services, or battery-powered sensors, sometimes while preserving the nominal prediction. This survey unifies a fragmented literature spanning perception pipelines (including physical attacks on autonomous-driving detection and tracking), input-adaptive neural inference (sponge examples, dynamic networks), and autoregressive and agentic systems (output-length, verbose-image, and reasoning denial-of-service attacks on LLMs, VLMs, mixture-of-experts models, and tool-using agents). We organize attacks by exploited computational bottleneck rather than formulation, separating what makes a computation expensive from how the attacker triggers it; the delivery channel (input, prompt or retrieved content, message, poisoning, or weight tampering) is an orthogonal attribute. Many attacks share one mechanism, intermediate-work amplification, motivating a work-budget defense abstraction; we distinguish caps on the work entering an expensive stage from caps on the results leaving it. We further analyze when a model-level cost increase becomes a system-level availability failure, which depends on critical-path share, slack, existing ceilings, accumulation, resource sharing, and fallback policy, not on the amplification factor alone. We also provide a threat-model taxonomy, consolidated quantitative comparisons, a defense review by control mechanism, and open challenges such as standardized evaluation, physical realizability, and whole-system availability. Companion website: https://github.com/guzonghua/awesome-latency-attacks.

cs.CR↗

TrackFlood: Relocating Latency Attacks from NMS-Free Detectors to Real-Time Trackers

We consider latency attacks on object detectors, where the attacker's goal is not to corrupt a prediction but to make the system fail to respond in time, targeting real-time applications such as autonomous driving. Modern object detectors eliminate Non-Maximum Suppression (NMS) through one-to-one assignment or set prediction, removing the classical detector-side latency bottleneck exploited by prior latency (``sponge'') attacks. We show that NMS-free does not mean latency-robust: this architectural change does not eliminate the attack surface but relocates it downstream to multi-object tracking, whose data-association cost remains content dependent. We present \emph{TrackFlood}, a unified white-box overload attack against NMS-free detect-then-track pipelines spanning both one-to-one detectors (YOLOv10 and YOLO26) and query-based detectors (RT-DETR). TrackFlood recovers differentiable confidence tensors and optimizes perturbations that flood the tracker with spatially distributed phantom detections while leaving detector inference unchanged. Evaluated entirely on an NVIDIA Jetson AGX Orin (TensorRT FP16), detector latency remains essentially constant, whereas tracker latency increases substantially. At a standard imperceptible budget ($L_\infty{=}8/255$), a universal perturbation produces clearly measurable tracker overload but only modest end-to-end slowdown, without deadline misses for the association-dominated trackers; a separate higher-budget stress test drives severe end-to-end slowdowns and sustained deadline misses. We further evaluate a lightweight, architecture-agnostic bounded-admission layer that caps the tracker workload and largely restores end-to-end latency, at a non-trivial cost in admitted clean detections. Our results demonstrate that evaluating NMS-free perception systems requires considering downstream tracking and end-to-end timing, not detector inference alone.

cs.CR↗