SearcharxivSearch

arXiv subjects

Thanh Le

Publications and source records attributed to Thanh Le.

12 recordsLinked to original sources

Retrieved But Not Reliable: A Survey on Attacks, and Defenses in Retrieval-Augmented Generation

Retrieval-Augmented Generation (RAG) enhances large language models by grounding outputs in external knowledge, improving factuality and reducing hallucinations. At the same time, the retrieval-augmented pipeline introduces new robustness and security risks, including corpus poisoning, backdoor attacks, privacy leakage, and fairness violations. Despite rapid progress in this area, existing surveys remain limited in their treatment of attacker objectives, threat models, and stage-specific defenses across the full RAG pipeline. This survey presents a unified and pipeline-aware overview of RAG robustness. We formalize threat models over the corpus, retriever, and generator, and organize attacks into three main objectives: accuracy, privacy, and fairness. We further review defenses from a pipeline-aware perspective, covering the retrieval, rerank, generation, and traceback stages. In addition, we summarize robustness benchmarks and explainability methods for more deeply evaluating and explaining RAG robustness.

cs.CR

Cost-Aware Uplink MPQUIC Scheduling via Multi-Objective Bayesian Optimization

Multipath QUIC (MPQUIC) enables simultaneous uplink transmission over heterogeneous access networks such as Wi-Fi and LTE, improving reliability and performance. However, aggressive LTE utilization increases operational cost, creating an inherent trade-off between upload delay and cellular usage. Existing MPQUIC schedulers typically optimize a single performance objective and operate at fixed points within this trade-off space, without explicitly supporting cost-aware operation. This paper formulates uplink MPQUIC scheduling as a multi-objective optimization problem that jointly considers maximum upload completion time and total LTE usage. We propose a Bayesian Optimization-based framework that treats the MPQUIC system as a black box and systematically explores probabilistic path selection configurations to uncover Pareto-efficient operating points. Rather than committing to a predefined scheduling policy, the framework exposes a spectrum of delay--cost trade-offs without modifying protocol internals. Experiments conducted using the Mininet-WiFi emulator show that the proposed approach characterizes a wide delay--cost region and identifies configurations that achieve substantial LTE savings (up to 80%) with controlled increases in upload time. The results further indicate that, under higher contention levels, systematic multi-objective exploration provides increased flexibility compared to fixed-policy schedulers in cost-aware heterogeneous uplink deployments.

cs.NI

Formal Verification for Deep Learning-based Power Control in Massive MIMO

Deep learning is a promising approach to optimize wireless communication by simplifying the search for near-optimal solutions. Prior studies on deep learning-based wireless communication optimization have explored supervised learning approaches that map raw user information, such as location or channel state information, to optimal power allocation vectors. While this approach demonstrates competitive performance, it is susceptible to adversarial attacks via input perturbations. Current defense mechanisms primarily rely on empirical methods, which do not provide formal guarantees of robustness. We fill this gap by proposing a formal verification framework to evaluate the robustness of deep learning-based power allocation in multi-cell massive multiple-input multiple-output (MIMO) systems against a wide range of potential adversarial input manipulations. To the best of our knowledge, this is the first attempt to formally verify deep neural networks in a regression setting with non-linear output constraints. We model the adversary's capabilities using hyper-rectangle constraints on their perturbation, adopt the abstraction-based bound-propagation technique (DeepPoly) to bound the interval of potential allocated powers, and formulate the minimum performance requirements as a constrained program for numerical feasibility analysis. Evaluation on publicly available datasets for power allocation in multi-cell massive MIMO indicates that a well-trained model can guarantee the local robustness under location perturbation by +-1m while retaining a maximum 1% optimality gap.

cs.NI

FGGM: Formal Grey-box Gradient Method for Attacking DRL-based MU-MIMO Scheduler

In 5G mobile communication systems, MU-MIMO has been applied to enhance spectral efficiency and support high data rates. To maximize spectral efficiency while providing fairness among users, the base station (BS) needs to selects a subset of users for data transmission. Given that this problem is NP-hard, DRL-based methods have been proposed to infer the near-optimal solutions in real-time, yet this approach has an intrinsic security problem. This paper investigates how a group of adversarial users can exploit unsanitized raw CSIs to launch a throughput degradation attack. Most existing studies only focused on systems in which adversarial users can obtain the exact values of victims' CSIs, but this is impractical in the case of uplink transmission in LTE/5G mobile systems. We note that the DRL policy contains an observation normalizer which has the mean and variance of the observation to improve training convergence. Adversarial users can then estimate the upper and lower bounds of the local observations including the CSIs of victims based solely on that observation normalizer. We develop an attacking scheme FGGM by leveraging polytope abstract domains, a technique used to bound the outputs of a neural network given the input ranges. Our goal is to find one set of intentionally manipulated CSIs which can achieve the attacking goals for the whole range of local observations of victims. Experimental results demonstrate that FGGM can determine a set of adversarial CSI vector controlled by adversarial users, then reuse those CSIs throughout the simulation to reduce the network throughput of a victim up to 70\% without knowing the exact value of victims' local observations. This study serves as a case study and can be applied to many other DRL-based problems, such as a knapsack-oriented resource allocation problems.

cs.NI

Verifying DNN-based Semantic Communication Against Generative Adversarial Noise

Safety-critical applications like autonomous vehicles and industrial IoT are adopting semantic communication (SemCom) systems using deep neural networks to reduce bandwidth and increase transmission speed by transmitting only task-relevant semantic features. However, adversarial attacks against these DNN-based SemCom systems can cause catastrophic failures by manipulating transmitted semantic features. Existing defense mechanisms rely on empirical approaches provide no formal guarantees against the full spectrum of adversarial perturbations. We present VSCAN, a neural network verification framework that provides mathematical robustness guarantees by formulating adversarial noise generation as mixed integer programming and verifying end-to-end properties across multiple interconnected networks (encoder, decoder, and task model). Our key insight is that realistic adversarial constraints (power limitations and statistical undetectability) can be encoded as logical formulae to enable efficient verification using state-of-the-art DNN verifiers. Our evaluation on 600 verification properties characterizing various attacker's capabilities shows VSCAN matches attack methods in finding vulnerabilities while providing formal robustness guarantees for 44% of properties -- a significant achievement given the complexity of multi-network verification. Moreover, we reveal a fundamental security-efficiency tradeoff: compact 16-dimensional latent spaces achieve 50% verified robustness compared to 64-dimensional spaces.

cs.LO

SwiftBrush v2: Make Your One-step Diffusion Model Better Than Its Teacher

In this paper, we aim to enhance the performance of SwiftBrush, a prominent one-step text-to-image diffusion model, to be competitive with its multi-step Stable Diffusion counterpart. Initially, we explore the quality-diversity trade-off between SwiftBrush and SD Turbo: the former excels in image diversity, while the latter excels in image quality. This observation motivates our proposed modifications in the training methodology, including better weight initialization and efficient LoRA training. Moreover, our introduction of a novel clamped CLIP loss enhances image-text alignment and results in improved image quality. Remarkably, by combining the weights of models trained with efficient LoRA and full training, we achieve a new state-of-the-art one-step diffusion model, achieving an FID of 8.14 and surpassing all GAN-based and multi-step Stable Diffusion models. The project page is available at https://swiftbrushv2.github.io.

cs.CV

PlanarTrack: A Large-scale Challenging Benchmark for Planar Object Tracking

Planar object tracking is a critical computer vision problem and has drawn increasing interest owing to its key roles in robotics, augmented reality, etc. Despite rapid progress, its further development, especially in the deep learning era, is largely hindered due to the lack of large-scale challenging benchmarks. Addressing this, we introduce PlanarTrack, a large-scale challenging planar tracking benchmark. Specifically, PlanarTrack consists of 1,000 videos with more than 490K images. All these videos are collected in complex unconstrained scenarios from the wild, which makes PlanarTrack, compared with existing benchmarks, more challenging but realistic for real-world applications. To ensure the high-quality annotation, each frame in PlanarTrack is manually labeled using four corners with multiple-round careful inspection and refinement. To our best knowledge, PlanarTrack, to date, is the largest and most challenging dataset dedicated to planar object tracking. In order to analyze the proposed PlanarTrack, we evaluate 10 planar trackers and conduct comprehensive comparisons and in-depth analysis. Our results, not surprisingly, demonstrate that current top-performing planar trackers degenerate significantly on the challenging PlanarTrack and more efforts are needed to improve planar tracking in the future. In addition, we further derive a variant named PlanarTrack$_{\mathbf{BB}}$ for generic object tracking from PlanarTrack. Our evaluation of 10 excellent generic trackers on PlanarTrack$_{\mathrm{BB}}$ manifests that, surprisingly, PlanarTrack$_{\mathrm{BB}}$ is even more challenging than several popular generic tracking benchmarks and more attention should be paid to handle such planar objects, though they are rigid. All benchmarks and evaluations will be released at the project webpage.

cs.CV

Ensemble Learning of Myocardial Displacements for Myocardial Infarction Detection in Echocardiography

Early detection and localization of myocardial infarction (MI) can reduce the severity of cardiac damage through timely treatment interventions. In recent years, deep learning techniques have shown promise for detecting MI in echocardiographic images. However, there has been no examination of how segmentation accuracy affects MI classification performance and the potential benefits of using ensemble learning approaches. Our study investigates this relationship and introduces a robust method that combines features from multiple segmentation models to improve MI classification performance by leveraging ensemble learning. Our method combines myocardial segment displacement features from multiple segmentation models, which are then input into a typical classifier to estimate the risk of MI. We validated the proposed approach on two datasets: the public HMC-QU dataset (109 echocardiograms) for training and validation, and an E-Hospital dataset (60 echocardiograms) from a local clinical site in Vietnam for independent testing. Model performance was evaluated based on accuracy, sensitivity, and specificity. The proposed approach demonstrated excellent performance in detecting MI. The results showed that the proposed approach outperformed the state-of-the-art feature-based method. Further research is necessary to determine its potential use in clinical settings as a tool to assist cardiologists and technicians with objective assessments and reduce dependence on operator subjectivity. Our research codes are available on GitHub at https://github.com/vinuni-vishc/mi-detection-echo.

cs.CV

Measurement of 2x2 LoS MIMO Terahertz Channel

This paper examines the performance of a 2x2 Line of Sight (LoS) Multiple Input Multiple Output (MIMO) channel at three terahertz frequencies-340 GHz, 410 Ghz, and 460 GHz. While theoretical models predict very high channel capacities, we observe lower capacity which is explained by asymmetric transmit-to-receive signal strengths as well as due to signal attenuation over longer distances. Overall, however, we note that at 460 Ghz, channel capacity of higher than 12 bps/hz is possible even at sub-optimal inter-antenna spacings (for different distances). An important observation is also that we need to maintain appropriate receive signal levels at receive antennas in order to improve capacity.

eess.SP

Reflection Channel Model for Terahertz Communications

Terahertz frequencies are an untapped resource for providing high-speed short-range communications. As a result, it is of interest to study the propagation characteristics of terahertz waves and to develop channel models. In previous work we used a measurement-based approach to develop an accurate channel model for line of sight (LoS) links. In this paper we extend that work by developing channel models for non-line of sight (NLoS) links where the signal suffers one reflection. We study reflections that occur off a metal plate as well as a piece of wood. Our model for received magnitude includes the effects of standing waves that develop between the transmitter and receiver. Measurements show an excellent agreement between empirical data and the model. In addition, we have analyzed the received phase of the reflected signal at frequencies in the range 320- 480 GHz. We observed a linear error between the predicted and actual phase and developed a model to accommodate that discrepancy. The final model we have developed for predicting received phase is very accurate for the entire range 320 - 480 GHz and for both materials.

eess.SP

Effect of Standing Wave on Terahertz Channel Model

There is a growing interest in exploiting the terahertz frequency band for future communication systems that demand high data rates. Given the complex propagation behavior of this frequency band, various researchers have developed channel models that can be utilized in the development of communication systems. These models however do not include a crucial aspect of terahertz propagation at short distances: the presence of standing waves. Our measurements show that at specific distances, the effect of standing waves is significant. In this paper, we extend previous terahertz channel models to include the effect of standing waves and show a good fit with our measurements. Our measurements and modeling cover the five most promising terahertz frequency bands: 140, 220, 340, 410, 460 GHz.

eess.SP

The Dynamical Gaussian Process Latent Variable Model in the Longitudinal Scenario

The Dynamical Gaussian Process Latent Variable Models provide an elegant non-parametric framework for learning the low dimensional representations of the high-dimensional time-series. Real world observational studies, however, are often ill-conditioned: the observations can be noisy, not assuming the luxury of relatively complete and equally spaced like those in time series. Such conditions make it difficult to learn reasonable representations in the high dimensional longitudinal data set by way of Gaussian Process Latent Variable Model as well as other dimensionality reduction procedures. In this study, we approach the inference of Gaussian Process Dynamical Systems in Longitudinal scenario by augmenting the bound in the variational approximation to include systematic samples of the unseen observations. We demonstrate the usefulness of this approach on synthetic as well as the human motion capture data set.

cs.LG