SearcharxivSearch

arXiv subjects

Xiaolei Dong

Publications and source records attributed to Xiaolei Dong.

4 recordsLinked to original sources

DoubtProbe: Black-Box Jailbreak Defense via Structural Verification and Semantic Auditing

As large language models (LLMs) are increasingly deployed in user-facing systems, black-box jailbreak defense has become an important practical problem. Existing defenses often rely on known-attack coverage, prompt-level semantic judgment, or local runtime control, yet these paths can become unstable under evolving prompt packaging, expression rewriting, and structure manipulation. We observe that many black-box jailbreaks do not remove the harmful goal, but reorganize the information needed to express and execute it, thereby evading safety alignment while remaining recoverable during generation. Motivated by this observation, we propose DoubtProbe, a dual-branch inference-time defense framework that combines structural verification with semantic auditing and formulates black-box jailbreak defense as consistency checking under controlled transformation. The structural branch extracts a structured representation from the original request, reconstructs the request under representation constraints, and detects information-preservation failures between the original and reconstructed requests; the semantic branch audits the original prompt directly. We evaluate DoubtProbe against representative black-box defenses on jailbreak and benign-request benchmarks, and further test backbone transfer from Qwen2.5-72B to Llama-3.1-70B. Results show that DoubtProbe achieves a stronger and more stable defense-utility trade-off: on Qwen2.5-72B, it reduces the JBB attack success rate from 0.293 to 0.100 and the CodeAttack attack success rate from 0.152 to 0.001, while maintaining false positive rates of 0.022 and 0.016 on AlpacaEval and OR-Bench; the same pattern remains stable on Llama-3.1-70B. These findings show that structural inconsistency signals provide a practical and generalizable basis for black-box jailbreak defense, especially when combined with semantic auditing.

cs.CR

FedRW: Efficient Privacy-Preserving Data Reweighting for Enhancing Federated Learning of Language Models

Data duplication within large-scale corpora often impedes large language models' (LLMs) performance and privacy. In privacy-concerned federated learning scenarios, conventional deduplication methods typically rely on trusted third parties to perform uniform deletion, risking loss of informative samples while introducing privacy vulnerabilities. To address these gaps, we propose Federated ReWeighting (FedRW), the first privacy-preserving framework, to the best of our knowledge, that performs soft deduplication via sample reweighting instead of deletion in federated LLM training, without assuming a trusted third party. At its core, FedRW proposes a secure, frequency-aware reweighting protocol through secure multi-party computation, coupled with a parallel orchestration strategy to ensure efficiency and scalability. During training, FedRW utilizes an adaptive reweighting mechanism with global sample frequencies to adjust individual loss contributions, effectively improving generalization and robustness. Empirical results demonstrate that FedRW outperforms the state-of-the-art method by achieving up to 28.78x speedup in preprocessing and approximately 11.42% improvement in perplexity, while offering enhanced security guarantees. FedRW thus establishes a new paradigm for managing duplication in federated LLM training.

cs.CR

Prandtl Equations and Related Boundary Layer Equations

This book aims to present some recent results on Prandtl equations and MHD boundary layer equations. This book is essentially divided into two parts. Chapter 1 as the first part systematically surveys the results till 2020 on Prandtl equations and MHD boundary layer equations. Chapter 2 to 6 are the main part of the book, which presents the local and the global well-posedness of solutions to the Prandtl equations and MHD boundary layer equations. In detail, Chapter 2 is concerned with global well-posedness of solutions to the 2D Prandtl-Hartmann equations in an analytic framework. Chapter 3 investigates the local existence of solutions to the 2D Prandtl equations in a weighted Sobolev space. Chapter 4 studies the local well-posedness of solutions to the 2D mixed Prandtl equations in a Sobolev space without monotonicity and lower bound. Chapter 5 is concerned with global existence of solutions to the 2D magnetic Prandtl equations in the Prandtl-Hartmann regime. Chapter 6 proves the local existence of solutions to the 3D Prandtl equations with a special structure. Mathematicians and physicists who are interested in fluid dynamics will find this book helpful.

math.AP

Strong global attractors for a three dimensional nonclassical diffusion equation with memory

In this paper, we study the strong global attractors for a three dimensional nonclassical diffusion equation with memory. First, we prove the existence and uniqueness of strong solutions for the equations by the Galerkin method. Then we prove the existence of global attractors for the equations in $H^2(Ω)\cap H^1_0(Ω)\times L^2_μ(\mathbb{R}^+;H^2(Ω)\cap H^1_0(Ω))$ by the condition (C).

math.AP