SearcharxivSearch

arXiv subjects

Xiaodan Li

Publications and source records attributed to Xiaodan Li.

At least 19 recordsLinked to original sources

Weak Equilibrium Measures and Capacity--Hitting Identities for the Hypoelliptic Third-Order Langevin Diffusion

We construct weak equilibrium measures and weak capacities for the hypoelliptic third-order Langevin diffusion motivated by an accelerated sampling algorithm (Mou et al. (2021) \textit{J. Mach. Learn. Res.}, \textbf{22}(42), 1--41). In this process, the Brownian noise acts only in the highest-order auxiliary variable and reaches the physical variables through a step-three H\"ormander chain, so the standard uniformly elliptic boundary-flux theory is not directly applicable at characteristic points of phase-space balls. We prove an elliptic-regularization stability theorem for the corresponding hitting laws and then define the weak equilibrium measure and weak capacity. The proof combines the boundary-hitting stability strategy of Lee--Ramil--Seo (2026, \textit{arXiv:2503.12610v2}) with localized hypoelliptic heat-kernel estimates (Pigato (2022) \textit{Stoch. Process. Appl.}, \textbf{145}, 117--142) adapted to the third-order chain. We obtain the bounded-domain weak capacity--hitting identity and a Lyapunov drift argument in the spirit of Lee--Ramil--Seo that yields positive Harris recurrence and extends the construction to a whole-space weak equilibrium measure, and whole-space capacity--hitting identity.

math.PR

From Cannings model to Brownian motion conditioned on local time profile

We study the scaling limits of genealogical trees arising from Cannings models. Under suitable moment conditions, we show that the rescaled contour and height functions converge to a time change of Brownian motion conditioned on a given local time profile. This conditioned Brownian motion is a self-interacting diffusion constructed independently by Warren--Yor (1998) and Aldous (1998). A key ingredient in our proof is a sequential version of the coming-down-from-infinity property.

math.PR

Kelp: A Streaming Safeguard for Large Models via Latent Dynamics-Guided Risk Detection

Large models (LMs) are powerful content generators, yet their open-ended nature can also introduce potential risks, such as generating harmful or biased content. Existing guardrails mostly perform post-hoc detection that may expose unsafe content before it is caught, and the latency constraints further push them toward lightweight models, limiting detection accuracy. In this work, we propose Kelp, a novel plug-in framework that enables streaming risk detection within the LM generation pipeline. Kelp leverages intermediate LM hidden states through a Streaming Latent Dynamics Head (SLD), which models the temporal evolution of risk across the generated sequence for more accurate real-time risk detection. To ensure reliable streaming moderation in real applications, we introduce an Anchored Temporal Consistency (ATC) loss to enforce monotonic harm predictions by embedding a benign-then-harmful temporal prior. Besides, for a rigorous evaluation of streaming guardrails, we also present StreamGuardBench-a model-grounded benchmark featuring on-the-fly responses from each protected model, reflecting real-world streaming scenarios in both text and vision-language tasks. Across diverse models and datasets, Kelp consistently outperforms state-of-the-art post-hoc guardrails and prior plug-in probes (15.61% higher average F1), while using only 20M parameters and adding less than 0.5 ms of per-token latency.

cs.LG

Refiner: Data Refining against Gradient Leakage Attacks in Federated Learning

Recent works have brought attention to the vulnerability of Federated Learning (FL) systems to gradient leakage attacks. Such attacks exploit clients' uploaded gradients to reconstruct their sensitive data, thereby compromising the privacy protection capability of FL. In response, various defense mechanisms have been proposed to mitigate this threat by manipulating the uploaded gradients. Unfortunately, empirical evaluations have demonstrated limited resilience of these defenses against sophisticated attacks, indicating an urgent need for more effective defenses. In this paper, we explore a novel defensive paradigm that departs from conventional gradient perturbation approaches and instead focuses on the construction of robust data. Intuitively, if robust data exhibits low semantic similarity with clients' raw data, the gradients associated with robust data can effectively obfuscate attackers. To this end, we design Refiner that jointly optimizes two metrics for privacy protection and performance maintenance. The utility metric is designed to promote consistency between the gradients of key parameters associated with robust data and those derived from clients' data, thus maintaining model performance. Furthermore, the privacy metric guides the generation of robust data towards enlarging the semantic gap with clients' data. Theoretical analysis supports the effectiveness of Refiner, and empirical evaluations on multiple benchmark datasets demonstrate the superior defense effectiveness of Refiner at defending against state-of-the-art attacks.

cs.LG

Comprehensive Evaluation and Analysis for NSFW Concept Erasure in Text-to-Image Diffusion Models

Text-to-image diffusion models have gained widespread application across various domains, demonstrating remarkable creative potential. However, the strong generalization capabilities of diffusion models can inadvertently lead to the generation of not-safe-for-work (NSFW) content, posing significant risks to their safe deployment. While several concept erasure methods have been proposed to mitigate the issue associated with NSFW content, a comprehensive evaluation of their effectiveness across various scenarios remains absent. To bridge this gap, we introduce a full-pipeline toolkit specifically designed for concept erasure and conduct the first systematic study of NSFW concept erasure methods. By examining the interplay between the underlying mechanisms and empirical observations, we provide in-depth insights and practical guidance for the effective application of concept erasure methods in various real-world scenarios, with the aim of advancing the understanding of content safety in diffusion models and establishing a solid foundation for future research and development in this critical area.

cs.CV

Many-body localization properties of one-dimensional anisotropic spin-1/2 chains

In this paper, we theoretically investigate the many-body localization (MBL) properties of one-dimensional anisotropic spin-1/2 chains by using the exact matrix diagonalization method. Starting from the Ising spin-1/2 chain, we introduce different forms of external fields and spin coupling interactions, and construct three distinct anisotropic spin-1/2 chain models. The influence of these interactions on the MBL phase transition is systematically explored. We first analyze the eigenstate properties by computing the excited-state fidelity. The results show that MBL phase transitions occur in all three models, and that both the anisotropy parameter and the finite system size significantly affect the critical disorder strength of the transition. Moreover, we calculated the bipartite entanglement entropy of the system, and the critical points determined by the intersection of curves for different system sizes are basically consistent with those obtained from the excited-state fidelity. Then, the dynamical characteristics of the systems are studied through the time evolution of diagonal entropy (DE), local magnetization, and fidelity. These observations further confirm the occurrence of the MBL phase transition and allow for a clear distinction between the ergodic (thermal) phase and the many-body localized phase. Finally, to examine the effect of additional interactions on the transition, we incorporate Dzyaloshinskii-Moriya (DM) interactions into the three models. The results demonstrate that the MBL phase transition still occurs in the presence of DM interactions. However, the anisotropy parameter and finite system size significantly affect the critical disorder strength. Moreover, the critical behavior is somewhat suppressed, indicating that DM interactions tend to inhibit the onset of localization.

cond-mat.dis-nn

Comprehensive Assessment and Analysis for NSFW Content Erasure in Text-to-Image Diffusion Models

Text-to-image diffusion models have gained widespread application across various domains, demonstrating remarkable creative potential. However, the strong generalization capabilities of diffusion models can inadvertently lead to the generation of not-safe-for-work (NSFW) content, posing significant risks to their safe deployment. While several concept erasure methods have been proposed to mitigate the issue associated with NSFW content, a comprehensive evaluation of their effectiveness across various scenarios remains absent. To bridge this gap, we introduce a full-pipeline toolkit specifically designed for concept erasure and conduct the first systematic study of NSFW concept erasure methods. By examining the interplay between the underlying mechanisms and empirical observations, we provide in-depth insights and practical guidance for the effective application of concept erasure methods in various real-world scenarios, with the aim of advancing the understanding of content safety in diffusion models and establishing a solid foundation for future research and development in this critical area.

cs.CV

Transferability Bound Theory: Exploring Relationship between Adversarial Transferability and Flatness

A prevailing belief in attack and defense community is that the higher flatness of adversarial examples enables their better cross-model transferability, leading to a growing interest in employing sharpness-aware minimization and its variants. However, the theoretical relationship between the transferability of adversarial examples and their flatness has not been well established, making the belief questionable. To bridge this gap, we embark on a theoretical investigation and, for the first time, derive a theoretical bound for the transferability of adversarial examples with few practical assumptions. Our analysis challenges this belief by demonstrating that the increased flatness of adversarial examples does not necessarily guarantee improved transferability. Moreover, building upon the theoretical analysis, we propose TPA, a Theoretically Provable Attack that optimizes a surrogate of the derived bound to craft adversarial examples. Extensive experiments across widely used benchmark datasets and various real-world applications show that TPA can craft more transferable adversarial examples compared to state-of-the-art baselines. We hope that these results can recalibrate preconceived impressions within the community and facilitate the development of stronger adversarial attack and defense mechanisms. The source codes are available in .

cs.LG

The Harmonic Descent Chain

The decreasing Markov chain on \{1,2,3, \ldots\} with transition probabilities $p(j,j-i) \propto 1/i$ arises as a key component of the analysis of the beta-splitting random tree model. We give a direct and almost self-contained "probability" treatment of its occupation probabilities, as a counterpart to a more sophisticated but perhaps opaque derivation using a limit continuum tree structure and Mellin transforms.

math.PR

Many-body Localization Transition of Ising Spin-1 Chains

In this paper, we theoretically investigate the many-body localization properties of one-dimensional Ising spin-1 chains by using the methods of exact matrix diagonalization. We compare it with the MBL properties of the Ising spin-1/2 chains. The results indicate that the one-dimensional Ising spin-1 chains can also undergo MBL phase transition. There are various forms of disorder, and we compare the effects of different forms of quasi-disorder and random disorder on many-body localization in this paper. First, we calculate the exctied-state fidelity to study the MBL phase transtion. By changing the form of the quasi-disorder, we study the MBL transition of the system with different forms of quasi-disorder and compare them with those of the random disordered system. The results show that both random disorder and quasi-disorder can cause the MBL phase transition in the one-dimensional Ising spin-1 chains. In order to study the effect of spin interactions, we compare Ising spin-1 chains and spin-1/2 chains with the next-nearest-neighbour(N-N) two-body interactions and the next-next-nearest-neighbour (N-N-N)interactions. The results show that the critical point increases with the addition of the interaction. Then we study the dynamical properties of the model by the dynamical behavior of diagonal entropy (DE), local magnetization and the time evolution of fidelity to further prove the occurrence of MBL phase transition in the disordered Ising spin-1 chains with the (N-N) coupling term and distinguish the ergodic phase (thermal phase) and the many-body localized phase. Lastly, we delve into the impact of periodic driving on one-dimensional Ising spin-1 chains. And we compare it with the results obtained from the Ising spin-1/2 chains. It shows that periodic driving can cause Ising spin-1 chains and Ising spin-1/2 chains to occur the MBL transition.

cond-mat.dis-nn

Model Inversion Attack via Dynamic Memory Learning

Model Inversion (MI) attacks aim to recover the private training data from the target model, which has raised security concerns about the deployment of DNNs in practice. Recent advances in generative adversarial models have rendered them particularly effective in MI attacks, primarily due to their ability to generate high-fidelity and perceptually realistic images that closely resemble the target data. In this work, we propose a novel Dynamic Memory Model Inversion Attack (DMMIA) to leverage historically learned knowledge, which interacts with samples (during the training) to induce diverse generations. DMMIA constructs two types of prototypes to inject the information about historically learned knowledge: Intra-class Multicentric Representation (IMR) representing target-related concepts by multiple learnable prototypes, and Inter-class Discriminative Representation (IDR) characterizing the memorized samples as learned prototypes to capture more privacy-related information. As a result, our DMMIA has a more informative representation, which brings more diverse and discriminative generated results. Experiments on multiple benchmarks show that DMMIA performs better than state-of-the-art MI attack methods.

cs.CV

A note on $α$-permanent and loop soup

In this paper, it is shown that $α$-permanent in algebra is closely related to loop soup in probability. We give explicit expansions of $α$-permanents of the block matrices obtained from matrices associated to $*$-forests, which are a special class of matrices containing tridiagonal matrices. It is proved in two ways, one is the direct combinatorial proof, and the other is the probabilistic proof via loop soup.

math.PR

ImageNet-E: Benchmarking Neural Network Robustness via Attribute Editing

Recent studies have shown that higher accuracy on ImageNet usually leads to better robustness against different corruptions. Therefore, in this paper, instead of following the traditional research paradigm that investigates new out-of-distribution corruptions or perturbations deep models may encounter, we conduct model debugging in in-distribution data to explore which object attributes a model may be sensitive to. To achieve this goal, we create a toolkit for object editing with controls of backgrounds, sizes, positions, and directions, and create a rigorous benchmark named ImageNet-E(diting) for evaluating the image classifier robustness in terms of object attributes. With our ImageNet-E, we evaluate the performance of current deep learning models, including both convolutional neural networks and vision transformers. We find that most models are quite sensitive to attribute changes. A small change in the background can lead to an average of 9.23\% drop on top-1 accuracy. We also evaluate some robust models including both adversarially trained models and other robust trained models and find that some models show worse robustness against attribute changes than vanilla models. Based on these findings, we discover ways to enhance attribute robustness with preprocessing, architecture designs, and training strategies. We hope this work can provide some insights to the community and open up a new avenue for research in robust computer vision. The code and dataset are available at https://github.com/alibaba/easyrobust.

cs.CV

TransAudio: Towards the Transferable Adversarial Audio Attack via Learning Contextualized Perturbations

In a transfer-based attack against Automatic Speech Recognition (ASR) systems, attacks are unable to access the architecture and parameters of the target model. Existing attack methods are mostly investigated in voice assistant scenarios with restricted voice commands, prohibiting their applicability to more general ASR related applications. To tackle this challenge, we propose a novel contextualized attack with deletion, insertion, and substitution adversarial behaviors, namely TransAudio, which achieves arbitrary word-level attacks based on the proposed two-stage framework. To strengthen the attack transferability, we further introduce an audio score-matching optimization strategy to regularize the training process, which mitigates adversarial example over-fitting to the surrogate model. Extensive experiments and analysis demonstrate the effectiveness of TransAudio against open-source ASR models and commercial APIs.

cs.SD

Information-containing Adversarial Perturbation for Combating Facial Manipulation Systems

With the development of deep learning technology, the facial manipulation system has become powerful and easy to use. Such systems can modify the attributes of the given facial images, such as hair color, gender, and age. Malicious applications of such systems pose a serious threat to individuals' privacy and reputation. Existing studies have proposed various approaches to protect images against facial manipulations. Passive defense methods aim to detect whether the face is real or fake, which works for posterior forensics but can not prevent malicious manipulation. Initiative defense methods protect images upfront by injecting adversarial perturbations into images to disrupt facial manipulation systems but can not identify whether the image is fake. To address the limitation of existing methods, we propose a novel two-tier protection method named Information-containing Adversarial Perturbation (IAP), which provides more comprehensive protection for {facial images}. We use an encoder to map a facial image and its identity message to a cross-model adversarial example which can disrupt multiple facial manipulation systems to achieve initiative protection. Recovering the message in adversarial examples with a decoder serves passive protection, contributing to provenance tracking and fake image detection. We introduce a feature-level correlation measurement that is more suitable to measure the difference between the facial images than the commonly used mean squared error. Moreover, we propose a spectral diffusion method to spread messages to different frequency channels, thereby improving the robustness of the message against facial manipulation. Extensive experimental results demonstrate that our proposed IAP can recover the messages from the adversarial examples with high average accuracy and effectively disrupt the facial manipulation systems.

cs.CV

Rethinking Out-of-Distribution Detection From a Human-Centric Perspective

Out-Of-Distribution (OOD) detection has received broad attention over the years, aiming to ensure the reliability and safety of deep neural networks (DNNs) in real-world scenarios by rejecting incorrect predictions. However, we notice a discrepancy between the conventional evaluation vs. the essential purpose of OOD detection. On the one hand, the conventional evaluation exclusively considers risks caused by label-space distribution shifts while ignoring the risks from input-space distribution shifts. On the other hand, the conventional evaluation reward detection methods for not rejecting the misclassified image in the validation dataset. However, the misclassified image can also cause risks and should be rejected. We appeal to rethink OOD detection from a human-centric perspective, that a proper detection method should reject the case that the deep model's prediction mismatches the human expectations and adopt the case that the deep model's prediction meets the human expectations. We propose a human-centric evaluation and conduct extensive experiments on 45 classifiers and 8 test datasets. We find that the simple baseline OOD detection method can achieve comparable and even better performance than the recently proposed methods, which means that the development in OOD detection in the past years may be overestimated. Additionally, our experiments demonstrate that model selection is non-trivial for OOD detection and should be considered as an integral of the proposed method, which differs from the claim in existing works that proposed methods are universal across different models.

cs.CV

Enhanced compound-protein binding affinity prediction by representing protein multimodal information via a coevolutionary strategy

Due to the lack of a method to efficiently represent the multimodal information of a protein, including its structure and sequence information, predicting compound-protein binding affinity (CPA) still suffers from low accuracy when applying machine learning methods. To overcome this limitation, in a novel end-to-end architecture (named FeatNN), we develop a coevolutionary strategy to jointly represent the structure and sequence features of proteins and ultimately optimize the mathematical models for predicting CPA. Furthermore, from the perspective of data-driven approach, we proposed a rational method that can utilize both high- and low-quality databases to optimize the accuracy and generalization ability of FeatNN in CPA prediction tasks. Notably, we visually interpret the feature interaction process between sequence and structure in the rationally designed architecture. As a result, FeatNN considerably outperforms the state-of-the-art (SOTA) baseline in virtual drug screening tasks, indicating the feasibility of this approach for practical use. FeatNN provides an outstanding method for higher CPA prediction accuracy and better generalization ability by efficiently representing multimodal information of proteins via a coevolutionary strategy.

q-bio.BM

Inverting Ray-Knight identities on trees

In this paper, we first introduce the Ray-Knight identity and percolation Ray-Knight identity related to loop soup with intensity $α(\ge 0)$ on trees. Then we present the inversions of the above identities, which are expressed in terms of repelling jump processes. In particular, the inversion in the case of $α=0$ gives the conditional law of a Markov jump process given its local time field. We further show that the fine mesh limits of these repelling jump processes are the self-repelling diffusions \cite{Aidekon} involved in the inversion of the Ray-Knight identity on the corresponding metric graph. This is a generalization of results in \cite{2016Inverting,lupu2019inverting,LupuEJP657}, where the authors explore the case of $α=1/2$ on a general graph. Our construction is different from \cite{2016Inverting,lupu2019inverting} and based on the link between random networks and loop soups.

math.PR