SearcharxivSearch

arXiv subjects

Xianglong Du

Publications and source records attributed to Xianglong Du.

5 recordsLinked to original sources

When Sample Selection Bias Precipitates Model Collapse

The proliferation of recursive training on synthetic data can alleviate data scarcity but risks model collapse, where repeated training erodes distributional tails and homogenizes outputs. Data selection is widely viewed as a remedy, yet its reliability depends critically on the reference distribution used by the verifier. We show that in low-resource verification regimes, where each verifier observes only a small, fragmented, and biased slice of the target manifold, selection itself becomes biased. This situation naturally arises in low-resource data silos such as healthcare consortia or proprietary financial institutions, where raw data cannot be pooled and local references are inherently incomplete. As a result, selection preferentially retains samples aligned with the local manifold while pruning globally relevant tail modes, turning from a safeguard against collapse into a mechanism that precipitates it. We theoretically prove that such siloed selection accelerates collapse and induces power-law diversity decay. As an initial mitigation, we construct Wasserstein proxy references from multiple silos without sharing raw data. Empirical results confirm that local-reference selection fails on skewed distributions, whereas collaborative proxy references mitigate diversity degradation, suggesting that recursive synthetic-data pipelines require particular caution when real-data coverage is fragmented or scarce.

cs.AI

SharedRequest: Privacy-Preserving Model-Agnostic Inference for Large Language Models

With the widespread deployment of public large language models (LLMs) such as ChatGPT, protecting user prompt privacy has become an increasingly critical issue. Existing privacy-preserving inference methods sacrifice either utility or efficiency, and often require model-specific modifications that limit their compatibility. In this paper, we propose SharedRequest, a model-agnostic framework for privacy-preserving LLM inference that reformulates privacy protection at the batch level rather than the individual-prompt level. The key idea is to obscure sensitive information by mixing original prompts with noisy variants, while grouping semantically equivalent instructions to amortize the inference cost over a large batch of queries with minimal impact on LLM response quality. This design is independent of the LLM architecture, requiring no access to model parameters or architectural modification. Empirical results demonstrate that SharedRequest achieves over $20\%$ higher utility compared to prior differential privacy baselines, and its shared-prompt mechanism reduces query cost by up to $5\times$ compared to non-batched inference.

cs.CR

Towards Stealthy and Effective Backdoor Attacks on Lane Detection: A Naturalistic Data Poisoning Approach

Deep learning-based lane detection (LD) plays a critical role in autonomous driving and advanced driver assistance systems. However, its vulnerability to backdoor attacks presents a significant security concern. Existing backdoor attack methods on LD often exhibit limited practical utility due to the artificial and conspicuous nature of their triggers. To address this limitation and investigate the impact of more ecologically valid backdoor attacks on LD models, we examine the common data poisoning attack and introduce DBALD, a novel diffusion-based data poisoning framework for generating naturalistic backdoor triggers. DBALD comprises two key components: optimal trigger position finding and stealthy trigger generation. Given the insight that attack performance varies depending on the trigger position, we propose a heatmap-based method to identify the optimal trigger location, with gradient analysis to generate attack-specific heatmaps. A region-based editing diffusion process is then applied to synthesize visually plausible triggers within the most susceptible regions identified previously. Furthermore, to ensure scene integrity and stealthy attacks, we introduce two loss strategies: one for preserving lane structure and another for maintaining the consistency of the driving scene. Consequently, compared to existing attack methods, DBALD achieves both a high attack success rate and superior stealthiness. Extensive experiments on 4 mainstream LD models show that DBALD exceeds state-of-the-art methods, with an average success rate improvement of +10.87% and significantly enhanced stealthiness. The experimental results highlight significant practical challenges in ensuring model robustness against real-world backdoor threats in LD.

cs.CR

Machine Learning Accelerated Computational Surface-Specific Vibrational Spectroscopy Reveals Oxidation Level of Graphene in Contact with Water

Precise characterization of the graphene/water interface has been hindered by experimental inconsistencies and limited molecular-level access to interfacial structures. In this work, we present a novel integrated computational approach that combines machine-learning-driven molecular dynamics simulations with first-principles vibrational spectroscopy calculations to reveal how graphene oxidation alters interfacial water structure. Our simulations demonstrate that pristine graphene leaves the hydrogen-bond network of interfacial water largely unperturbed, whereas graphene oxide (GO) with surface hydroxyls induces a pronounced $\Delta f \sim 100 cm^{-1}$ redshift of the free-OH vibrational band and a dramatic reduction in its amplitude. These spectral shifts in the computed surface-specific sum-frequency generation spectrum serve as sensitive molecular markers of the GO oxidation level, reconciling previously conflicting experimental observations. By providing a quantitative spectroscopic fingerprint of GO oxidation, our findings have broad implications for catalysis and electrochemistry, where the structuring of interfacial water is critical to performance.

physics.chem-ph

Revealing the molecular structures of a-Al2O3(0001)-water interface by machine learning based computational vibrational spectroscopy

Solid-water interfaces are crucial to many physical and chemical processes and are extensively studied using surface-specific sum-frequency generation (SFG) spectroscopy. To establish clear correlations between specific spectral signatures and distinct interfacial water structures, theoretical calculations using molecular dynamics (MD) simulations are required. These MD simulations typically need relatively long trajectories (a few nanoseconds) to achieve reliable SFG response function calculations via the dipole-polarizability time correlation function. However, the requirement for long trajectories limits the use of computationally expensive techniques such as ab initio MD (AIMD) simulations, particularly for complex solid-water interfaces. In this work, we present a pathway for calculating vibrational spectra (IR, Raman, SFG) of solid-water interfaces using machine learning (ML)-accelerated methods. We employ both the dipole moment-polarizability correlation function and the surface-specific velocity-velocity correlation function approaches to calculate SFG spectra. Our results demonstrate the successful acceleration of AIMD simulations and the calculation of SFG spectra using ML methods. This advancement provides an opportunity to calculate SFG spectra for the complicated solid-water systems more rapidly and at a lower computational cost with the aid of ML.

cond-mat.mtrl-sci