SearcharxivSearch

arXiv subjects

Yiping Zhang

Publications and source records attributed to Yiping Zhang.

At least 19 recordsLinked to original sources

Sign-preserving solutions to the Tzitz\'eica equation on lattice graphs

On the lattice graph $\mathbb{Z}^n$, we establish the existence of sign-preserving solutions to the Tzitz\'eica equation. We prove the existence of positive solutions and two classes of negative solutions under different assumptions, and derive their decay estimates. The proof is based on a suitable approximation scheme, the monotone convergence theorem, and an exhaustion argument. These results extend those of Hua, Huang, and Wang (Anal. PDE, 2026) by establishing the existence of both positive and negative sign-preserving solutions together with their decay estimates.

math.AP

The lifespan of positive solutions of heat equation with power-logarithmic nonlinearity on locally finite graph

On a locally finite connected graph $G=(V,E)$, using the first eigenvalue method introduced by Kaplan \cite{MR160044} and the discrete Phragm\'{e}n-Lindel\"{o}f principle developed by Hu-Wang \cite{cvhuyuanyang}, we first establish the asymptotic behaviour of the lifespan of positive solutions to a semilinear heat equation with the power-logarithmic nonlinearity $u^p|\log u|^q$, provided that the initial datum is bounded below by a positive constant. These results extend those of Hu-Wang \cite{cvhuyuanyang} to equations with a power-logarithmic source term. Moreover, by means of a more direct argument, we show that analogous lifespan estimates remain valid for nonnegative initial datum $u(x,0)$, provided that $u(x_i,0)$ is suitably large at some vertex $x_i\in V$.

math.AP

DSBA: Dynamic Stealthy Backdoor Attack with Collaborative Optimization in Self-Supervised Learning

Self-Supervised Learning (SSL) has emerged as a significant paradigm in representation learning thanks to its ability to learn without extensive labeled data, its strong generalization capabilities, and its potential for privacy preservation. However, recent research reveals that SSL models are also vulnerable to backdoor attacks. Existing backdoor attack methods in the SSL context commonly suffer from issues such as high detectability of triggers, feature entanglement, and pronounced out-of-distribution properties in poisoned samples, all of which compromises attack effectiveness and stealthiness. To that, we propose a Dynamic Stealthy Backdoor Attack (DSBA) backed by a new technique we term Collaborative Optimization. This method decouples the attack process into two collaborative optimization layers: the outer-layer optimization trains a backdoor encoder responsible for global feature space remodeling, aiming to achieve precise backdoor implantation while preserving core functionality; meanwhile, the inner-layer optimization employs a dynamically optimized generator to adaptively produce optimally concealed triggers for individual samples, achieving coordinated concealment across feature space and visual space. We also introduce multiple loss functions to dynamically balance attack performance and stealthiness, in which we employ an adaptive weight scheduling mechanism to enhance training stability. Extensive experiments on various mainstream SSL algorithms and five public datasets demonstrate that: (i) DSBA significantly enhances Attack Success Rate (ASR) and stealthiness while maintaining downstream task accuracy; and (ii) DSBA exhibits superior robustness against existing mainstream defense methods.

cs.CR

BadRSSD: Backdoor Attacks on Regularized Self-Supervised Diffusion Models

Self-supervised diffusion models learn high-quality visual representations via latent space denoising. However, their representation layer poses a distinct threat: unlike traditional attacks targeting generative outputs, its unconstrained latent semantic space allows for stealthy backdoors, permitting malicious control upon triggering. In this paper, we propose BadRSSD, the first backdoor attack targeting the representation layer of self-supervised diffusion models. Specifically, it hijacks the semantic representations of poisoned samples with triggers in Principal Component Analysis (PCA) space toward those of a target image, then controls the denoising trajectory during diffusion by applying coordinated constraints across latent, pixel, and feature distribution spaces to steer the model toward generating the specified target. Additionally, we integrate representation dispersion regularization into the constraint framework to maintain feature space uniformity, significantly enhancing attack stealth. This approach preserves normal model functionality (high utility) while achieving precise target generation upon trigger activation (high specificity). Experiments on multiple benchmark datasets demonstrate that BadRSSD substantially outperforms existing attacks in both FID and MSE metrics, reliably establishing backdoors across different architectures and configurations, and effectively resisting state-of-the-art backdoor defenses.

cs.CR

ADCA: Attention-Driven Multi-Party Collusion Attack in Federated Self-Supervised Learning

Federated Self-Supervised Learning (FSSL) integrates the privacy advantages of distributed training with the capability of self-supervised learning to leverage unlabeled data, showing strong potential across applications. However, recent studies have shown that FSSL is also vulnerable to backdoor attacks. Existing attacks are limited by their trigger design, which typically employs a global, uniform trigger that is easily detected, gets diluted during aggregation, and lacks robustness in heterogeneous client environments. To address these challenges, we propose the Attention-Driven multi-party Collusion Attack (ADCA). During local pre-training, malicious clients decompose the global trigger to find optimal local patterns. Subsequently, these malicious clients collude to form a malicious coalition and establish a collaborative optimization mechanism within it. In this mechanism, each submits its model updates, and an attention mechanism dynamically aggregates them to explore the best cooperative strategy. The resulting aggregated parameters serve as the initial state for the next round of training within the coalition, thereby effectively mitigating the dilution of backdoor information by benign updates. Experiments on multiple FSSL scenarios and four datasets show that ADCA significantly outperforms existing methods in Attack Success Rate (ASR) and persistence, proving its effectiveness and robustness.

cs.CR

CoT is Not the Chain of Truth: An Empirical Internal Analysis of Reasoning LLMs for Fake News Generation

From generating headlines to fabricating news, the Large Language Models (LLMs) are typically assessed by their final outputs, under the safety assumption that a refusal response signifies safe reasoning throughout the entire process. Challenging this assumption, our study reveals that during fake news generation, even when a model rejects a harmful request, its Chain-of-Thought (CoT) reasoning may still internally contain and propagate unsafe narratives. To analyze this phenomenon, we introduce a unified safety-analysis framework that systematically deconstructs CoT generation across model layers and evaluates the role of individual attention heads through Jacobian-based spectral metrics. Within this framework, we introduce three interpretable measures: stability, geometry, and energy to quantify how specific attention heads respond or embed deceptive reasoning patterns. Extensive experiments on multiple reasoning-oriented LLMs show that the generation risk rises significantly when the thinking mode is activated, where the critical routing decisions are concentrated in only a few contiguous mid-depth layers. By precisely identifying the attention heads responsible for this divergence, our work challenges the assumption that refusal implies safety and provides a new understanding perspective for mitigating latent reasoning risks.

cs.CL

Decay Properties of Invariant Measure and Application to Elliptic Homogenization of Non-divergence Form with an Interface

Using the self-contained PDE analysis, this paper investigates the existence and the decay properties of the invariant measure in elliptic homogenization of non-divergence form with an interface assumptions on the leading coefficient $A$ and the drift $b$ for $b_1\equiv 0$, which partially provides an alternative proof of the previous work by Hairer and Manson [Ann. Probab. 39(2011) 648-682]. Moreover, as a direct application after using the analysis by the second author [Calc. Var. Partial Differ. Equ. 64(2025) No. 114], we obtain the quantitative estimates for the homogenization problem.

math.AP

LRA-GNN: Latent Relation-Aware Graph Neural Network with Initial and Dynamic Residual for Facial Age Estimation

Face information is mainly concentrated among facial key points, and frontier research has begun to use graph neural networks to segment faces into patches as nodes to model complex face representations. However, these methods construct node-to-node relations based on similarity thresholds, so there is a problem that some latent relations are missing. These latent relations are crucial for deep semantic representation of face aging. In this novel, we propose a new Latent Relation-Aware Graph Neural Network with Initial and Dynamic Residual (LRA-GNN) to achieve robust and comprehensive facial representation. Specifically, we first construct an initial graph utilizing facial key points as prior knowledge, and then a random walk strategy is employed to the initial graph for obtaining the global structure, both of which together guide the subsequent effective exploration and comprehensive representation. Then LRA-GNN leverages the multi-attention mechanism to capture the latent relations and generates a set of fully connected graphs containing rich facial information and complete structure based on the aforementioned guidance. To avoid over-smoothing issues for deep feature extraction on the fully connected graphs, the deep residual graph convolutional networks are carefully designed, which fuse adaptive initial residuals and dynamic developmental residuals to ensure the consistency and diversity of information. Finally, to improve the estimation accuracy and generalization ability, progressive reinforcement learning is proposed to optimize the ensemble classification regressor. Our proposed framework surpasses the state-of-the-art baselines on several age estimation benchmarks, demonstrating its strength and effectiveness.

cs.CV

GroupFace: Imbalanced Age Estimation Based on Multi-hop Attention Graph Convolutional Network and Group-aware Margin Optimization

With the recent advances in computer vision, age estimation has significantly improved in overall accuracy. However, owing to the most common methods do not take into account the class imbalance problem in age estimation datasets, they suffer from a large bias in recognizing long-tailed groups. To achieve high-quality imbalanced learning in long-tailed groups, the dominant solution lies in that the feature extractor learns the discriminative features of different groups and the classifier is able to provide appropriate and unbiased margins for different groups by the discriminative features. Therefore, in this novel, we propose an innovative collaborative learning framework (GroupFace) that integrates a multi-hop attention graph convolutional network and a dynamic group-aware margin strategy based on reinforcement learning. Specifically, to extract the discriminative features of different groups, we design an enhanced multi-hop attention graph convolutional network. This network is capable of capturing the interactions of neighboring nodes at different distances, fusing local and global information to model facial deep aging, and exploring diverse representations of different groups. In addition, to further address the class imbalance problem, we design a dynamic group-aware margin strategy based on reinforcement learning to provide appropriate and unbiased margins for different groups. The strategy divides the sample into four age groups and considers identifying the optimum margins for various age groups by employing a Markov decision process. Under the guidance of the agent, the feature representation bias and the classification margin deviation between different groups can be reduced simultaneously, balancing inter-class separability and intra-class proximity. After joint optimization, our architecture achieves excellent performance on several age estimation benchmark datasets.

cs.CV

Large-scale boundary estimates of parabolic homogenization over rough boundaries

In this paper, for a family of second-order parabolic system or equation with rapidly oscillating and time-dependent periodic coefficients over rough boundaries, we obtain the large-scale boundary estimates, by a quantitative approach. The quantitative approach relies on approximating twice: we first approximate the original parabolic problem over rough boundary by the same equation over a non-oscillating boundary and then approximate the oscillating equation over a non-oscillating boundary by its homogenized equation over the same non-oscillating boundary.

math.AP

Reward-Robust RLHF in LLMs

As Large Language Models (LLMs) continue to progress toward more advanced forms of intelligence, Reinforcement Learning from Human Feedback (RLHF) is increasingly seen as a key pathway toward achieving Artificial General Intelligence (AGI). However, the reliance on reward-model-based (RM-based) alignment methods introduces significant challenges due to the inherent instability and imperfections of Reward Models (RMs), which can lead to critical issues such as reward hacking and misalignment with human intentions. In this paper, we introduce a reward-robust RLHF framework aimed at addressing these fundamental challenges, paving the way for more reliable and resilient learning in LLMs. Our approach introduces a novel optimization objective that carefully balances performance and robustness by incorporating Bayesian Reward Model Ensembles (BRME) to model the uncertainty set of reward functions. This allows the framework to integrate both nominal performance and minimum reward signals, ensuring more stable learning even with imperfect RMs. Empirical results demonstrate that our framework consistently outperforms baselines across diverse benchmarks, showing improved accuracy and long-term stability. We also provide a theoretical analysis, demonstrating that reward-robust RLHF approaches the stability of constant reward settings, which proves to be acceptable even in a stochastic-case analysis. Together, these contributions highlight the framework potential to enhance both the performance and stability of LLM alignment.

cs.LG

A Multi-view Mask Contrastive Learning Graph Convolutional Neural Network for Age Estimation

The age estimation task aims to use facial features to predict the age of people and is widely used in public security, marketing, identification, and other fields. However, the features are mainly concentrated in facial keypoints, and existing CNN and Transformer-based methods have inflexibility and redundancy for modeling complex irregular structures. Therefore, this paper proposes a Multi-view Mask Contrastive Learning Graph Convolutional Neural Network (MMCL-GCN) for age estimation. Specifically, the overall structure of the MMCL-GCN network contains a feature extraction stage and an age estimation stage. In the feature extraction stage, we introduce a graph structure to construct face images as input and then design a Multi-view Mask Contrastive Learning (MMCL) mechanism to learn complex structural and semantic information about face images. The learning mechanism employs an asymmetric siamese network architecture, which utilizes an online encoder-decoder structure to reconstruct the missing information from the original graph and utilizes the target encoder to learn latent representations for contrastive learning. Furthermore, to promote the two learning mechanisms better compatible and complementary, we adopt two augmentation strategies and optimize the joint losses. In the age estimation stage, we design a Multi-layer Extreme Learning Machine (ML-IELM) with identity mapping to fully use the features extracted by the online encoder. Then, a classifier and a regressor were constructed based on ML-IELM, which were used to identify the age grouping interval and accurately estimate the final age. Extensive experiments show that MMCL-GCN can effectively reduce the error of age estimation on benchmark datasets such as Adience, MORPH-II, and LAP-2016.

cs.CV

Convergence Rates for the Stationary and Non-stationary Navier-Stokes Equations over Non-Lipschitz Boundaries

In this paper, we consider the higher-order convergence rates for the 2D stationary and non-stationary Navier-Stokes Equations over highly oscillating periodic bumpy John domains with $C^{2}$ regularity in some neighborhood of the boundary point (0,0). For the stationary case and any $γ\in (0,1/2)$, using the variational equation satisfied by the solution and the correctors for the bumpy John domains obtained by Higaki, Prange and Zhuge \cite{higaki2021large,MR4619004} after correcting the values on the inflow/outflow boundaries $(\{0\}\cup\{1\})\times(0,1)$, we can obtain an $O(\varepsilon^{2-γ})$ approximation in $L^2$ for the velocity and an $O(\varepsilon^{2-γ})$ convergence rates in $L^1$ approximated by the so called Navier's wall laws, which generalized the results obtained by Jäger and Mikelić \cite{MR1813101}. Moreover, for the non-stationary case, using the energy method, we can obtain an $O(\varepsilon^{2-γ}+\exp(-Ct))$ convergence rate for the velocity in $L_x^1$.

math.AP

On the periodic homogenization of elliptic equations in non-divergence form with large drifts

We study the quantitative homogenization of linear second order elliptic equations in non-divergence form with highly oscillating periodic diffusion coefficients and with large drifts, in the so-called ``centered'' setting where homogenization occurs and the large drifts contribute to the effective diffusivity. Using the centering condition and the invariant measures associated to the underlying diffusion process, we transform the equation into divergence form with modified diffusion coefficients but without drift. The latter is in the standard setting for which quantitative homogenization results have been developed systematically. An application of those results then yields quantitative estimates, such as the convergence rates and uniform Lipschitz regularity, for equations in non-divergence form with large drifts.

math.AP

Parabolic Homogenization with an Interface

This paper considers a family of second-order parabolic equations in divergence form with rapidly oscillating and time-dependent periodic coefficients and an interface between two periodic structures. Following a framework initiated by Blanc, Le Bris and Lions and a generalized two-scale expansion in divergence form of elliptic homogenization with an interface by Josien, we can determine the effective (or homogenized) equation with the coefficient matrix being piecewise constant and discontinuous across the interface. Moreover, we obtain the $O(\varepsilon)$ convergence rates in $L^{2(d+2)/{d}}_{x,t}$ with $\varepsilon$-smoothing method and the uniform interior Lipschitz estimates via compactness argument.

math.AP

Quantitative Estimates in Elliptic Homogenization of Non-divergence Form with Unbounded Drift and an Interface

This paper investigates quantitative estimates in elliptic homogenization of non-divergence form with unbounded drift and an interface, which continues the study of the previous work by Hairer and Manson [Ann. Probab. 39(2011) 648-682], where they investigated the limiting long time/large scale behavior of such a process under diffusive rescaling. We determine the effective equation and obtain the size estimates of the gradient of Green functions as well as the optimal (in general) convergence rates. The proof relies on transferring the non-divergence form into the divergence-form with the coefficient matrix decaying exponentially to some (different) periodic matrix on the different sides of the interface first and then investigating this special structure in homogenization of divergence form.

math.AP

Homogenization and Convergence Rates for Periodic Parabolic Equations with Highly Oscillating Potentials

This paper considers a family of second-order periodic parabolic equations with highly oscillating potentials, which have been considered many times for the time-varying potentials in stochastic homogenization. Following a standard two-scale expansions illusion, we can guess and succeed in determining the homogenized equation in different cases that the potentials satisfy the corresponding assumptions, based on suitable uniform estimates of the $L^2(0,T;H^1(Ω))$-norm for the solutions. To handle the more singular case and obtain the convergence rates in $L^\infty(0,T;L^2(Ω))$, we need to estimate the Hessian term as well as the t-derivative term more exactly, which may be depend on $\varepsilon$. The difficulty is to find suitable uniform estimates for the $L^2(0,T;H^1(Ω))$-norm and suitable estimates for the higher order derivative terms.

math.AP

A Carleman-Type Inequality in Elliptic Periodic Homogenization

In this paper, for a family of second-order elliptic equations with rapidly oscillating periodic coefficients, we are interested in a Carleman-type inequality for these solutions satisfying an additional growth condition in elliptic periodic homogenization, which implies a three-ball inequality without an error term at a macroscopic scale. Moreover, if we replace the additional growth condition by the doubling condition at a macroscopic scale, then the three-ball inequality without an error term holds at any scale. The proof relies on the convergence of $H^1$-norm for the solution and the compactness argument.

math.AP