Searcharxiv⌕ Search

arXiv subjects

Zhuoqin Yang

Publications and source records attributed to Zhuoqin Yang.

8 recordsLinked to original sources

Maximizing Incremental Information Entropy for Contrastive Learning

Contrastive learning has achieved remarkable success in self-supervised representation learning, often guided by information-theoretic objectives such as mutual information maximization. Motivated by the limitations of static augmentations and rigid invariance constraints, we propose IE-CL (Incremental-Entropy Contrastive Learning), a framework that explicitly optimizes the entropy gain between augmented views while preserving semantic consistency. Our theoretical framework reframes the challenge by identifying the encoder as an information bottleneck and proposes a joint optimization of two components: a learnable transformation for entropy generation and an encoder regularizer for its preservation. Experiments on CIFAR-10/100, STL-10, and ImageNet demonstrate that IE-CL consistently improves performance under small-batch settings. Moreover, our core modules can be seamlessly integrated into existing frameworks. This work bridges theoretical principles and practice, offering a new perspective in contrastive learning.

cs.LG↗

Vision KAN: Towards an Attention-Free Backbone for Vision with Kolmogorov-Arnold Networks

Attention mechanisms have become a key module in modern vision backbones due to their ability to model long-range dependencies. However, their quadratic complexity in sequence length and the difficulty of interpreting attention weights limit both scalability and clarity. Recent attention-free architectures demonstrate that strong performance can be achieved without pairwise attention, motivating the search for alternatives. In this work, we introduce Vision KAN (ViK), an attention-free backbone inspired by the Kolmogorov-Arnold Networks. At its core lies MultiPatch-RBFKAN, a unified token mixer that combines (a) patch-wise nonlinear transform with Radial Basis Function-based KANs, (b) axis-wise separable mixing for efficient local propagation, and (c) low-rank global mapping for long-range interaction. Employing as a drop-in replacement for attention modules, this formulation tackles the prohibitive cost of full KANs on high-resolution features by adopting a patch-wise grouping strategy with lightweight operators to restore cross-patch dependencies. Experiments on ImageNet-1K show that ViK achieves competitive accuracy with linear complexity, demonstrating the potential of KAN-based token mixing as an efficient and theoretically grounded alternative to attention.

cs.CV↗

Enhancing Federated Learning with Kolmogorov-Arnold Networks: A Comparative Study Across Diverse Aggregation Strategies

Multilayer Perceptron (MLP), as a simple yet powerful model, continues to be widely used in classification and regression tasks. However, traditional MLPs often struggle to efficiently capture nonlinear relationships in load data when dealing with complex datasets. Kolmogorov-Arnold Networks (KAN), inspired by the Kolmogorov-Arnold representation theorem, have shown promising capabilities in modeling complex nonlinear relationships. In this study, we explore the performance of KANs within federated learning (FL) frameworks and compare them to traditional Multilayer Perceptrons. Our experiments, conducted across four diverse datasets demonstrate that KANs consistently outperform MLPs in terms of accuracy, stability, and convergence efficiency. KANs exhibit remarkable robustness under varying client numbers and non-IID data distributions, maintaining superior performance even as client heterogeneity increases. Notably, KANs require fewer communication rounds to converge compared to MLPs, highlighting their efficiency in FL scenarios. Additionally, we evaluate multiple parameter aggregation strategies, with trimmed mean and FedProx emerging as the most effective for optimizing KAN performance. These findings establish KANs as a robust and scalable alternative to MLPs for federated learning tasks, paving the way for their application in decentralized and privacy-preserving environments.

cs.LG↗

MedKAN: An Advanced Kolmogorov-Arnold Network for Medical Image Classification

Recent advancements in deep learning for image classification predominantly rely on convolutional neural networks (CNNs) or Transformer-based architectures. However, these models face notable challenges in medical imaging, particularly in capturing intricate texture details and contextual features. Kolmogorov-Arnold Networks (KANs) represent a novel class of architectures that enhance nonlinear transformation modeling, offering improved representation of complex features. In this work, we present MedKAN, a medical image classification framework built upon KAN and its convolutional extensions. MedKAN features two core modules: the Local Information KAN (LIK) module for fine-grained feature extraction and the Global Information KAN (GIK) module for global context integration. By combining these modules, MedKAN achieves robust feature modeling and fusion. To address diverse computational needs, we introduce three scalable variants--MedKAN-S, MedKAN-B, and MedKAN-L. Experimental results on nine public medical imaging datasets demonstrate that MedKAN achieves superior performance compared to CNN- and Transformer-based models, highlighting its effectiveness and generalizability in medical image analysis.

cs.CV↗

Activation Space Selectable Kolmogorov-Arnold Networks

The multilayer perceptron (MLP), a fundamental paradigm in current artificial intelligence, is widely applied in fields such as computer vision and natural language processing. However, the recently proposed Kolmogorov-Arnold Network (KAN), based on nonlinear additive connections, has been proven to achieve performance comparable to MLPs with significantly fewer parameters. Despite this potential, the use of a single activation function space results in reduced performance of KAN and related works across different tasks. To address this issue, we propose an activation space Selectable KAN (S-KAN). S-KAN employs an adaptive strategy to choose the possible activation mode for data at each feedforward KAN node. Our approach outperforms baseline methods in seven representative function fitting tasks and significantly surpasses MLP methods with the same level of parameters. Furthermore, we extend the structure of S-KAN and propose an activation space selectable Convolutional KAN (S-ConvKAN), which achieves leading results on four general image classification datasets. Our method mitigates the performance variability of the original KAN across different tasks and demonstrates through extensive experiments that feedforward KANs with selectable activations can achieve or even exceed the performance of MLP-based methods. This work contributes to the understanding of the data-centric design of new AI paradigms and provides a foundational reference for innovations in KAN-based network architectures.

cs.LG↗

Using single-cell entropy to describe the dynamics of reprogramming and differentiation of induced pluripotent stem cells

Induced pluripotent stem cells (iPSCs) provide a great model to study the process of reprogramming and differentiation of stem cells. Single-cell RNA sequencing (scRNA-seq) enables us to investigate the reprogramming process at single-cell level. Here, we introduce single-cell entropy (scEntropy) as a macroscopic variable to quantify the cellular transcriptome from scRNA-seq data during reprogramming and differentiation of iPSCs. scEntropy measures the relative order parameter of genomic transcriptions at single cell level during the cell fate change process, which shows increasing during differentiation, and decreasing upon reprogramming. Moreover, based on the scEntropy dynamics, we construct a phenomenological stochastic differential equation model and the corresponding Fokker-Plank equation for cell state transitions during iPSC differentiation, which provide insights to infer cell fates changes and stem cell differentiation. This study is the first to introduce the novel concept of scEntropy to the biological process of iPSC, and suggests that the scEntropy can provide a suitable quantify to describe cell fate transition in differentiation and reprogramming of stem cells.

q-bio.CB↗

DNA methylation heterogeneity induced by collaborations between enhancers

During mammalian embryo development, reprogramming of DNA methylation plays important roles in the erasure of parental epigenetic memory and the establishment of naïve pluripogent cells. Multiple enzymes that regulate the processes of methylation and demethylation work together to shape the pattern of genome-scale DNA methylation and guid the process of cell differentiation. Recent availability of methylome information from single-cell whole genome bisulfite sequencing (scBS-seq) provides an opportunity to study DNA methylation dynamics in the whole genome in individual cells, which reveal the heterogeneous methylation distributions of enhancers in embryo stem cells (ESCs). In this study, we developed a computational model of enhancer methylation inheritance to study the dynamics of genome-scale DNA methylation reprogramming during exit from pluripotency. The model enables us to track genome-scale DNA methylation reprogramming at single-cell level during the embryo development process, and reproduce the DNA methylation heterogeneity reported by scBS-seq. Model simulations show that DNA methylation heterogeneity is an intrinsic property driven by cell division along the development process, and the collaboration between neighboring enhancers is required for heterogeneous methylation. Our study suggest that the mechanism of genome-scale oscillation proposed by Rulands et al. (2018) might not necessary to the DNA methylation during exit from pluripotency.

q-bio.GN↗

Bifurcation analysis and potential landscape of the p53-Mdm2 oscillator regulated by the co-activator PDCD5

Dynamics of p53 is known to play important roles in the regulation of cell fate decisions in response to various stresses, and PDCD5 functions as a co-activator of p53 to modulate the p53 dynamics. In the present paper, we investigate how p53 dynamics are modulated by PDCD5 during the DNA damage response using methods of bifurcation analysis and potential landscape. Our results reveal that p53 activities can display rich dynamics under different PDCD5 levels, including monostability, bistability with two stable steady states, oscillations, and co-existence of a stable steady state and an oscillatory state. Physical properties of the p53 oscillations are further shown by the potential landscape, in which the potential force attracts the system state to the limit cycle attractor, and the curl flux force drives the coherent oscillation along the cyclic. We also investigate the effect of PDCD5 efficiency on inducing the p53 oscillations. We show that Hopf bifurcation is induced by increasing the PDCD5 efficiency, and the system dynamics show clear transition features in both barrier height and energy dissipation when the efficiency is close to the bifurcation point. This study provides a global picture of how PDCD5 regulates p53 dynamics via the interaction with the p53-Mdm2 oscillator and can be helpful in understanding the complicate p53 dynamics in a more complete p53 pathway.

q-bio.MN↗