SearcharxivSearch

arXiv subjects

Yunjiao Zhang

Publications and source records attributed to Yunjiao Zhang.

3 recordsLinked to original sources

Beyond Sharpness: A Flatness Decomposition Framework for Efficient Continual Learning

Continual Learning (CL) aims to enable models to sequentially learn multiple tasks without forgetting previous knowledge. Recent studies have shown that optimizing towards flatter loss minima can improve model generalization. However, existing sharpness-aware methods for CL suffer from two key limitations: (1) they treat sharpness regularization as a unified signal without distinguishing the contributions of its components. and (2) they introduce substantial computational overhead that impedes practical deployment. To address these challenges, we propose FLAD, a novel optimization framework that decomposes sharpness-aware perturbations into gradient-aligned and stochastic-noise components, and show that retaining only the noise component promotes generalization. We further introduce a lightweight scheduling scheme that enables FLAD to maintain significant performance gains even under constrained training time. FLAD can be seamlessly integrated into various CL paradigms and consistently outperforms standard and sharpness-aware optimizers in diverse experimental settings, demonstrating its effectiveness and practicality in CL.

cs.LG

Information-Theoretic Generalization Bounds of Replay-based Continual Learning

Continual learning (CL) has emerged as a dominant paradigm for acquiring knowledge from sequential tasks while avoiding catastrophic forgetting. Although many CL methods have been proposed to show impressive empirical performance, the theoretical understanding of their generalization behavior remains limited, particularly for replay-based approaches. This paper establishes a unified theoretical framework for replay-based CL, deriving a series of information-theoretic generalization bounds that explicitly elucidate the impact of the memory buffer alongside the current task on generalization performance. Specifically, our hypothesis-based bounds capture the trade-off between the number of selected exemplars and the information dependency between the hypothesis and the memory buffer. Our prediction-based bounds yield tighter and computationally tractable upper bounds on the generalization error by leveraging low-dimensional variables. Theoretical analysis is general and broadly applicable to a wide range of learning algorithms, exemplified by stochastic gradient Langevin dynamics (SGLD) as a representative method. Comprehensive experimental evaluations demonstrate the effectiveness of our derived bounds in capturing the generalization dynamics in replay-based CL settings.

cs.LG

Attention-aware convolutional neural networks for identification of magnetic islands in the tearing mode on EAST tokamak

The tearing mode, a large-scale MHD instability in tokamak, typically disrupts the equilibrium magnetic surfaces, leads to the formation of magnetic islands, and reduces core electron temperature and density, thus resulting in significant energy losses and may even cause discharge termination. This process is unacceptable for ITER. Therefore, the accurate identification of a magnetic island in real time is crucial for the effective control of the tearing mode in ITER in the future. In this study, based on the characteristics induced by tearing modes, an attention-aware convolutional neural network (AM-CNN) is proposed to identify the presence of magnetic islands in tearing mode discharge utilizing the data from ECE diagnostics in the EAST tokamak. A total of 11 ECE channels covering the range of core is used in the tearing mode dataset, which includes 2.5*10^9 data collected from 68 shots from 2016 to 2021 years. We split the dataset into training, validation, and test sets (66.5%, 5.7%, and 27.8%), respectively. An attention mechanism is designed to couple with the convolutional neural networks to improve the capability of feature extraction of signals. During the model training process, we utilized adaptive learning rate adjustment and early stopping mechanisms to optimize performance of AM-CNN. The model results show that a classification accuracy of 91.96% is achieved in tearing mode identification. Compared to CNN without AM, the attention-aware convolutional neural networks demonstrate great performance across accuracy, recall metrics, and F1 score. By leveraging the deep learning model, which incorporates a physical understanding of the tearing process to identify tearing mode behaviors, the combination of physical mechanisms and deep learning is emphasized, significantly laying an important foundation for the future intelligent control of tearing mode dynamics.

physics.plasm-ph