SearcharxivSearch

arXiv subjects

Chuan Sun

Publications and source records attributed to Chuan Sun.

10 recordsLinked to original sources

"Anomalous Solid Solution" in Ultra-High Melting Point Oxides: A New Strategy for Developing Ultra-High Temperature Thermal Protection Coatings

The high-temperature performance of ultra-high temperature ceramics (UHTCs) in atmospheric environment is fundamentally governed by their melting points of oxidation products. Typical high-melting-point oxides, such as ZrO2, undergo phase transformations at elevated temperatures, leading to structural instability. Although doping with rare-earth or transition-metal cations can suppress these transformations, it often results in a reduction in melting point, thereby limiting practical service temperature. Here, ytterbia-stabilized zirconia (YbSZ) coatings are prepared via atmospheric plasma spraying, achieving a remarkable increase in the melting point of ZrO2 to approximately 2850 $^\circ\mathrm{C}$ and raising the ultimate plasma and oxyacetylene ablation temperature up to nearly 2780 $^\circ\mathrm{C}$ and 3200 $^\circ\mathrm{C}$, which is the highest temperature resistance property as reported. Notably, this performance enhancement originates from a synergistic mechanism of strengthened ionic-covalent mixed bonding and improved oxygen vacancy stability. Based on these findings, the concept of "anomalous solid solution" is firstly proposed to be used in the area of ultra-high temperature protection, which provides new insights into the compositional design of UHTC systems.

cond-mat.mtrl-sci

Efficient Shapley Value-based Non-Uniform Pruning of Large Language Models

Pruning large language models (LLMs) is a promising solution for reducing model sizes and computational complexity while preserving performance. Traditional layer-wise pruning methods often adopt a uniform sparsity approach across all layers, which leads to suboptimal performance due to the varying significance of individual transformer layers within the model not being accounted for. To this end, we propose the Shapley Value-based Non-Uniform Pruning (SV-NUP) method for LLMs. This approach quantifies the contribution of each transformer layer to the overall model performance, enabling the assignment of tailored pruning budgets to different layers to retain critical parameters. To further improve efficiency, we design the Sliding Window-based Shapley Value approximation method. It substantially reduces computational overhead compared to exact SV calculation methods. Extensive experiments on various LLMs including LLaMA-v1, LLaMA-v2 and OPT demonstrate the effectiveness of the proposed approach. The results reveal that non-uniform pruning significantly enhances the performance of pruned models. Notably, SV-NUP achieves a reduction in perplexity (PPL) of 18.01% and 19.55% on LLaMA-7B and LLaMA-13B, respectively, compared to SparseGPT at 70% sparsity.

cs.CL

Ten Challenging Problems in Federated Foundation Models

Federated Foundation Models (FedFMs) represent a distributed learning paradigm that fuses general competences of foundation models as well as privacy-preserving capabilities of federated learning. This combination allows the large foundation models and the small local domain models at the remote clients to learn from each other in a teacher-student learning setting. This paper provides a comprehensive summary of the ten challenging problems inherent in FedFMs, encompassing foundational theory, utilization of private data, continual learning, unlearning, Non-IID and graph data, bidirectional knowledge transfer, incentive mechanism design, game mechanism design, model watermarking, and efficiency. The ten challenging problems manifest in five pivotal aspects: ``Foundational Theory," which aims to establish a coherent and unifying theoretical framework for FedFMs. ``Data," addressing the difficulties in leveraging domain-specific knowledge from private data while maintaining privacy; ``Heterogeneity," examining variations in data, model, and computational resources across clients; ``Security and Privacy," focusing on defenses against malicious attacks and model theft; and ``Efficiency," highlighting the need for improvements in training, communication, and parameter efficiency. For each problem, we offer a clear mathematical definition on the objective function, analyze existing methods, and discuss the key challenges and potential solutions. This in-depth exploration aims to advance the theoretical foundations of FedFMs, guide practical implementations, and inspire future research to overcome these obstacles, thereby enabling the robust, efficient, and privacy-preserving FedFMs in various real-world applications.

cs.LG

Highly stable modular-assembled laser system for a dual-atom-interferometer gyroscope

Operating atom-interferometer gyroscopes outside a laboratory environment is challenging primarily owing to the instability of laser systems. To enhance the thermal stability of free-space laser systems, a compact laser system using fiber lasers and all-quartz-jointed optical modules was developed for a dual-atom-interferometer gyroscope. Millimeter-scale optical elements jointed on quartz plates with identical quartz supports, ensure laser power stability and facilitate component upgrades. The primary diode laser was locked to the modulation transfer spectrum of Rb atoms, and Raman lasers were phase-locked to the primary laser. Frequencies for repumping, blow-away, and detection lasers were adjusted with acousto-optic modulators. At room temperature, laser power fluctuation was under 1:1000, polarization extinction ratio exceeded 30 dB, frequency fluctuation was below 91 kHz, and phase noise reached to -100 dBc/Hz @ 1 kHz. The optical modules were tested at 5--50 $^{\circ}$C and applied to a dual-atom-interferometer gyroscope. The fringe contrast was tested over the temperature range. The proposed system paves the way for promoting field applications of atom-interferometer sensors.

physics.atom-ph

SPD-CFL: Stepwise Parameter Dropout for Efficient Continual Federated Learning

Federated Learning (FL) is a collaborative machine learning paradigm for training models on local sensitive data with privacy protection. Pre-trained transformer-based models have emerged as useful foundation models (FMs) to be fine-tuned for a wide range of downstream tasks. However, large-scale pre-trained models make it challenging for traditional FL due to high communication overhead in the resource-constrained IoT. This has inspired the field of parameter-efficient fine-tuning (PEFT) research. Existing PEFT methods attempt to optimize model performance at the given dropout level. Such an approach places the burden on human users to find a dropout rate that provides a satisfactory level of performance through trial-and-error, which is time consuming and resource intensive. To address this limitation, we propose the Step-wise Parameter Dropout for Continual Federated Learning (SPD-CFL) approach. Instead of pre-defining a desired dropout rate, it allows users to specify the target level of performance and then attempts to find the most suitable dropout rate for the given FL model. Specifically, on the server side, SPD-CFL drops trainable parameters in a stepwise manner to improve communication efficiency by reducing the rank of low-rank adaptation (LoRA). The sensitivity-based gradient consistency (SGC) measure is designed to facilitate the adaptive adjustment of parameter dropout. In addition, SPD-CFL introduces continual learning (CL) on the client side to mitigate performance degradation due to the inconsistent optima with distinct parameter dropout rates under heterogeneous FL. Extensive experiments on the public benchmark dataset CIFAR-10 and a real-world medical Face dataset demonstrate significant superiority of SPD-CFL over state-of-the-art methods. Compared to the best-performing baseline, it achieves a 2.07% higher test AUC while reducing communication overhead by 29.53%.

cs.LG

Self-Calibrated Atom-Interferometer Gyroscope by Modulating Atomic Velocities

Atom-interferometer gyroscopes have attracted much attention for their potential superior long-term stability and extremely low drift. For such high precision instrument, a self-calibration to achieve an absolute rotation measurement is highly demanded. Here we propose and demonstrate a self-calibration of the atomic gyroscope. The calibration is realized by using the detuning of laser frequency to control the atomic velocity thus to modulate the scale factor of the gyroscope. The modulation determines the order and the initial phase of the interference stripe, thus eliminates the ambiguity caused by the periodicity of the interferometric signal. The calibration method is verified by measuring the Earth's rotation. Long-term stable and self-calibrated atom-interferometer gyros can find important applications in the fields of fundamental physics and long-time navigation.

physics.atom-ph

Compact multi-channel radio frequency pulse sequence generator with fast switching capability for cold atom interferometers

Cold atom interferometers have matured to a powerful tool in fundamental physics research, and they are currently on their way from realizations in the laboratory to applications in the real world. The radio frequency (RF) generator is an indispensable component to control lasers and then manipulate atoms. We developed a highly compact RF generator for fast switching and sweeping frequencies/amplitudes of atomic interference pulse sequences. Multi-channel RF signals are generated by using a field-programmable gate array (FPGA) to control eight direct digital synthesizers (DDSs). We further proposed and demonstrated a method of preloading the parameters of all RF pulse sequences to the DDS registers before the execution of the pulse sequences, which eliminates the data transfer between the FPGA and DDSs to change RF signals and thus sharply shortens the delay of frequency switching when the pulse sequences are running. The characterized performance shows the generated RF signals achieve 119 ns frequency switching delay, and 40 dB harmonic rejection ratio. The generated RF pulse sequences are applied to a cold atom-interferometer gyroscope, and the contrast of atomic interference fringes reaches 38%. This compact multi-channel generator with the fast frequency/amplitude switching or/and sweeping capability has beneficial applications for the real-world atom interferometers.

physics.atom-ph

Self-Alignment of a Large-Area Dual-Atom-Interferometer Gyroscope Using Parameter Decoupled Phase Seeking Calibrations

We realize a Mach-Zehnder-type dual-atom-interferometer gyroscope with an interrogation arm of 40 cm length and the interference area up to 1.2 cm$^2$. The precise angular alignment of the large-scale separated Raman lasers is demonstrated by seeking the phase intersection of Ramsey-Bord$\acute{e}$ interferometers after the gravity effect is compensated and by decoupling the velocity dependent crosstalk phase shifts, and applied to build the Mach-Zehnder atom interferometer. Then a compact inertial rotation sensor is realized based on dual large-area Mach-Zehnder atom interferometers by precisely aligning the large-scale separated Raman lasers, in which the coherence is well preserved and the common noise is differentially suppressed. The sensor presents a sensitivity of $1.5\times10^{-7}$ rad/s/Hz$^{1/2}$, and a stability of $9.5\times10^{-10}$ rad/s at 23000 s. The absolute rotation measurement is carried out by adjusting the atomic velocity which corresponds to modulating the scale factor.

physics.atom-ph

Co-Correcting: Noise-tolerant Medical Image Classification via mutual Label Correction

With the development of deep learning, medical image classification has been significantly improved. However, deep learning requires massive data with labels. While labeling the samples by human experts is expensive and time-consuming, collecting labels from crowd-sourcing suffers from the noises which may degenerate the accuracy of classifiers. Therefore, approaches that can effectively handle label noises are highly desired. Unfortunately, recent progress on handling label noise in deep learning has gone largely unnoticed by the medical image. To fill the gap, this paper proposes a noise-tolerant medical image classification framework named Co-Correcting, which significantly improves classification accuracy and obtains more accurate labels through dual-network mutual learning, label probability estimation, and curriculum label correcting. On two representative medical image datasets and the MNIST dataset, we test six latest Learning-with-Noisy-Labels methods and conduct comparative studies. The experiments show that Co-Correcting achieves the best accuracy and generalization under different noise ratios in various tasks. Our project can be found at: https://github.com/JiarunLiu/Co-Correcting.

eess.IV

DeceFL: A Principled Decentralized Federated Learning Framework

Traditional machine learning relies on a centralized data pipeline, i.e., data are provided to a central server for model training. In many applications, however, data are inherently fragmented. Such a decentralized nature of these databases presents the biggest challenge for collaboration: sending all decentralized datasets to a central server raises serious privacy concerns. Although there has been a joint effort in tackling such a critical issue by proposing privacy-preserving machine learning frameworks, such as federated learning, most state-of-the-art frameworks are built still in a centralized way, in which a central client is needed for collecting and distributing model information (instead of data itself) from every other client, leading to high communication pressure and high vulnerability when there exists a failure at or attack on the central client. Here we propose a principled decentralized federated learning algorithm (DeceFL), which does not require a central client and relies only on local information transmission between clients and their neighbors, representing a fully decentralized learning framework. It has been further proven that every client reaches the global minimum with zero performance gap and achieves the same convergence rate $O(1/T)$ (where $T$ is the number of iterations in gradient descent) as centralized federated learning when the loss function is smooth and strongly convex. Finally, the proposed algorithm has been applied to a number of applications to illustrate its effectiveness for both convex and nonconvex loss functions, demonstrating its applicability to a wide range of real-world medical and industrial applications.

cs.LG