Searcharxiv⌕ Search

arXiv subjects

Yan Sun

Publications and source records attributed to Yan Sun.

At least 181 records · Page 10Linked to original sources

Understanding How Consistency Works in Federated Learning via Stage-wise Relaxed Initialization

Federated learning (FL) is a distributed paradigm that coordinates massive local clients to collaboratively train a global model via stage-wise local training processes on the heterogeneous dataset. Previous works have implicitly studied that FL suffers from the ``client-drift'' problem, which is caused by the inconsistent optimum across local clients. However, till now it still lacks solid theoretical analysis to explain the impact of this local inconsistency. To alleviate the negative impact of the ``client drift'' and explore its substance in FL, in this paper, we first design an efficient FL algorithm \textit{FedInit}, which allows employing the personalized relaxed initialization state at the beginning of each local training stage. Specifically, \textit{FedInit} initializes the local state by moving away from the current global state towards the reverse direction of the latest local state. This relaxed initialization helps to revise the local divergence and enhance the local consistency level. Moreover, to further understand how inconsistency disrupts performance in FL, we introduce the excess risk analysis and study the divergence term to investigate the test error of the proposed \textit{FedInit} method. Our studies show that optimization error is not sensitive to this local inconsistency, while it mainly affects the generalization error bound in \textit{FedInit}. Extensive experiments are conducted to validate this conclusion. Our proposed \textit{FedInit} could achieve state-of-the-art~(SOTA) results compared to several advanced benchmarks without any additional costs. Meanwhile, stage-wise relaxed initialization could also be incorporated into the current advanced algorithms to achieve higher performance in the FL paradigm.

cs.LG↗

Dense gas and star formation in the Outer Milky Way

We present maps and spectra of the HCN(1-0) and HCO$^+$(1-0) lines in the extreme outer Galaxy, at galactocentric radii between 14 and 22 kpc, with the 13.7 meter Delingha telescope. The 9 molecular clouds were selected from a CO/$^{13}$CO survey of the outer quadrants. The goal is to better understand the structure of molecular clouds in these poorly studied subsolar metallicity regions and the relation with star formation. The lines are all narrow, less than 2km/s at half power, enabling detection of the HCN hyperfine structure in the stronger sources and allowing us to observationally test hyperfine collision rates. The hyperfine line ratios show that the HCN emission is optically thin with column densities estimated at N(HCN)~$3x10^{12}$\scm. The HCO$^+$ emission is approximately twice as strong as the HCN (taken as the sum of all components), in contrast with the inner Galaxy and nearby galaxies where they are similarly strong. For an abundance ratio $χ_{HCN}/χ_{HCO^+} = 3$, this requires a relatively low density solution for the dense gas, with n(H2) $\sim 10^3 - 10^4$\ccm. The $^{12}$CO/$^{13}$CO line ratios are similar to solar neighborhood values, roughly 7.5, despite the low $^{13}$CO abundance expected at such large radii. The HCO$^+$/CO and HCO$^+$/$^{13}$CO integrated intensity ratios are also standard at about 1/35 and 1/5 respectively. HCN is weak compared to the CO emission, with HCN/CO $\sim 1/70$ even after summing all hyperfine components. At the parsec scales observed here, the correlation between star formation, as traced by 24~$μ$m emission as is standard in extragalactic work, and dense gas via the HCN or HCO$^+$ emission, is poor, perhaps due to the lack of dynamic range. We find that the lowest dense gas fractions are in the sources at high galactic latitude (b>2, h>300pc above the plane), possibly due to lower pressure.

astro-ph.GA↗

Large shift current, $π$ Zak phase and unconventional nature in Se and Te

Recently, unconventional materials (or obstructed atomic insulators) have attracted much attention owing to the unconventional feature of mismatch between Wannier centers and atomic positions. In this paper, we demonstrate that the trigonal selenium and tellurium host an unconventional nature in both electronic and phonon spectra. In electronic band structures, the band representation (BR) decomposition for occupied bands has to contain the essential BR of $A@3b$, and the real-space invariant is $δ_1@3b=-1$. The $π$ Zak phase suggests that the one-dimensional Se/Te chain is a chiral Su-Schrieffer-Heeger chain. The effective magnetism can be induced by $p$ states at ends. More importantly, a large shift current is obtained in Se quantum well. In addtion, in phonon spectra, three sets of phonon bands are well separated and assigned to $B@3b$, $B@3a$, and $A@3b$ BRs, respectively. Thus, the obstructed phonon states are predicted on the (0001) surface. As the prototypes of unconventional materials in both electronic and phonon spectra, our findings could create much interest in the study of obstructed surface electronic and phonon states in these novel materials.

cond-mat.mtrl-sci↗

Molecular Clouds in the Galactic Plane from $l$ = [59.75$^\circ$, 74.75$^\circ$] and $b$ = [$-$5.25$^\circ$, +5.25$^\circ$]

In this paper we present the distribution of molecular gas in the Milky Way Galactic plane from $l$ = [59.75, 74.75]$^{\circ}$ and $b$ = [${-}$5.25, +5.25]$^{\circ}$, using the MWISP $^{12}$CO/$^{13}$CO/$\rm {C}^{18}{O}$ emission line data. The molecular gas in this region can be mainly attributed to the Local spur, Local arm, Perseus arm, and Outer arm. Statistics of the physical properties of the molecular gas in each arm, such as excitation temperature, optical depth, and column density, are presented. Using the DBSCAN algorithm, we identified 15 extremely distant molecular clouds with kinematic distances of 14.72$-$17.77 kpc and masses of 363$-$520 M$_{\odot}$, which we find could be part of the Outer Scutum-Centaurus (OSC) arm identified by \cite{2011ApJ...734L..24D} and \cite{2015ApJ...798L..27S}. It is also possible that, 12 of these 15 extremely distant molecular clouds constitute an independent structure between the Outer and the OSC arms or a spur. There exist two Gaussian components in the vertical distribution of the molecular gas in the Perseus spiral arm. These two Gaussian components correspond to two giant filaments parallel to the Galactic plane. We find an upward warping of the molecular gas in the Outer spiral arm with a displacement of around 270 pc with respect to the Galactic mid-plane.

astro-ph.GA↗

Towards More Suitable Personalization in Federated Learning via Decentralized Partial Model Training

Personalized federated learning (PFL) aims to produce the greatest personalized model for each client to face an insurmountable problem--data heterogeneity in real FL systems. However, almost all existing works have to face large communication burdens and the risk of disruption if the central server fails. Only limited efforts have been used in a decentralized way but still suffers from inferior representation ability due to sharing the full model with its neighbors. Therefore, in this paper, we propose a personalized FL framework with a decentralized partial model training called DFedAlt. It personalizes the "right" components in the modern deep models by alternately updating the shared and personal parameters to train partially personalized models in a peer-to-peer manner. To further promote the shared parameters aggregation process, we propose DFedSalt integrating the local Sharpness Aware Minimization (SAM) optimizer to update the shared parameters. It adds proper perturbation in the direction of the gradient to overcome the shared model inconsistency across clients. Theoretically, we provide convergence analysis of both algorithms in the general non-convex setting for decentralized partial model training in PFL. Our experiments on several real-world data with various data partition settings demonstrate that (i) decentralized training is more suitable for partial personalization, which results in state-of-the-art (SOTA) accuracy compared with the SOTA PFL baselines; (ii) the shared parameters with proper perturbation make partial personalized FL more suitable for decentralized training, where DFedSalt achieves most competitive performance.

cs.LG↗

Prokaryotic genome editing based on the subtype I-B-Svi CRISPR-Cas system

Type I CRISPR-Cas systems are the most common among six types of CRISPR-Cas systems, however, non-self-targeting genome editing based on a single Cas3 of type I CRISPR-Cas systems has not been reported. Here, we present the subtype I-B-Svi CRISPR-Cas system (with three confirmed CRISPRs and a cas gene cluster) and genome editing based on this system found in Streptomyces virginiae IBL14. Importantly, like the animal-derived bacterial protein SpCas9 (1368 amino-acids), the single, compact, non-animal-derived bacterial protein SviCas3 (771 amino-acids) can also direct template-based microbial genome editing through the target cell's own homology-directed repair system, which breaks the view that the genome editing based on type I CRISPR-Cas systems requires a full Cascade. Notably, no off-target changes or indel-formation were detected in the analysis of potential off-target sites. This discovery broadens our understanding of the diversity of type I CRISPR-Cas systems and will facilitate new developments in genome editing tools.

q-bio.GN↗

VLBI Astrometry of Radio Stars to Link Radio and Optical Celestial Reference Frames. I. HD 199178 $\&$ AR Lacertae

To accurately link the radio and optical Celestial Reference Frames (CRFs) at optical bright end, i.e., with Gaia G band magnitude < 13, increasing number and improving sky distribution of radio stars with accurate astrometric parameters from both Very Long Baseline Interferometry (VLBI) and Gaia measurements are mandatory. We selected two radio stars HD 199178 and AR Lacertae as the target for a pilot program for the frame link, using the Very Long Baseline Array (VLBA) at 15 GHz at six epochs spanning about 1 year, to measure their astrometric parameters. The measured parallax of HD 199178 is $8.949 \pm 0.059$ mas and the proper motion is $μ_αcos δ= 26.393 \pm 0.093$, $μ_δ= -0.950 \pm 0.083~mas~yr^{-1}$, while the parallax of AR Lac is $23.459 \pm 0.094$ mas and the proper motion is $μ_αcos δ= -51.906 \pm 0.138$, $μ_δ= 46.732 \pm 0.131~mas~yr^{-1}$. Our VLBI measured astrometric parameters have accuracies about 4-5 times better than the corresponding historic VLBI measurements and comparable accuracies with those from Gaia, validating the feasibility of frame link using radio stars. With the updated astrometric parameters for these two stars, there is a 25% reduction of the uncertainties on the Y axis for both orientation and spin parameters.

astro-ph.SR↗

On Efficient Training of Large-Scale Deep Learning Models: A Literature Review

The field of deep learning has witnessed significant progress, particularly in computer vision (CV), natural language processing (NLP), and speech. The use of large-scale models trained on vast amounts of data holds immense promise for practical applications, enhancing industrial productivity and facilitating social development. With the increasing demands on computational capacity, though numerous studies have explored the efficient training, a comprehensive summarization on acceleration techniques of training deep learning models is still much anticipated. In this survey, we present a detailed review for training acceleration. We consider the fundamental update formulation and split its basic components into five main perspectives: (1) data-centric: including dataset regularization, data sampling, and data-centric curriculum learning techniques, which can significantly reduce the computational complexity of the data samples; (2) model-centric, including acceleration of basic modules, compression training, model initialization and model-centric curriculum learning techniques, which focus on accelerating the training via reducing the calculations on parameters; (3) optimization-centric, including the selection of learning rate, the employment of large batchsize, the designs of efficient objectives, and model average techniques, which pay attention to the training policy and improving the generality for the large-scale models; (4) budgeted training, including some distinctive acceleration methods on source-constrained situations; (5) system-centric, including some efficient open-source distributed libraries/systems which provide adequate hardware support for the implementation of acceleration algorithms. By presenting this comprehensive taxonomy, our survey presents a comprehensive review to understand the general mechanisms within each component and their joint interaction.

cs.LG↗

Comprehensive ab initio study of effects of alloying elements on generalized stacking fault energies of Ni and Ni$_3$Al

Excellent high-temperature mechanical properties of Ni-based single crystal superalloys (NSCSs) are attributed to the yield strength anomaly of Ni$_{3}$Al that is intimately related to generalized stacking fault energies (GSFEs). Therefore, clarifying the effects of alloying elements on the GSFEs is of great significance for alloys design. Here, by means of ab initio density functional theory calculations, we systematically calculated the GSFEs of different slip systems of Ni and Ni$_{3}$Al without and with alloying elements using the alias shear method. We obtained that for Ni, except for magnetic elements Mn, Fe, and Co, most of alloying elements decrease the unstable stacking fault energy ($γ_{usf}$) of the $[01\bar{1}](111)$ and $[11\bar{2}](111)$ slip systems and also decrease the stable stacking fault energy ($γ_{sf}$) of the $[11\bar{2}](111)$ slip system. For Ni$_{3}$Al, most of alloying elements in groups IIIB-VIIB show a strong Al site preference. Except for Mn and Fe, the elements in groups VB-VIIB and the first column of group VIII increase the values of $γ_{usf}$ of different slip systems of Ni$_{3}$Al. On the other hand, the elements in groups IIIB-VIIB also increase the value of $γ_{sf}$. We found that Re is an excellent strengthening alloying element that significantly increases the slip barrier of the tailing slip process for Ni, and also enhances the slip barrier of the leading slip process of three slip systems for Ni$_{3}$Al. W and Mo exhibit similar effects as Re. We predicted that Os, Ru, and Ir are good strengthening alloying elements as well, since they show the strengthening effects on both the leading and tailing slip process for Ni and Ni$_{3}$Al.

cond-mat.mtrl-sci↗

The parallax and 3D kinematics of water masers in the massive star-forming region G034.43+0.24

We report a trigonometric parallax measurement of 22 GHz water masers in the massive star-forming region G034.43+0.24 as part of the Bar and Spiral Structure Legacy (BeSSeL) Survey using the Very Long Baseline Array. The parallax is 0.330$\pm$50.018 mas, corresponding to a distance of $3.03^{+0.17}_{-0.16}$ kpc. This locates G034.43+0.24 near the inner edge of the Sagittarius spiral arm and at one end of a linear distribution of massive young stars which cross nearly the full width of the arm. The measured 3-dimensional motion of G034.43+0.24 indicates a near-circular Galactic orbit. The water masers display arc-like distributions, possibly bow shocks, associated with winds from one or more massive young stars.

astro-ph.GA↗

Visual Prompt Based Personalized Federated Learning

As a popular paradigm of distributed learning, personalized federated learning (PFL) allows personalized models to improve generalization ability and robustness by utilizing knowledge from all distributed clients. Most existing PFL algorithms tackle personalization in a model-centric way, such as personalized layer partition, model regularization, and model interpolation, which all fail to take into account the data characteristics of distributed clients. In this paper, we propose a novel PFL framework for image classification tasks, dubbed pFedPT, that leverages personalized visual prompts to implicitly represent local data distribution information of clients and provides that information to the aggregation model to help with classification tasks. Specifically, in each round of pFedPT training, each client generates a local personalized prompt related to local data distribution. Then, the local model is trained on the input composed of raw data and a visual prompt to learn the distribution information contained in the prompt. During model testing, the aggregated model obtains prior knowledge of the data distributions based on the prompts, which can be seen as an adaptive fine-tuning of the aggregation model to improve model performances on different clients. Furthermore, the visual prompt can be added as an orthogonal method to implement personalization on the client for existing FL methods to boost their performance. Experiments on the CIFAR10 and CIFAR100 datasets show that pFedPT outperforms several state-of-the-art (SOTA) PFL algorithms by a large margin in various settings.

cs.LG↗

Reply to: Low-frequency quantum oscillations in LaRhIn$_5$: Dirac point or nodal line?

We thank G.P. Mikitik and Yu.V. Sharlai for contributing this note and the cordial exchange about it. First and foremost, we note that the aim of our paper is to report a methodology to diagnose topological (semi)metals using magnetic quantum oscillations. Thus far, such diagnosis has been based on the phase offset of quantum oscillations, which is extracted from a "Landau fan plot". A thorough analysis of the Onsager-Lifshitz-Roth quantization rules has shown that the famous $π$-phase shift can equally well arise from orbital- or spin magnetic moments in topologically trivial systems with strong spin-orbit coupling or small effective masses. Therefore, the "Landau fan plot" does not by itself constitute a proof of a topologically nontrivial Fermi surface. In the paper at hand, we report an improved analysis method that exploits the strong energy-dependence of the effective mass in linearly dispersing bands. This leads to a characteristic temperature dependence of the oscillation frequency which is a strong indicator of nontrivial topology, even for multi-band metals with complex Fermi surfaces. Three materials, Cd$_3$As$_2$, Bi$_2$O$_2$Se and LaRhIn$_5$ served as test cases for this method. Linear band dispersions were detected for Cd$_3$As$_2$, as well as the $F$ $\approx$ 7 T pocket in LaRhIn$_5$.

cond-mat.str-el↗

On parent structures of near-ambient nitrogen-doped lutetium hydride superconductor

Recently, near-ambient superconductivity has been experimentally evidenced in a nitrogen-doped lutetium hydride by Dasenbrock-Gammon \emph{et al.} [Nature 615, 244 (2023)], which yields a remarkable maximum $T_c$ of 294 K at just 10 kbar. However, due to the difficulty of x-ray diffraction (XRD) in identifying light elements such as hydrogen and nitrogen, the crystal structure of the superconductor remains elusive, in particular for the actual stoichiometry of hydrogen and nitrogen and their atomistic positions. This holds even for its parent structure. Here, we set out to address this issue by performing a thorough density functional theory study on the structural, electronic, dynamical, and optical properties of lutetium hydrides. Through thermal and lattice dynamic analysis as well as XRD and superconductor color comparisons, we unambiguously clarified that the parent structures are a mixture of dominant LuH$_2$ phase of the CaF$_2$-type (instead of originally proposed LuH$_3$ structure of $Fm\bar{3}m$ space group) and minor LuH phase of the NaCl-type.

cond-mat.mtrl-sci↗

Spin-dependent recombination mechanisms for quintet bi-excitons generated through singlet fission

We investigate the physical mechanisms for spin-dependent recombination of a strongly bound pair of triplet excitons generated by singlet fission and forming a spin quintet (total spin of two) bi-exciton. For triplet excitons the spin-dependent recombination pathways can involve intersystem crossing or triplet-triplet annihilation back to the singlet ground state. However the modeling of spin-dependent recombination for quintets is still an open question. Here we introduce two theoretical models and compare their predictions with the broadband optically detected magnetic resonance spectrum of a long lived quintet bi-exciton with known molecular structure. This spectrum measures the change in the fluorescence signal induced by microwave excitation of each of the ten possible spin transitions within the quintet manifold as function of the magnetic field. While most of the experimental features can be reproduced for both models, the behavior of some of the transitions is only consistent with the quintet spin-recombination model inspired by triplet intersystem crossing which can reproduce accurately the experimental two-dimensional spectrum with a small number of kinetic parameters. Thus quantitative analysis of the broadband optically detected magnetic resonance signal enables quantitative understanding of the dominant spin-recombination processes and estimation of the out-of equilibrium spin populations.

cond-mat.mes-hall↗

Subspace based Federated Unlearning

Federated learning (FL) enables multiple clients to train a machine learning model collaboratively without exchanging their local data. Federated unlearning is an inverse FL process that aims to remove a specified target client's contribution in FL to satisfy the user's right to be forgotten. Most existing federated unlearning algorithms require the server to store the history of the parameter updates, which is not applicable in scenarios where the server storage resource is constrained. In this paper, we propose a simple-yet-effective subspace based federated unlearning method, dubbed SFU, that lets the global model perform gradient ascent in the orthogonal space of input gradient spaces formed by other clients to eliminate the target client's contribution without requiring additional storage. Specifically, the server first collects the gradients generated from the target client after performing gradient ascent, and the input representation matrix is computed locally by the remaining clients. We also design a differential privacy method to protect the privacy of the representation matrix. Then the server merges those representation matrices to get the input gradient subspace and updates the global model in the orthogonal subspace of the input gradient subspace to complete the forgetting task with minimal model performance degradation. Experiments on MNIST, CIFAR10, and CIFAR100 show that SFU outperforms several state-of-the-art (SOTA) federated unlearning algorithms by a large margin in various settings.

cs.LG↗

Fusion of Global and Local Knowledge for Personalized Federated Learning

Personalized federated learning, as a variant of federated learning, trains customized models for clients using their heterogeneously distributed data. However, it is still inconclusive about how to design personalized models with better representation of shared global knowledge and personalized pattern. To bridge the gap, we in this paper explore personalized models with low-rank and sparse decomposition. Specifically, we employ proper regularization to extract a low-rank global knowledge representation (GKR), so as to distill global knowledge into a compact representation. Subsequently, we employ a sparse component over the obtained GKR to fuse the personalized pattern into the global knowledge. As a solution, we propose a two-stage proximal-based algorithm named \textbf{Fed}erated learning with mixed \textbf{S}parse and \textbf{L}ow-\textbf{R}ank representation (FedSLR) to efficiently search for the mixed models. Theoretically, under proper assumptions, we show that the GKR trained by FedSLR can at least sub-linearly converge to a stationary point of the regularized problem, and that the sparse component being fused can converge to its stationary point under proper settings. Extensive experiments also demonstrate the superior empirical performance of FedSLR. Moreover, FedSLR reduces the number of parameters, and lowers the down-link communication complexity, which are all desirable for federated learning algorithms. Source code is available in \url{https://github.com/huangtiansheng/fedslr}.

cs.LG↗

The NTSC VLBI System and its application in UT1 measurement

In order to measure the Universal Time (UT1) in real time, National Time Service Center (NTSC) has built a VGOS-like (VLBI Global Observing System) broadband VLBI network, which includes three 13-m radio telescopes located in Jilin, Sanya and Kashi, and a data analysis center in Xi'an. Each station is equipped with a highly stable hydrogen atomic clock and a self-developed VLBI backend, and is co-located with two GPS receivers. This VGOS-like VLBI network may play an important role in improving the Chinese broadband VLBI technology and making valuable contributions to domestic VLBI measurements of UT1. In this paper, we introduce the specifications of this VLBI network, and present the UT1 measurements at C-band conducted in 2018 using the Jilin-Kashi baseline of this network. The comparisons between our UT1 estimates and those provided by IERS suggest that the NTSC VLBI network is capable to determine UT1 accurate at the level of 58.8 microseconds.

astro-ph.IM↗

Enhance Local Consistency in Federated Learning: A Multi-Step Inertial Momentum Approach

Federated learning (FL), as a collaborative distributed training paradigm with several edge computing devices under the coordination of a centralized server, is plagued by inconsistent local stationary points due to the heterogeneity of the local partial participation clients, which precipitates the local client-drifts problems and sparks off the unstable and slow convergence, especially on the aggravated heterogeneous dataset. To address these issues, we propose a novel federated learning algorithm, named FedMIM, which adopts the multi-step inertial momentum on the edge devices and enhances the local consistency for free during the training to improve the robustness of the heterogeneity. Specifically, we incorporate the weighted global gradient estimations as the inertial correction terms to guide both the local iterates and stochastic gradient estimation, which can reckon the global objective optimization on the edges' heterogeneous dataset naturally and maintain the demanding consistent iteration locally. Theoretically, we show that FedMIM achieves the $\mathcal{O}(\frac{1}{\sqrt{SKT}})$ convergence rate with a linear speedup property with respect to the number of selected clients $S$ and proper local interval $K$ in communication round $T$ without convex assumption. Empirically, we conduct comprehensive experiments on various real-world datasets and demonstrate the efficacy of the proposed FedMIM against several state-of-the-art baselines.

eess.SY↗