Searcharxiv⌕ Search

arXiv subjects

Yan Sun

Publications and source records attributed to Yan Sun.

At least 163 records · Page 9Linked to original sources

Sparse Deep Learning for Time Series Data: Theory and Applications

Sparse deep learning has become a popular technique for improving the performance of deep neural networks in areas such as uncertainty quantification, variable selection, and large-scale network compression. However, most existing research has focused on problems where the observations are independent and identically distributed (i.i.d.), and there has been little work on the problems where the observations are dependent, such as time series data and sequential data in natural language processing. This paper aims to address this gap by studying the theory for sparse deep learning with dependent data. We show that sparse recurrent neural networks (RNNs) can be consistently estimated, and their predictions are asymptotically normally distributed under appropriate assumptions, enabling the prediction uncertainty to be correctly quantified. Our numerical results show that sparse deep learning outperforms state-of-the-art methods, such as conformal predictions, in prediction uncertainty quantification for time series data. Furthermore, our results indicate that the proposed method can consistently identify the autoregressive order for time series data and outperform existing methods in large-scale model compression. Our proposed method has important practical implications in fields such as finance, healthcare, and energy, where both accurate point estimates and prediction uncertainty quantification are of concern.

stat.ML↗

Demystifying CXL Memory with Genuine CXL-Ready Systems and Devices

The ever-growing demands for memory with larger capacity and higher bandwidth have driven recent innovations on memory expansion and disaggregation technologies based on Compute eXpress Link (CXL). Especially, CXL-based memory expansion technology has recently gained notable attention for its ability not only to economically expand memory capacity and bandwidth but also to decouple memory technologies from a specific memory interface of the CPU. However, since CXL memory devices have not been widely available, they have been emulated using DDR memory in a remote NUMA node. In this paper, for the first time, we comprehensively evaluate a true CXL-ready system based on the latest 4th-generation Intel Xeon CPU with three CXL memory devices from different manufacturers. Specifically, we run a set of microbenchmarks not only to compare the performance of true CXL memory with that of emulated CXL memory but also to analyze the complex interplay between the CPU and CXL memory in depth. This reveals important differences between emulated CXL memory and true CXL memory, some of which will compel researchers to revisit the analyses and proposals from recent work. Next, we identify opportunities for memory-bandwidth-intensive applications to benefit from the use of CXL memory. Lastly, we propose a CXL-memory-aware dynamic page allocation policy, Caption to more efficiently use CXL memory as a bandwidth expander. We demonstrate that Caption can automatically converge to an empirically favorable percentage of pages allocated to CXL memory, which improves the performance of memory-bandwidth-intensive applications by up to 24% when compared to the default page allocation policy designed for traditional NUMA systems.

cs.PF↗

Efficient Federated Prompt Tuning for Black-box Large Pre-trained Models

With the blowout development of pre-trained models (PTMs), the efficient tuning of these models for diverse downstream applications has emerged as a pivotal research concern. Although recent investigations into prompt tuning have provided promising avenues, three salient challenges persist: (1) memory constraint: the continuous growth in the size of open-source PTMs renders fine-tuning, even a fraction of their parameters, challenging for many practitioners. (2) model privacy: existing PTMs often function as public API services, with their parameters inaccessible for effective or tailored fine-tuning. (3) data privacy: the fine-tuning of PTMs necessitates high-quality datasets, which are typically localized and not shared to public. To optimally harness each local dataset while navigating memory constraints and preserving privacy, we propose Federated Black-Box Prompt Tuning (Fed-BBPT). This innovative approach eschews reliance on parameter architectures and private dataset access, instead capitalizing on a central server that aids local users in collaboratively training a prompt generator through regular aggregation. Local users leverage API-driven learning via a zero-order optimizer, obviating the need for PTM deployment. Relative to extensive fine-tuning, Fed-BBPT proficiently sidesteps memory challenges tied to PTM storage and fine-tuning on local machines, tapping into comprehensive, high-quality, yet private training datasets. A thorough evaluation across 40 datasets spanning CV and NLP tasks underscores the robustness of our proposed model.

cs.LG↗

Merging Filaments and Hub Formation in the G083.097$+$03.270 Molecular Complex

We uncover a hub-filament system associated with massive star formation in the G083.097$+$03.270. Diagnosed with simultaneous $^{12}$CO, $^{13}$CO, and C$^{18}$O line observations, the region is found to host two distinct and elongated filaments having separate velocity components, interacting spatially and kinematically, that appear to have seeded the formation of a dense hub at the intersection. A large velocity spread at the hub in addition to clear bridging feature connecting the filaments in velocity are indicating merging of filaments. Along the filaments axis, the velocity gradient reveals a global gas motion with an increasing velocity dispersion inward to the hub signifying turbulence. Altogether, the clustering of Class I sources, a high excitation temperature, a high column density, and presence of a massive outflow at the central hub suggest enhanced star formation. We propose that merging of large-scale filaments and velocity gradients along filaments are the driving factors in the mass accumulation process at the hub that have sequentially led to the massive star formation. With two giant filaments merging to coincide with a hub therein with ongoing star formation, this site serves as a benchmark for the `filaments to clusters' star-forming paradigm.

astro-ph.GA↗

Superconductivity in the bcc-type High-entropy Alloy TiHfNbTaMo

X-ray powder diffraction, electrical resistivity, magnetization, and thermodynamic measurements were conducted to investigate the structure and superconducting properties of TiHfNbTaMo, a novel high-entropy alloy possessing a valence electron count (VEC) of 4.8. The TiHfNbTaMo HEA was discovered to have a body-centered cubic structure and a microscopically homogeneous distribution of the constituent elements. This material shows type-II superconductivity with Tc = 3.42 K, lower critical field with 22.8 mT, and upper critical field with 3.95 T. Low-temperature specific heat measurements show that the alloy is a conventional s-wave type with a moderately coupled superconductor. First-principles calculations show that the density of states (DOS) of the TiHfNbTaMo alloy is dominated by hybrid d orbitals of these five metal elements. Additionally, the TiHfNbTaMo HEA exhibits three van Hove singularities. Furthermore, the VEC and the composition of the elements (especially the Nb elemental content) affect the Tc of the bcc-type HEA.

cond-mat.supr-con↗

Understanding the Kinetic Energy deposition within Molecular Clouds

According to the structures traced by $^{13}$CO spectral lines within the $^{12}$CO molecular clouds (MCs), we investigate the contributions of their internal gas motions and relative motions to the total velocity dispersions of $^{12}$CO MCs. Our samples of 2851 $^{12}$CO MCs harbor a total of 9556 individual $^{13}$CO structures, among which 1848 MCs ($\sim$ 65$\%$) have one individual $^{13}$CO structure and the other 1003 MCs ($\sim$ 35$\%$) have multiple $^{13}$CO structures. We find that the contribution of the relative motion between $^{13}$CO structures ($σ_{\rm ^{13}CO, re}$) is larger than that from their internal gas motion ($σ_{\rm ^{13}CO, in}$) in $\sim$ 62$\%$ of 1003 MCs in the `multiple' regime. In addition, we find the $σ_{\rm ^{13}CO, re}$ tends to increase with the total velocity dispersion($σ_{\rm ^{12}CO, tot}$) in our samples, especially for the MCs having multiple $^{13}$CO structures. This result provides a manifestation of the macro-turbulent within MCs, which gradually becomes the dominant way to store the kinetic energy along with the development of MC scales.

astro-ph.GA↗

A Systematic Study of Associations between Supernova Remnants and Molecular Clouds

We universally search for evidence of kinematic and spatial correlation of supernova remnant (SNR) and molecular cloud (MC) associations for nearly all SNRs in the coverage of the MWISP CO survey, i.e. 149 SNRs, 170 SNR candidates, and 18 pure pulsar wind nebulae (PWNe) in 1 deg < l < 230 deg and -5.5 deg < b < 5.5 deg. Based on high-quality and unbiased 12CO/13CO/C18O (J = 1--0) survey data, we apply automatic algorithms to identify broad lines and spatial correlations for molecular gas in each SNR region. The 91% of SNR-MC associations detected previously are identified in this paper by CO line emission. Overall, there could be as high as 80% of SNRs associated with MCs. The proportion of SNRs associated with MCs is high within the Galactic longitude less than ~50 deg. Kinematic distances of all SNRs that are associated with MCs are estimated based on systemic velocities of associated MCs. The radius of SNRs associated with MCs follows a lognormal distribution, which peaks at ~8.1 pc. The progenitor initial mass of these SNRs follows a power-law distribution with an index of ~-2.3 that is consistent with the Salpeter index of -2.35. We find that SNR-MC associations are mainly distributed in a thin disk along the Galactic plane, while a small amount distributed in a thick disk. With the height of these SNRs from the Galactic plane below ~45 pc, the distribution of the average radius relative to the height of them is roughly flat, and the average radius increases with the height when above ~45 pc.

astro-ph.GA↗

TriGait: Aligning and Fusing Skeleton and Silhouette Gait Data via a Tri-Branch Network

Gait recognition is a promising biometric technology for identification due to its non-invasiveness and long-distance. However, external variations such as clothing changes and viewpoint differences pose significant challenges to gait recognition. Silhouette-based methods preserve body shape but neglect internal structure information, while skeleton-based methods preserve structure information but omit appearance. To fully exploit the complementary nature of the two modalities, a novel triple branch gait recognition framework, TriGait, is proposed in this paper. It effectively integrates features from the skeleton and silhouette data in a hybrid fusion manner, including a two-stream network to extract static and motion features from appearance, a simple yet effective module named JSA-TC to capture dependencies between all joints, and a third branch for cross-modal learning by aligning and fusing low-level features of two modalities. Experimental results demonstrate the superiority and effectiveness of TriGait for gait recognition. The proposed method achieves a mean rank-1 accuracy of 96.0% over all conditions on CASIA-B dataset and 94.3% accuracy for CL, significantly outperforming all the state-of-the-art methods. The source code will be available at https://github.com/feng-xueling/TriGait/.

cs.CV↗

Symmetry breaking induced insulating electronic state in Pb$_{9}$Cu(PO$_4$)$_6$O

The recent experimental claim of room-temperature ambient-pressure superconductivity in a Cu-doped lead-apatite (LK-99) has ignited substantial research interest in both experimental and theoretical domains. Previous density functional theory (DFT) calculations with the inclusion of an on-site Hubbard interaction $U$ consistently predict the presence of flat bands crossing the Fermi level. This is in contrast to DFT plus dynamical mean field theory calculations, which reveal the Mott insulating behavior for the stoichiometric Pb$_{9}$Cu(PO$_4$)$_6$O compound. Nevertheless, the existing calculations are all based on the $P6_3/m$ structure, which is argued to be not the ground-state structure. Here, we revisit the electronic structure of Pb$_{9}$Cu(PO$_4$)$_6$O with the energetically more favorable $P\bar{3}$ structure, fully taking into account electronic symmetry breaking. We examine all possible configurations for Cu substituting the Pb sites. Our results show that the doped Cu atoms exhibit a preference for substituting the Pb2 sites than the Pb1 sites. In both cases, the calculated substitutional formation energies are large, indicating the difficulty in incorporating Cu at the Pb sites. We find that most of structures with Cu at the Pb2 site tend to be insulating, while the structures with both two Cu atoms at the Pb1 sites (except one configuration) are predicted to be metallic by DFT+$U$ calculations. However, when accounting for the electronic symmetry breaking, some Cu-doped configurations previously predicted to be metallic (including the structure studied in previous DFT+$U$ calculations) become insulating. Our work highlights the importance of symmetry breaking in obtaining correct electronic state for Pb$_{9}$Cu(PO$_4$)$_6$O, thereby reconciling previous DFT+$U$ and DFT+DMFT calculations.

cond-mat.mtrl-sci↗

Distributions and Physical Properties of Molecular Clouds in the Third Galactic Quadrant: $l$ = [219.75, 229.75]$^\circ$ and $b$ = [-5.25, 5.25]$^\circ$

We present the results of an unbiased $^{12}$CO/$^{13}$CO/C$^{18}$O ($J$ = 1-0) survey in a portion of the third Galactic quadrant (TGQ): $l$ = [219.75, 229.75]$^\circ$ and $b$ = [-5.25, 5.25]$^\circ$. The high-resolution and high-sensitivity data sets help to unravel the distributions and physical properties of the molecular clouds (MCs) in the mapped area. In the LSR velocity range from -1 to 85 km/s, the molecular material successfully traces the Local, Perseus, and Outer arms. In the TGQ, the Outer arm appears to be more prominent than that in the second Galactic quadrant (SGQ), but the Perseus arm is not as conspicuous as that in the SGQ. A total of 1,502 $^{12}$CO, 570 $^{13}$CO, and 53 C$^{18}$O molecular structures are identified, spanning over $\sim2$ and $\sim6$ orders of magnitude in size and mass, respectively. Tight mass-radius correlations and virial parameter-mass anticorrelations are observable. Yet, it seems that no clear correlations between velocity dispersion and effective radius can be found over the full dynamic range. The vertical distribution of the MCs renders evident pictures of the Galactic warp and flare.

astro-ph.GA↗

Temporal Sentence Grounding in Streaming Videos

This paper aims to tackle a novel task - Temporal Sentence Grounding in Streaming Videos (TSGSV). The goal of TSGSV is to evaluate the relevance between a video stream and a given sentence query. Unlike regular videos, streaming videos are acquired continuously from a particular source, and are always desired to be processed on-the-fly in many applications such as surveillance and live-stream analysis. Thus, TSGSV is challenging since it requires the model to infer without future frames and process long historical frames effectively, which is untouched in the early methods. To specifically address the above challenges, we propose two novel methods: (1) a TwinNet structure that enables the model to learn about upcoming events; and (2) a language-guided feature compressor that eliminates redundant visual frames and reinforces the frames that are relevant to the query. We conduct extensive experiments using ActivityNet Captions, TACoS, and MAD datasets. The results demonstrate the superiority of our proposed methods. A systematic ablation study also confirms their effectiveness.

cs.CV↗

Hardness and fracture toughness models by symbolic regression

Superhard materials with good fracture toughness have found wide industrial applications, which necessitates the development of accurate hardness and fracture toughness models for efficient materials design. Although several macroscopic models have been proposed, they are mostly semiempirical based on prior knowledge or assumptions, and obtained by fitting limited experimental data. Here, through an unbiased and explanatory symbolic regression technique, we built a macroscopic hardness model and fracture toughness model, which only require shear and bulk moduli as inputs. The developed hardness model was trained on an extended dataset, which not only includes cubic systems, but also contains non-cubic systems with anisotropic elastic properties. The obtained models turned out to be simple, accurate, and transferable. Moreover, we assessed the performance of three popular deep learning models for predicting bulk and shear moduli, and found that the crystal graph convolutional neural network and crystal explainable property predictor perform almost equally well, both better than the atomistic line graph neural network. By combining the machine-learned bulk and shear moduli with the hardness and fracture toughness prediction models, potential superhard materials with good fracture toughness can be efficiently screened out through high-throughput calculations.

cond-mat.mtrl-sci↗

First-principles study on the electronic structure of Pb$_{10-x}$Cu$_x$(PO$_4$)$_6$O ($x$=0, 1)

Recently, Lee et al. reported the experimental discovery of room-temperature ambient-pressure superconductivity in a Cu-doped lead-apatite (LK-99) (arXiv:2307.12008, arXiv:2307.12037). Remarkably, the superconductivity persists up to 400 K at ambient pressure. Despite strong experimental evidence, the electronic structure of LK-99 has not yet been studied. Here, we investigate the electronic structures of LK-99 and its parent compound using first-principles calculations, aiming to elucidate the doping effects of Cu. Our results reveal that the parent compound Pb$_{10}$(PO$_4$)$_6$O is an insulator, while Cu doping induces an insulator-metal transition and thus volume contraction. The band structures of LK-99 around the Fermi level are featured by a half-filled flat band and a fully-occupied flat band. These two flat bands arise from both the $2p$ orbitals of $1/4$-occupied O atoms and the hybridization of the $3d$ orbitals of Cu with the $2p$ orbitals of its nearest-neighboring O atoms. Interestingly, we observe four van Hove singularities on these two flat bands. Furthermore, we show that the flat band structures can be tuned by including electronic correlation effects or by doping different elements. We find that among the considered doping elements (Ni, Cu, Zn, Ag, and Au), both Ni and Zn doping result in the gap opening, whereas Au exhibits doping effects more similar to Cu than Ag. Our work provides a foundation for future studies on the role of unique electronic structures of LK-99 in superconductivity.

cond-mat.mtrl-sci↗

Efficient Federated Learning via Local Adaptive Amended Optimizer with Linear Speedup

Adaptive optimization has achieved notable success for distributed learning while extending adaptive optimizer to federated Learning (FL) suffers from severe inefficiency, including (i) rugged convergence due to inaccurate gradient estimation in global adaptive optimizer; (ii) client drifts exacerbated by local over-fitting with the local adaptive optimizer. In this work, we propose a novel momentum-based algorithm via utilizing the global gradient descent and locally adaptive amended optimizer to tackle these difficulties. Specifically, we incorporate a locally amended technique to the adaptive optimizer, named Federated Local ADaptive Amended optimizer (\textit{FedLADA}), which estimates the global average offset in the previous communication round and corrects the local offset through a momentum-like term to further improve the empirical training speed and mitigate the heterogeneous over-fitting. Theoretically, we establish the convergence rate of \textit{FedLADA} with a linear speedup property on the non-convex case under the partial participation settings. Moreover, we conduct extensive experiments on the real-world dataset to demonstrate the efficacy of our proposed \textit{FedLADA}, which could greatly reduce the communication rounds and achieves higher accuracy than several baselines.

cs.LG↗

Anisotropic in-plane heat transport of Kitaev magnet Na$_2$Co$_2$TeO$_6$

We report a study on low-temperature heat transport of Kitaev magnet Na$_2$Co$_2$TeO$_6$, with the heat current and magnetic fields along the honeycomb spin layer (the $ab$ plane). The zero-field thermal conductivity of $κ^a_{xx}$ and $κ^{a*}_{xx}$ display similar temperature dependence and small difference in their magnitudes; whereas, their magnetic field (parallel to the heat current) dependence are quite different and are related to the field-induced magnetic transitions. The $κ^a_{xx}(B)$ data for $B \parallel a$ at very low temperatures have an anomaly at 10.25--10.5 T, which reveals an unexplored magnetic transition. The planar thermal Hall conductivity $κ^a_{xy}$ and $κ^{a*}_{xy}$ show very weak signals at low fields and rather large values with sign change at high fields. This may point to a possible magnetic structure transition or the change of the magnon band topology that induces a radical change of magnon Berry curvature distribution before entering the spin polarized state. These results put clear constraints on the high-field phase and the theoretical models for Na$_2$Co$_2$TeO$_6$.

cond-mat.str-el↗

Robust anomalous Hall effect in ferromagnetic metal under high pressure

Recently, the giant intrinsic anomalous Hall effect (AHE) has been observed in the materials with kagome lattice. In this study, we systematically investigate the influence of high pressure on the AHE in the ferromagnet LiMn6Sn6 with clean Mn kagome lattice. Our in-situ high-pressure Raman spectroscopy indicates that the crystal structure of LiMn6Sn6 maintains a hexagonal phase under high pressures up to 8.51 GPa. The anomalous Hall conductivity (AHC) σxyA remains around 150 Ω-1 cm-1, dominated by the intrinsic mechanism. Combined with theoretical calculations, our results indicate that the stable AHE under pressure in LiMn6Sn6 originates from the robust electronic and magnetic structure.

cond-mat.mtrl-sci↗

FedSpeed: Larger Local Interval, Less Communication Round, and Higher Generalization Accuracy

Federated learning is an emerging distributed machine learning framework which jointly trains a global model via a large number of local devices with data privacy protections. Its performance suffers from the non-vanishing biases introduced by the local inconsistent optimal and the rugged client-drifts by the local over-fitting. In this paper, we propose a novel and practical method, FedSpeed, to alleviate the negative impacts posed by these problems. Concretely, FedSpeed applies the prox-correction term on the current local updates to efficiently reduce the biases introduced by the prox-term, a necessary regularizer to maintain the strong local consistency. Furthermore, FedSpeed merges the vanilla stochastic gradient with a perturbation computed from an extra gradient ascent step in the neighborhood, thereby alleviating the issue of local over-fitting. Our theoretical analysis indicates that the convergence rate is related to both the communication rounds $T$ and local intervals $K$ with a upper bound $\small \mathcal{O}(1/T)$ if setting a proper local interval. Moreover, we conduct extensive experiments on the real-world dataset to demonstrate the efficiency of our proposed FedSpeed, which performs significantly faster and achieves the state-of-the-art (SOTA) performance on the general FL experimental settings than several baselines. Our code is available at \url{https://github.com/woodenchild95/FL-Simulator.git}.

cs.LG↗

Improving the Model Consistency of Decentralized Federated Learning

To mitigate the privacy leakages and communication burdens of Federated Learning (FL), decentralized FL (DFL) discards the central server and each client only communicates with its neighbors in a decentralized communication network. However, existing DFL suffers from high inconsistency among local clients, which results in severe distribution shift and inferior performance compared with centralized FL (CFL), especially on heterogeneous data or sparse communication topology. To alleviate this issue, we propose two DFL algorithms named DFedSAM and DFedSAM-MGS to improve the performance of DFL. Specifically, DFedSAM leverages gradient perturbation to generate local flat models via Sharpness Aware Minimization (SAM), which searches for models with uniformly low loss values. DFedSAM-MGS further boosts DFedSAM by adopting Multiple Gossip Steps (MGS) for better model consistency, which accelerates the aggregation of local flat models and better balances communication complexity and generalization. Theoretically, we present improved convergence rates $\small \mathcal{O}\big(\frac{1}{\sqrt{KT}}+\frac{1}{T}+\frac{1}{K^{1/2}T^{3/2}(1-λ)^2}\big)$ and $\small \mathcal{O}\big(\frac{1}{\sqrt{KT}}+\frac{1}{T}+\frac{λ^Q+1}{K^{1/2}T^{3/2}(1-λ^Q)^2}\big)$ in non-convex setting for DFedSAM and DFedSAM-MGS, respectively, where $1-λ$ is the spectral gap of gossip matrix and $Q$ is the number of MGS. Empirically, our methods can achieve competitive performance compared with CFL methods and outperform existing DFL methods.

cs.LG↗