SearcharxivSearch

arXiv subjects

Le Qin

Publications and source records attributed to Le Qin.

7 recordsLinked to original sources

DualSparse-MoE: Coordinating Tensor/Neuron-Level Sparsity with Expert Partition and Reconstruction

Mixture of Experts (MoE) has become a mainstream architecture for building Large Language Models (LLMs) by reducing per-token computation while enabling model scaling. It can be viewed as partitioning a large Feed-Forward Network (FFN) at the tensor level into fine-grained sub-FFNs, or experts, and activating only a sparse subset for each input. While this sparsity improves efficiency, MoE still faces substantial challenges due to their massive computational scale and unpredictable activation patterns. To enable efficient MoE deployment, we identify dual sparsity at the tensor and neuron levels in pre-trained MoE modules as a key factor for both accuracy and efficiency. Unlike prior work that increases tensor-level sparsity through finer-grained expert design during pre-training, we introduce post-training expert partitioning to induce such sparsity without retraining. This preserves the mathematical consistency of model transformations and enhances both efficiency and accuracy in subsequent fine-tuning and inference. Building upon this, we propose DualSparse-MoE, an inference system that integrates dynamic tensor-level computation dropping with static neuron-level reconstruction to deliver significant efficiency gains with minimal accuracy loss. Experimental results show that enforcing an approximate 25% drop rate with our approach reduces average accuracy by only 0.08%-0.28% across three prevailing MoE models, while nearly all degrees of computation dropping consistently yield proportional computational speedups. Furthermore, incorporating load-imbalance awareness into expert parallelism achieves a 1.41x MoE module speedup with just 0.5% average accuracy degradation.

cs.LG

MoC-System: Efficient Fault Tolerance for Sparse Mixture-of-Experts Model Training

As large language models continue to scale up, distributed training systems have expanded beyond 10k nodes, intensifying the importance of fault tolerance. Checkpoint has emerged as the predominant fault tolerance strategy, with extensive studies dedicated to optimizing its efficiency. However, the advent of the sparse Mixture-of-Experts (MoE) model presents new challenges due to the substantial increase in model size, despite comparable computational demands to dense models. In this work, we propose the Mixture-of-Checkpoint System (MoC-System) to orchestrate the vast array of checkpoint shards produced in distributed training systems. MoC-System features a novel Partial Experts Checkpointing (PEC) mechanism, an algorithm-system co-design that strategically saves a selected subset of experts, effectively reducing the MoE checkpoint size to levels comparable with dense models. Incorporating hybrid parallel strategies, MoC-System involves fully sharded checkpointing strategies to evenly distribute the workload across distributed ranks. Furthermore, MoC-System introduces a two-level checkpointing management method that asynchronously handles in-memory snapshots and persistence processes. We build MoC-System upon the Megatron-DeepSpeed framework, achieving up to a 98.9% reduction in overhead for each checkpointing process compared to the original method, during MoE model training with ZeRO-2 data parallelism and expert parallelism. Additionally, extensive empirical analyses substantiate that our methods enhance efficiency while maintaining comparable model accuracy, even achieving an average accuracy increase of 1.08% on downstream tasks.

cs.DC

Shortcut-connected Expert Parallelism for Accelerating Mixture-of-Experts

Expert parallelism has emerged as a key strategy for distributing the computational workload of sparsely-gated mixture-of-experts (MoE) models across multiple devices, enabling the processing of increasingly large-scale models. However, the All-to-All communication inherent to expert parallelism poses a significant bottleneck, limiting the efficiency of MoE models. Although existing optimization methods partially mitigate this issue, they remain constrained by the sequential dependency between communication and computation operations. To address this challenge, we propose ScMoE, a novel shortcut-connected MoE architecture integrated with an overlapping parallelization strategy. ScMoE decouples communication from its conventional sequential ordering, enabling up to 100% overlap with computation. Compared to the prevalent top-2 MoE baseline, ScMoE achieves speedups of 1.49 times in training and 1.82 times in inference. Moreover, our experiments and analyses indicate that ScMoE not only achieves comparable but in some instances surpasses the model quality of existing approaches.

cs.LG

Manipulating Hubbard-type Coulomb blockade effect of metallic wires embedded in an insulator

Correlated states emerge in low-dimensional systems owing to enhanced Coulomb interactions. Elucidating these states requires atomic scale characterization and delicate control capabilities. In this study, spectroscopic imaging-scanning tunneling microscopy was employed to investigate the correlated states residing in the one-dimensional electrons of the monolayer and bilayer MoSe2 mirror twin boundary (MTB). The Coulomb energies, determined by the wire length, drive the MTB into two types of ground states with distinct respective out-of-phase and in-phase charge orders. The two ground states can be reversibly converted through a metastable zero-energy state with in situ voltage pulses, which tunes the electron filling of the MTB via a polaronic process, as substantiated by first-principles calculations. Our modified Hubbard model reveals the ground states as correlated insulators from an on-site U-originated Coulomb interaction, dubbed Hubbard-type Coulomb blockade effect. Our work sets a foundation for understanding correlated physics in complex systems and for tailoring quantum states for nano-electronics applications.

cond-mat.mes-hall

Realization of AlSb in the double layer honeycomb structure: a robust new class of two-dimensional material

Exploring new two-dimensional (2D) van der Waals (vdW) systems is at the forefront of materials physics. Here, through molecular beam epitaxy on graphene-covered SiC(0001), we report successful growth of AlSb in the double-layer honeycomb (DLHC) structure, a 2D vdW material which has no direct analogue to its 3D bulk and is predicted kinetically stable when freestanding. The structural morphology and electronic structure of the experimental 2D AlSb are characterized with spectroscopic imaging scanning tunneling microscopy and cross-sectional imaging scanning transmission electron microscopy, which compare well to the proposed DLHC structure. The 2D AlSb exhibits a bandgap of 0.93 eV versus the predicted 1.06 eV, which is substantially smaller than the 1.6 eV of bulk. We also attempt the less-stable InSb DLHC structure; however, it grows into bulk islands instead. The successful growth of a DLHC material here opens the door for the realization of a large family of novel 2D DLHC traditional semiconductors with unique excitonic, topological, and electronic properties.

cond-mat.mtrl-sci

Possible phason-polaron effect on purely one dimensional charge order of Mo6Se6 nanowires

In one-dimensional (1D) metallic systems, the diverging electron susceptibility and electron-phonon coupling collaboratively drive the electrons into a charge density wave (CDW) state. However, strictly 1D system is unstable against perturbations, whose effect on CDW order requires clarification ideally with altered coupling to surroundings. Here, we fabricate such a system with nanowires of Mo6Se6 bundles, which are either attached to edges of monolayer MoSe2 or isolated freely, by post-annealing the preformed MoSe2. Using scanning tunneling microscopy (STM), we visualized charge modulations and CDW gaps with prominent coherent peaks in the edge-attached nanowires. Astonishingly, the CDW order becomes suppressed in the isolated nanowires, showing CDW correlation gaps without coherent peaks. The contrasting behavior, as revealed with theoretical modeling, is interpreted as the effect of phason-polarons on the 1D CDW state. Our work elucidates a possibly unprecedented many body effect that may be generic to strictly 1D system but undermined in quasi-1D system.

cond-mat.mes-hall

Dimensional Crossover and Topological Phase Transition in Dirac Semimetal Na3Bi Films

Three-dimensional (3D) topological Dirac semimetal, when thinned down to 2D few layers, is expected to possess gapped Dirac nodes via quantum confinement effect and concomitantly display the intriguing quantum spin Hall (QSH) insulator phase. However, the 3D-to-2D crossover and the associated topological phase transition, which is valuable for understanding the topological quantum phases, remain unexplored. Here, we synthesize high-quality Na3Bi thin films with R3*R3 reconstruction on graphene, and systematically characterize their thickness-dependent electronic and topological properties by scanning tunneling microscopy/spectroscopy in combination with first-principles calculations. We demonstrate that Dirac gaps emerge in Na3Bi films, providing spectroscopic evidences of dimensional crossover from a 3D semimetal to a 2D topological insulator. Importantly, the Dirac gaps are revealed to be of sizable magnitudes on 3 and 4 monolayers (72 and 65 meV, respectively) with topologically nontrivial edge states. Moreover, the Fermi energy of a Na3Bi film can be tuned via certain growth process, thus offering a viable way for achieving charge neutrality in transport. The feasibility of controlling Dirac gap opening and charge neutrality enables realizing intrinsic high-temperature QSH effect in Na3Bi films and achieving potential applications in topological devices.

cond-mat.mtrl-sci