SearcharxivSearch

arXiv subjects

Weiwei Chu

Publications and source records attributed to Weiwei Chu.

7 recordsLinked to original sources

ScaleAcross Explorer: Exploring Communication Optimization for Scale-Across AI Model Training

The rapid scaling of large language model training requires distributing GPU resources across multiple data center buildings and regions. We refer to such paradigm as "scale-across" training. As infrastructure expands, the system design space becomes increasingly intricate, encompassing new model architectures, hardware heterogeneity, and evolving communication patterns. Drawing from Meta's production experience, we highlight the complexities of deploying training jobs across a few data centers housing hundreds of thousands of GPUs. To accelerate exploration of the large design space and to enable efficient training for frontier model development, we conduct in-depth characterization of three key design dimensions: parallelism placement, parallelism scheduling, and network layer technologies. We then propose ScaleAcross Explorer, an optimizer that considers the interplay of design dimensions and holistically optimizes scale-across training. Testbed experiments and simulations demonstrate up to 64.62% training speedups over production configuration and up to 37.59% training speedups over the state-of-the-art baseline across a wide range of design points.

cs.DC

WLB-LLM: Workload-Balanced 4D Parallelism for Large Language Model Training

In this work, we present WLB-LLM, a workLoad-balanced 4D parallelism for large language model training. We first thoroughly analyze the workload imbalance issue in LLM training and identify two primary sources of imbalance at the pipeline parallelism and context parallelism levels. Then, to address the imbalance issue, at the pipeline parallelism level, WLB-LLM incorporates a workload-aware variable-length document packing method to balance the computation and communication workload across micro-batches. Additionally, at the context parallelism level, WLB-LLM introduces a novel fine-grained per-document sharding strategy, ensuring each worker within a context parallelism group has an identical workload. Comprehensive experiments under different model scales demonstrate that WLB-LLM significantly mitigates the workload imbalance during 4D parallelism LLM training and achieves an average speedup of 1.23x when applying WLB-LLM in our internal LLM training framework.

cs.DC

CubicML: Automated ML for Large ML Systems Co-design with ML Prediction of Performance

Scaling up deep learning models has been proven effective to improve intelligence of machine learning (ML) models, especially for industry recommendation models and large language models. The co-design of large distributed ML systems and algorithms (to maximize training performance) plays a pivotal role for its success. As it scales, the number of co-design hyper-parameters grows rapidly which brings challenges to feasibly find the optimal setup for system performance maximization. In this paper, we propose CubicML which uses ML to automatically optimize training performance of large distributed ML systems. In CubicML, we use an ML model as a proxy to predict the training performance for search efficiency and performance modeling flexibility. We proved that CubicML can effectively optimize training speed of in-house ads recommendation models with 73 billion parameters and large language models up to 405 billion parameters at Meta.

cs.LG

Rankitect: Ranking Architecture Search Battling World-class Engineers at Meta Scale

Neural Architecture Search (NAS) has demonstrated its efficacy in computer vision and potential for ranking systems. However, prior work focused on academic problems, which are evaluated at small scale under well-controlled fixed baselines. In industry system, such as ranking system in Meta, it is unclear whether NAS algorithms from the literature can outperform production baselines because of: (1) scale - Meta ranking systems serve billions of users, (2) strong baselines - the baselines are production models optimized by hundreds to thousands of world-class engineers for years since the rise of deep learning, (3) dynamic baselines - engineers may have established new and stronger baselines during NAS search, and (4) efficiency - the search pipeline must yield results quickly in alignment with the productionization life cycle. In this paper, we present Rankitect, a NAS software framework for ranking systems at Meta. Rankitect seeks to build brand new architectures by composing low level building blocks from scratch. Rankitect implements and improves state-of-the-art (SOTA) NAS methods for comprehensive and fair comparison under the same search space, including sampling-based NAS, one-shot NAS, and Differentiable NAS (DNAS). We evaluate Rankitect by comparing to multiple production ranking models at Meta. We find that Rankitect can discover new models from scratch achieving competitive tradeoff between Normalized Entropy loss and FLOPs. When utilizing search space designed by engineers, Rankitect can generate better models than engineers, achieving positive offline evaluation and online A/B test at Meta scale.

cs.LG

Unusual Magnetic Properties in Layered Magnetic Topological Insulator EuSn2As2

EuSn2As2 with layered rhombohedral crystal structure is proposed to be a candidate of intrinsic antiferromagnetic (AFM) topological insulator. Here, we have investigated systematic magnetoresistance (MR) and magnetization measurements on the high quality EuSn2As2 single crystal with the magnetic field both parallel and perpendicular to (00l) plane. Both the kink of magnetic susceptibility and longitudinal resistivity reveal that EuSn2An2 undergoes an AFM transition at TN = 21 K. At T = 2 K, the magnetization exhibits two successive plateaus of ~ 5.6 μB/Eu and ~ 6.6 μB/Eu at the corresponding critical magnetic fields. Combined with the negative longitudinal MR and abnormal Hall resistance, we demonstrate that EuSn2An2 undergoes complicated magnetic transitions from an AFM state to a canted ferromagnetic (FM) state at Hc and then to a polarized FM state at Hs as the magnetic field increase.

cond-mat.mtrl-sci

Recognition of Fermi-arc states through the magnetoresistance quantum oscillations in Dirac semimetal Cd3As2 nanoplates

The disjointed Fermi-arcs in Weyl semimetals can intertwine with chiral bulk modes and participate in unusual closed magnetic orbits in the presence of a vertical magnetic field. Here we carry out the quantum oscillation study of such unusual Weyl magnetic orbits in Dirac semimetal Cd3As2, a close cousin of Weyl semimetals. We find that extra two-dimensional (2D) quantum oscillations emerge at high field, which superimpose on 3D bulk background and can be attributed to the Weyl magnetic orbits, when the thickness of nanoplates is smaller than the mean free path of the electrons. Further evidence of the 2D quantum oscillations from the Weyl magnetic orbits is provided by the nonlocal detection, which demonstrates an alternative way to study the quantum transport properties of Fermi-arcs under magnetic field.

cond-mat.mes-hall

Probing the Chiral Anomaly by Planar Hall Effect in Three-dimensional Dirac Semimetal Cd3As2 Nanoplates

Searching for exotic transport properties in new topological state of matters is an active topic. One of the most fascinating achievements is the chiral anomaly in recently discovered Weyl semimetals (WSMs), which is manifested as a negative longitudinal magnetoresistance (LMR) in the presence of a magnetic field B parallel to an electric field E. Another predicted key effect closely related to the chiral anomaly is the planar Hall effect (PHE), which has not been identified in WSMs so far. Here we carried out the planar Hall measurements on Cd3As2 nanoplates, and found that, accompanied by the large negative LMR, a PHE with non-zero transverse voltage can be developed while tilting the in-plane magnetic field B away from the electric field E. Further experiments reveal that both the PHE and the negative LMR can be suppressed synchronously by increasing the temperature, but still visible at room temperature, indicating the same origin of these two effects. The observation of PHE in Cd3As2 nanoplates gives another transport evidence for the chiral anomaly and provides a deep insight into the chiral charge pumping in Weyl Fermions system.

cond-mat.mes-hall