SearcharxivSearch

arXiv subjects

Xuan Peng

Publications and source records attributed to Xuan Peng.

6 recordsLinked to original sources

DawnPiper: A Memory-scablable Pipeline Parallel Training Framework

Pipeline parallelism is a crucial paradigm for large-scale model training. However, imbalances in memory footprint across stages can lead to significant GPU memory wastage, limiting the model sizes that pipeline parallelism can effectively support. In this paper, we introduce DawnPiper, a memory-scalable pipeline parallel training framework. Firstly, we develop a DL compilation-based profiling method that transforms the model into a fine-grained computation graph. This refinement gives us a finer granularity of model partitioning and memory optimization while facilitating automatic code generation. Based on observed memory usage characteristics, we derive a performance-optimal theorem for pipeline parallel partitioning that substantially reduces the partition search space. Secondly, we propose a binary pipeline partitioning algorithm and utilize a cost-model based memory optimization approach to efficiently identify nearly optimal pipeline parallel strategy. DawnPiper achieves up to a 4x and 11x increase in trainable maximum batch size compared to vPipe and PipeDream, respectively, and provides up to a 1.5x performance speedup compared to vPipe.

cs.DC

CFP: Efficient Optimization of Intra-Operator Parallelism Plans for Large Model Training

Optimizing the parallel training of large models requires exploring intra-operator parallelism plans for a computation graph that typically contains tens of thousands of primitive operators. While the optimization of parallel data processing graphs has been extensively researched in database systems, the vast search space makes it challenging to apply traditional database query optimization methods and algorithms. This paper introduces CFP, an optimization system for intra-operator parallelism that significantly reduces the complexity of searching for parallelism plans by leveraging two structural patterns found in large models. First, we identify parallel-preserving subgraphs, which ensure that the optimal global plan assigns the same parallel strategy to all operators within the subgraph. This approach allows us to avoid enumerating all possible combinations of parallel strategies for these operators. Second, we recognize repetitive subgraph patterns within the large computational graph, enabling us to profile a moderate number of representative subgraphs and accurately estimate the cost of parallelism plans with low overhead. With the significantly reduced search space, we can employ dynamic programming to search for the optimized parallelism plan. In our experiments, we demonstrate that CFP achieves significant speedups compared to the state-of-the-art framework for large models like GPT and LLAMA.

cs.DC

A broadband hyperspectral image sensor with high spatio-temporal resolution

Hyperspectral imaging provides high-dimensional spatial-temporal-spectral information revealing intrinsic matter characteristics. Here we report an on-chip computational hyperspectral imaging framework with high spatial and temporal resolution. By integrating different broadband modulation materials on the image sensor chip, the target spectral information is non-uniformly and intrinsically coupled on each pixel with high light throughput. Using intelligent reconstruction algorithms, multi-channel images can be recovered from each frame, realizing real-time hyperspectral imaging. Following such a framework, we for the first time fabricated a broadband VIS-NIR (400-1700 nm) hyperspectral imaging sensor using photolithography, with an average light throughput of 74.8% and 96 wavelength channels. The demonstrated resolution is 1,024*1,024 pixels at 124 fps. We demonstrated its wide applications including chlorophyll and sugar quantification for intelligent agriculture, blood oxygen and water quality monitoring for human health, textile classification and apple bruise detection for industrial automation, and remote lunar detection for astronomy. The integrated hyperspectral image sensor weighs only tens of grams, and can be assembled on various resource-limited platforms or equipped with off-the-shelf optical systems. The technique transforms the challenge of high-dimensional imaging from a high-cost manufacturing and cumbersome system to one that is solvable through on-chip compression and agile computation.

physics.optics

Ternary Nb3Sn superconductors with artificial pinning centers and high upper critical fields

In this letter we demonstrate the development of ternary Nb3Sn multifilamentary conductors with artificial pinning centers (APC) which achieve high critical fields. These recently-developed conductors were tested in a 31 T magnet, and the results showed that their upper critical field (Bc2) values at 4.2 K are 27-28 T, and irreversible field (Birr) values are above 26 T, values similar to or higher than those of best RRP conductors. The non-Cu Jc has been brought to nearly 1200 A/mm2 at 16 T and 4.2 K, comparable to RRP, in spite of the fact that the fine-grain Nb3Sn fractions in filaments are still low (20-30%) and the grain sizes are still not fully refined (70-80 nm) due to conductor designs and heat treatments that are not yet optimized. The Nb3Sn layer Jc at 4.2 K, 16 T is 4710 A/mm2 for the APC wire with 1%Zr, about 2.5 times higher than RRP conductors, in spite of the fact that its grain size is not yet fully refined due to insufficient oxygen and unoptimized heat treatment. An analysis is presented about the non-Cu Jc that can be achieved by further optimizing the APC conductors and their heat treatments.

cond-mat.supr-con

On better training the infinite restricted Boltzmann machines

The infinite restricted Boltzmann machine (iRBM) is an extension of the classic RBM. It enjoys a good property of automatically deciding the size of the hidden layer according to specific training data. With sufficient training, the iRBM can achieve a competitive performance with that of the classic RBM. However, the convergence of learning the iRBM is slow, due to the fact that the iRBM is sensitive to the ordering of its hidden units, the learned filters change slowly from the left-most hidden unit to right. To break this dependency between neighboring hidden units and speed up the convergence of training, a novel training strategy is proposed. The key idea of the proposed training strategy is randomly regrouping the hidden units before each gradient descent step. Potentially, a mixing of infinite many iRBMs with different permutations of the hidden units can be achieved by this learning method, which has a similar effect of preventing the model from over-fitting as the dropout. The original iRBM is also modified to be capable of carrying out discriminative training. To evaluate the impact of our method on convergence speed of learning and the model's generalization ability, several experiments have been performed on the binarized MNIST and CalTech101 Silhouettes datasets. Experimental results indicate that the proposed training strategy can greatly accelerate learning and enhance generalization ability of iRBMs.

cs.LG

Internally Oxidized Nb3Sn Strands with Fine Grain Size and High Critical Current Density

Nb3Sn superconducting strands are the most practical conductors to generate high magnetic fields (12-16 T), and thus have significant applications in nuclear magnetic resonance (NMR), and great potential for fusion reactors and particle accelerator magnets. High critical current density (Jc) is a key parameter for such applications. Significant efforts towards optimization of various factors led to an 80% improvement in Jc from the early 1990s to 2003, when the 4.2 K, 12 T non-matrix Jc reached 3000 A/mm2 (corresponding to 5000 A/mm2 in Nb3Sn layer Jc). However, further efforts over the past decade have failed to bring about further increase beyond this level, leading some researchers to conclude that the Jc of conventional Nb3Sn strands had reached its maximum. Here, however, by applying an internal oxidation method, we reduce the grain size by a factor of three and nearly double the 12 T Jc. In this method, a Nb3Sn strand is fabricated with Nb-Zr alloy as starting material; with oxygen supplied properly via an oxide powder, the Zr atoms in the Nb-Zr alloy are internally oxidized, forming fine intra-granular and inter-granular ZrO2 particles in Nb3Sn layer, which effectively refine Nb3Sn grain size. At a reaction temperature of 625 {\deg}C, grain size down to 20-50 nm (36 nm on average) has been achieved. For this sample the 4.2 K, 12 T Nb3Sn layer Jc reached 9600 A/mm2.

cond-mat.mtrl-sci