SearcharxivSearch

arXiv subjects

Pallab Bhattacharya

Publications and source records attributed to Pallab Bhattacharya.

12 recordsLinked to original sources

Training Video Foundation Models with NVIDIA NeMo

Video Foundation Models (VFMs) have recently been used to simulate the real world to train physical AI systems and develop creative visual experiences. However, there are significant challenges in training large-scale, high quality VFMs that can generate high-quality videos. We present a scalable, open-source VFM training pipeline with NVIDIA NeMo, providing accelerated video dataset curation, multimodal data loading, and parallelized video diffusion model training and inference. We also provide a comprehensive performance analysis highlighting best practices for efficient VFM training and inference.

cs.CV

Towards Universal Performance Modeling for Machine Learning Training on Multi-GPU Platforms

Characterizing and predicting the training performance of modern machine learning (ML) workloads on compute systems with compute and communication spread between CPUs, GPUs, and network devices is not only the key to optimization and planning but also a complex goal to achieve. The primary challenges include the complexity of synchronization and load balancing between CPUs and GPUs, the variance in input data distribution, and the use of different communication devices and topologies (e.g., NVLink, PCIe, network cards) that connect multiple compute devices, coupled with the desire for flexible training configurations. Built on top of our prior work for single-GPU platforms, we address these challenges and enable multi-GPU performance modeling by incorporating (1) data-distribution-aware performance models for embedding table lookup, and (2) data movement prediction of communication collectives, into our upgraded performance modeling pipeline equipped with inter-and intra-rank synchronization for ML workloads trained on multi-GPU platforms. Beyond accurately predicting the per-iteration training time of DLRM models with random configurations with a geomean error of 5.21% on two multi-GPU platforms, our prediction pipeline generalizes well to other types of ML workloads, such as Transformer-based NLP models with a geomean error of 3.00%. Moreover, even without actually running ML workloads like DLRMs on the hardware, it is capable of generating insights such as quickly selecting the fastest embedding table sharding configuration (with a success rate of 85%).

cs.DC

Nemotron-4 340B Technical Report

We release the Nemotron-4 340B model family, including Nemotron-4-340B-Base, Nemotron-4-340B-Instruct, and Nemotron-4-340B-Reward. Our models are open access under the NVIDIA Open Model License Agreement, a permissive model license that allows distribution, modification, and use of the models and its outputs. These models perform competitively to open access models on a wide range of evaluation benchmarks, and were sized to fit on a single DGX H100 with 8 GPUs when deployed in FP8 precision. We believe that the community can benefit from these models in various research studies and commercial applications, especially for generating synthetic data to train smaller language models. Notably, over 98% of data used in our model alignment process is synthetically generated, showcasing the effectiveness of these models in generating synthetic data. To further support open research and facilitate model development, we are also open-sourcing the synthetic data generation pipeline used in our model alignment process.

cs.CL

TPP: Transparent Page Placement for CXL-Enabled Tiered-Memory

The increasing demand for memory in hyperscale applications has led to memory becoming a large portion of the overall datacenter spend. The emergence of coherent interfaces like CXL enables main memory expansion and offers an efficient solution to this problem. In such systems, the main memory can constitute different memory technologies with varied characteristics. In this paper, we characterize memory usage patterns of a wide range of datacenter applications across the server fleet of Meta. We, therefore, demonstrate the opportunities to offload colder pages to slower memory tiers for these applications. Without efficient memory management, however, such systems can significantly degrade performance. We propose a novel OS-level application-transparent page placement mechanism (TPP) for CXL-enabled memory. TPP employs a lightweight mechanism to identify and place hot/cold pages to appropriate memory tiers. It enables a proactive page demotion from local memory to CXL-Memory. This technique ensures a memory headroom for new page allocations that are often related to request processing and tend to be short-lived and hot. At the same time, TPP can promptly promote performance-critical hot pages trapped in the slow CXL-Memory to the fast local memory, while minimizing both sampling overhead and unnecessary migrations. TPP works transparently without any application-specific knowledge and can be deployed globally as a kernel release. We evaluate TPP in the production server fleet with early samples of new x86 CPUs with CXL 1.1 support. TPP makes a tiered memory system performant as an ideal baseline (<1% gap) that has all the memory in the local tier. It is 18% better than today's Linux, and 5-17% better than existing solutions including NUMA Balancing and AutoTiering. Most of the TPP patches have been merged in the Linux v5.18 release.

cs.DC

Software-Hardware Co-design for Fast and Scalable Training of Deep Learning Recommendation Models

Deep learning recommendation models (DLRMs) are used across many business-critical services at Facebook and are the single largest AI application in terms of infrastructure demand in its data-centers. In this paper we discuss the SW/HW co-designed solution for high-performance distributed training of large-scale DLRMs. We introduce a high-performance scalable software stack based on PyTorch and pair it with the new evolution of Zion platform, namely ZionEX. We demonstrate the capability to train very large DLRMs with up to 12 Trillion parameters and show that we can attain 40X speedup in terms of time to solution over previous systems. We achieve this by (i) designing the ZionEX platform with dedicated scale-out network, provisioned with high bandwidth, optimal topology and efficient transport (ii) implementing an optimized PyTorch-based training stack supporting both model and data parallelism (iii) developing sharding algorithms capable of hierarchical partitioning of the embedding tables along row, column dimensions and load balancing them across multiple workers; (iv) adding high-performance core operators while retaining flexibility to support optimizers with fully deterministic updates (v) leveraging reduced precision communications, multi-level memory hierarchy (HBM+DDR+SSD) and pipelining. Furthermore, we develop and briefly comment on distributed data ingestion and other supporting services that are required for the robust and efficient end-to-end training in production environments.

cs.DC

FlexShard: Flexible Sharding for Industry-Scale Sequence Recommendation Models

Sequence-based deep learning recommendation models (DLRMs) are an emerging class of DLRMs showing great improvements over their prior sum-pooling based counterparts at capturing users' long term interests. These improvements come at immense system cost however, with sequence-based DLRMs requiring substantial amounts of data to be dynamically materialized and communicated by each accelerator during a single iteration. To address this rapidly growing bottleneck, we present FlexShard, a new tiered sequence embedding table sharding algorithm which operates at a per-row granularity by exploiting the insight that not every row is equal. Through precise replication of embedding rows based on their underlying probability distribution, along with the introduction of a new sharding strategy adapted to the heterogeneous, skewed performance of real-world cluster network topologies, FlexShard is able to significantly reduce communication demand while using no additional memory compared to the prior state-of-the-art. When evaluated on production-scale sequence DLRMs, FlexShard was able to reduce overall global all-to-all communication traffic by over 85%, resulting in end-to-end training communication latency improvements of almost 6x over the prior state-of-the-art approach.

cs.LG

Modeling Photocurrent Spectra of In$_{0.91}$Ga$_{0.09}$N/In$_{0.4}$Ga$_{0.6}$N Disk-in-Wire Photodiode on Silicon for $1.3$ $μ$m $-$ $1.55$ $μ$m Operation

This work reports comprehensive theoretical modeling of photocurrent spectra generated by an In$_{0.91}$Ga$_{0.09}$N/In$_{0.4}$Ga$_{0.6}$N disk-in-wire photodiode. The strain distribution is calculated by valence-force-field (VFF) model, while a realistic band structure of the InN/InGaN heterostructure is incorporated using an eight-band effective bond-orbital model (EBOM) with spin-orbit coupling neglected. The electrostatic potential is obtained from self-consistent calculation employing the non-equilibrium Green's function (NEGF) method. With the strain distribution and band profile determined, a multi-band transfer-matrix method (TMM) is used to calculate the tunneling coefficients of optically-pumped carriers in the absorbing region. The photocurrent spectra contributed by both single-photon absorption (SPA) and two-photon absorption (TPA) are calculated. The absorption coefficient is weighted by the carrier tunneling rate and the photon density-of-state (DOS) in the optical cavity formed in the nanowire region to produce the photocurrent. The calculated photocurrent spectra is in good agreement with experimental data, while physical mechanisms for the observed prominent peaks are identified and investigated.

cond-mat.mes-hall

Removing Stripes, Scratches, and Curtaining with Non-Recoverable Compressed Sensing

Highly-directional image artifacts such as ion mill curtaining, mechanical scratches, or image striping from beam instability degrade the interpretability of micrographs. These unwanted, aperiodic features extend the image along a primary direction and occupy a small wedge of information in Fourier space. Deleting this wedge of data replaces stripes, scratches, or curtaining, with more complex streaking and blurring artifacts-known within the tomography community as missing wedge artifacts. Here, we overcome this problem by recovering the missing region using total variation minimization, which leverages image sparsity based reconstruction techniques-colloquially referred to as compressed sensing-to reliably restore images corrupted by stripe like features. Our approach removes beam instability, ion mill curtaining, mechanical scratches, or any stripe features and remains robust at low signal-to-noise. The success of this approach is achieved by exploiting compressed sensings inability to recover directional structures that are highly localized and missing in Fourier Space.

cs.CV

Non-Linear Photocurrent Response to Bosonic Final State Stimulation in Microcavity Diodes

We report the optical excitation-dependent output photocurrent characteristics of GaN-based polariton diode lasers operated under reverse-bias at room temperature. The photocurrent demonstrates a non-linear enhancement at an incident optical power of ~ 1.6 mW, which is approximately equivalent to the value of polariton lasing threshold observed when the diodes are operated under forward bias conditions. This is explained in the framework of an Auger-like process of excitonic dissociation into its constituent electron-hole pairs, which can be stimulated by the occupation of the polariton lasing states. The observed effect is a remarkable manifestation of the bosonic final state stimulation in polariton lasers. A model based on the coupled kinetic equations for the free carriers, the excitonic reservoir and the polariton condensate shows a good agreement to the experimental data.

physics.app-ph

High resolution spectroscopy and narrow resonances from InGaN quantum dots in GaN nanowires

High resolution coherent nonlinear optical spectroscopy of an ensemble of red-emitting InGaN quantum dots in GaN nanowires is reported. The data show a pronounced atom-like interaction between resonant laser fields and quantum dot excitons at low temperature that is difficult to observe in the linear absorption spectrum due to inhomogeneous broadening from indium fluctuation effects. We find that the nonlinear signal persists strongly at room temperature. The robust atom-like room temperature response indicates the possibility that this material could serve as the platform for proposed excitonic based applications without the need of cryogenics.

cond-mat.mtrl-sci

Polariton Emission Characteristics of a Modulation-Doped Multiquantum-Well Microcavity Diode

The role of polariton-electron scattering on the performance characteristics of an electrically injected GaAs-based quantum well microcavity diode in the strong coupling regime has been investigated. An electron gas is introduced in the quantum wells by modulation doping with silicon dopants. It is observed that polariton-electron scattering suppresses the relaxation bottleneck in the lower polariton branch. However, it is not adequate to produce a degenerate coherent condensate at k|| ~ 0 and coherent emission.

cond-mat.mes-hall

Polariton Bose-Einstein condensate at room temperature in a Al(Ga)N nanowire-dielectric microcavity with a spatial potential trap

A spatial potential trap is formed in a 6.0 μm Al(Ga)N nanowire by varying the Al composition along its length during epitaxial growth. The polariton emission characteristics of a dielectric microcavity with the single nanowire embedded in-plane has been studied at room temperature. Excitation is provided at the Al(Ga)N end of the nanowire and polariton emission is observed from the lowest bandgap GaN region of the nanowire. Comparison of the results with those measured in an identical microcavity with an uniform GaN nanowire and having an identical exciton-photon detuning suggests evaporative cooling of the polaritons as they are transported across the trap in the Al(Ga)N nanowire. Measurement of the spectral characteristics of the polariton emission, their momentum distribution, first-order spatial coherence and time-resolved measurements of polariton cooling provide strong evidence of the formation of an equilibrium Bose-Einstein condensate, a unique state of matter in solid state systems, in the GaN region of the nanowire, at room temperature. An equilibrium condensate is not formed in the GaN nanowire dielectric microcavity without the spatial potential trap.

cond-mat.mes-hall