Searcharxiv⌕ Search

arXiv subjects

Naveen Kumar

Publications and source records attributed to Naveen Kumar.

At least 55 records · Page 3Linked to original sources

Scale MLPerf-0.6 models on Google TPU-v3 Pods

The recent submission of Google TPU-v3 Pods to the industry wide MLPerf v0.6 training benchmark demonstrates the scalability of a suite of industry relevant ML models. MLPerf defines a suite of models, datasets and rules to follow when benchmarking to ensure results are comparable across hardware, frameworks and companies. Using this suite of models, we discuss the optimizations and techniques including choice of optimizer, spatial partitioning and weight update sharding necessary to scale to 1024 TPU chips. Furthermore, we identify properties of models that make scaling them challenging, such as limited data parallelism and unscaled weights. These optimizations contribute to record performance in transformer, Resnet-50 and SSD in the Google MLPerf-0.6 submission.

cs.LG↗

Collisionless shocks in laboratory astrophysics experiments

Influence of the plasma collisions on the laser-driven collisionless shock formation and subsequent ion acceleration is studied on the basis of two different collisional algorithms and their implementations in two well-known particle-in-cell codes EPOCH and SMILEI. In this setup, an ultra-intense incident laser pulse generates hot-electrons in a thick target, launching an electrostatic shock at the laser-plasma interface while also pushing the interface through the hole-boring effect. We observe, to varying degrees, the weakening of the space-charge effects due to collisions and improvements ($\ge 10\%$) in the energy spectra of quasi-monoenergetic ions in both PIC codes EPOCH and SMILEI. These results establish the `collisionlessness' of the collisionless shocks in laboratory astrophysics experiments.

physics.plasm-ph↗

Multimodal Representation Learning using Deep Multiset Canonical Correlation

We propose Deep Multiset Canonical Correlation Analysis (dMCCA) as an extension to representation learning using CCA when the underlying signal is observed across multiple (more than two) modalities. We use deep learning framework to learn non-linear transformations from different modalities to a shared subspace such that the representations maximize the ratio of between- and within-modality covariance of the observations. Unlike linear discriminant analysis, we do not need class information to learn these representations, and we show that this model can be trained for complex data using mini-batches. Using synthetic data experiments, we show that dMCCA can effectively recover the common signal across the different modalities corrupted by multiplicative and additive noise. We also analyze the sensitivity of our model to recover the correlated components with respect to mini-batch size and dimension of the embeddings. Performance evaluation on noisy handwritten datasets shows that our model outperforms other CCA-based approaches and is comparable to deep neural network models trained end-to-end on this dataset.

cs.LG↗

Collisionless shock acceleration of quasi-monoenergetic ions in ultra-relativistic regime

Collisionless shock acceleration of carbon ions (C$^{6+}$) is investigated in the ultra-relativistic regime of laser-plasma interaction by accounting for the radiation reaction force and the pair production in particle-in-cell simulations. Both radiation reaction force and pair plasma formation tend to slow down the shock velocity, reducing the energy of the accelerated ions, albeit extending the time scales of the acceleration process. Slab plasma target achieves lower energy spread while target with a tailored density profile yields higher ion acceleration energies.

physics.plasm-ph↗

Polarized light from the transportation of a matter-antimatter beam in a plasma

A relativistic electron-positron beam propagating through a magnetized electron-ion plasma is shown to generate both circularly and linearly polarized synchrotron radiation. The degrees of circular and linear polarizations depend both on the density ratio of pair beam to background plasma and initial magnetization, and a maximum degree of circular polarization $\langle P_\textrm{circ}\rangle \approx 18\%$ is found to occur for a tenuous pair beam. We demonstrate that the generation of circularly polarized radiation is intrinsically linked to asymmetric energy dissipation of the pair beam during the filamentation instability dynamics in the electron-ion plasma. These results can help in understanding the recent observations of circularly polarized radiation from gamma-ray-bursts.

physics.plasm-ph↗

Ultraintense Attosecond Pulse Emission from Relativistic Laser-Plasma Interaction

We develop an analytical model for ultraintense attosecond pulse emission in the highly relativistic laser-plasma interaction. In this model, the attosecond pulse is emitted by a strongly compressed electron layer around the instant when the layer transverse current changes the sign and its longitudinal velocity approaches the maximum. The emitted attosecond pulse has a broadband exponential spectrum and a stabilized constant spectral phase $ψ(ω)=\pmπ/2-ψ_{A_m}$. The waveform of the attosecond pulse is also given explicitly, to our knowledge, for the first time. We validate the analytical model via particle-in-cell (PIC) simulations for both normal and oblique incidence. Based on this model, we highlight the potential to generate an isolated ultraintense phase-stabilized attosecond pulse

physics.plasm-ph↗

Super-intense Single Attosecond Pulse Generation by Plasma Gating

A robust plasma gating to generate a single ultra-intense attosecond pulse is developed. It is a manifestation of the hole-boring effect that limits the strongest attosecond pulse emission within one laser cycle. The generated pulse is characterized by a stabilized harmonic phase $ψ\approx \pmπ/2$ and a slowly decaying exponential spectrum bounded by $γ$-spike scaling and CSE scaling. The phase oscillations in low-frequency region and fluctuations in high-frequency region are discussed. We also show that the phase fluctuations in high-frequency region can be reduced by including radiation reaction force.

physics.plasm-ph↗

Grain-size dependent electric-field induced structural changes in relaxor-ferroelectric based unclamped piezoelectric grains

Polymer-piezoceramic 0-3 composites combine the flexibility of the polymers and the excellent piezoelectric properties of the ferroelectric based ceramic. While grain size of the ceramic powder is one of the important considerations in the fabrication of such composites, a correlation relating poling field induced structural changes and its possible influence on the overall piezoelectric response of the composite is still lacking. In this paper, we examine this issue on a 0-3 piezo-composite comprising of a ceramic powders of a low-lead piezoelectric alloy (x)Bi(Ni1/2Zr1/2O3-(1-x)PbTiO3 in proximity of its morphotropic phase boundary, and polyvinylidene fluoride (PVDF) as the polymer component. Composites were fabricated by fixing the volume fraction of the ceramic while varying the grain size. We found a non-monotonic variation in the piezo-response as a function of grain size. Structural analysis before and after poling of the piezo-composites revealed evidence of poling induced cubic-like to tetragonal irreversible transformation, the extent of which is dependent on the grain size.

cond-mat.mtrl-sci↗

Device Placement Optimization with Reinforcement Learning

The past few years have witnessed a growth in size and computational requirements for training and inference with neural networks. Currently, a common approach to address these requirements is to use a heterogeneous distributed environment with a mixture of hardware devices such as CPUs and GPUs. Importantly, the decision of placing parts of the neural models on devices is often made by human experts based on simple heuristics and intuitions. In this paper, we propose a method which learns to optimize device placement for TensorFlow computational graphs. Key to our method is the use of a sequence-to-sequence model to predict which subsets of operations in a TensorFlow graph should run on which of the available devices. The execution time of the predicted placements is then used as the reward signal to optimize the parameters of the sequence-to-sequence model. Our main result is that on Inception-V3 for ImageNet classification, and on RNN LSTM, for language modeling and neural machine translation, our model finds non-trivial device placements that outperform hand-crafted heuristics and traditional algorithmic methods.

cs.LG↗

Optimized plasma high harmonics generation from ultra-intense laser pulses

Plasma high harmonics generation from an extremely intense short-pulse laser is explored by including the effects of ion motion, electron-ion collisions and radiation reaction force in the plasma dynamics. The laser radiation pressure induces plasma ion motion through the hole-boring effect resulting into the frequency shifting and widening of the harmonic spectra. Classical radiation reaction force slightly mitigates the frequency broadening caused by the ion motion. Based on the results and physical considerations, parameter maps highlighting optimum regions for generating a single intense attosecond pulse and coherent XUV radiations are presented.

physics.plasm-ph↗

In-Datacenter Performance Analysis of a Tensor Processing Unit

Many architects believe that major improvements in cost-energy-performance must now come from domain-specific hardware. This paper evaluates a custom ASIC---called a Tensor Processing Unit (TPU)---deployed in datacenters since 2015 that accelerates the inference phase of neural networks (NN). The heart of the TPU is a 65,536 8-bit MAC matrix multiply unit that offers a peak throughput of 92 TeraOps/second (TOPS) and a large (28 MiB) software-managed on-chip memory. The TPU's deterministic execution model is a better match to the 99th-percentile response-time requirement of our NN applications than are the time-varying optimizations of CPUs and GPUs (caches, out-of-order execution, multithreading, multiprocessing, prefetching, ...) that help average throughput more than guaranteed latency. The lack of such features helps explain why, despite having myriad MACs and a big memory, the TPU is relatively small and low power. We compare the TPU to a server-class Intel Haswell CPU and an Nvidia K80 GPU, which are contemporaries deployed in the same datacenters. Our workload, written in the high-level TensorFlow framework, uses production NN applications (MLPs, CNNs, and LSTMs) that represent 95% of our datacenters' NN inference demand. Despite low utilization for some applications, the TPU is on average about 15X - 30X faster than its contemporary GPU or CPU, with TOPS/Watt about 30X - 80X higher. Moreover, using the GPU's GDDR5 memory in the TPU would triple achieved TOPS and raise TOPS/Watt to nearly 70X the GPU and 200X the CPU.

cs.AR↗

Active Target Localization using Low-Rank Matrix Completion and Unimodal Regression

The detection and localization of a target from samples of its generated field is a problem of interest in a broad range of applications. Often, the target field admits structural properties that enable the design of lower sample detection strategies with good performance. This paper designs a sampling and localization strategy which exploits separability and unimodality in target fields and theoretically analyzes the trade-off achieved between sampling density, noise level and convergence rate of localization. In particular, the strategy adopts an exploration-exploitation approach to target detection and utilizes the theory of low-rank matrix completion, coupled with unimodal regression, on decaying and approximately separable target fields. The assumptions on the field are fairly generic and are applicable to many decay profiles since no specific knowledge of the field is necessary, besides its admittance of an approximately rank-one representation. Extensive numerical experiments and comparisons are performed to test the efficacy and robustness of the presented approach. Numerical results suggest that the proposed strategy outperforms algorithms based on mean-shift clustering, surface interpolation and naive low-rank matrix completion with peak detection, under low sampling density.

cs.IT↗

Direct and secondary nuclear excitation with x-ray free-electron lasers

The direct and secondary nuclear excitation produced by an x-ray free electron laser when interacting with a solid-state nuclear target is investigated theoretically. When driven at the resonance energy, the x-ray free electron laser can produce direct photoexcitation. However, the dominant process in that interaction is the photoelectric effect producing a cold and very dense plasma in which also secondary processes such as nuclear excitation by electron capture may occur. We develop a realistic theoretical model to quantify the temporal dynamics of the plasma and the magnitude of the secondary excitation therein. Numerical results show that depending on the nuclear transition energy and the temperature and charge states reached in the plasma, secondary nuclear excitation by electron capture may dominate the direct photoexcitation by several orders of magnitude, as it is the case for the 4.8 keV transition from the isomeric state of $^{93}$Mo, or it can be negligible, as it is the case for the 14.4 keV Mössbauer transition in $^{57}\mathrm{Fe}$. These findings are most relevant for future nuclear quantum optics experiments at x-ray free electron laser facilities.

physics.plasm-ph↗

Toward Refactoring of DMARF and GIPSY Case Studies -- a Team 12 SOEN6471-S14 Project Report

The main significance of this document is two source systems namely GIPSY and DMARF. Intensional languages are required like GIPSY for absoluteness and forward practical investigations on the subject.DMARF mainly focuses on software architectural design and implementation on Distributed Audio recognition and its applications such as speaker identification which can run distributively on web services architecture. This mainly highlights security aspects in a distributed system, the Java data security framework (JDSF) in DMARF. ASSL (Autonomic System Specification Language) frame work is used to integrate a self-optimizing property for DMARF. GIPSY mainly depends on Higher-Order Intensional Logic (HOIL) and reflects three main goals Generality, Adaptability and Efficiency.

cs.SE↗

Novel aspects of radiation reaction in the classical and the quantum regime

This work is dedicated to the study of radiation reaction signatures in the framework of classical and quantum electrodynamics. Since there has been no distinct experimental validation of radiation reaction and its underlying equations so far and its impact is expected to be substantial for the construction of new experimental devices, e.g., quantum x-free electron lasers, a profound understanding of radiation reaction effects is of special interest. Here, we describe how the inclusion of quantum radiation reaction effects changes the dynamics of ultra-relativistic electron beams colliding with intense laser pulses significantly. Thereafter, the angular distribution of emitted radiation is demonstrated to be strongly altered in the quantum framework, if in addition to single photon emission also higher order photon emissions are considered. Furthermore, stimulated Raman scattering of an ultra-intense laser pulse in plasmas is examined and forward Raman scattering is found to be significantly increased by the inclusion of radiation reaction effects in the classical regime. The numerical simulations in this work show the feasibility of an experimental verification of the predicted effects with presently available lasers and electron accelerators.

hep-ph↗

Radiation reaction force induced nonlinear mixing of Raman sidebands of an ultra-intense laser pulse in a plasma

Stimulated Raman scattering of an ultra-intense laser pulse in plasmas is studied by perturbatively including the leading order term of the Landau-Lifshitz radiation reaction force in the equation of motion for plasma electrons. In this approximation, radiation reaction force causes phase shift in nonlinear current densities that drive the two Raman sidebands (anti-Stokes and Stokes waves), manifesting itself into the nonlinear mixing of two sidebands. This mixing results in a strong enhancement in the growth of the forward Raman scattering instability.

physics.plasm-ph↗

Influence of Surface Waves on Plasma High Harmonic Generation

The influence of surface plasma waves (SPW) on high harmonic generation (HHG) from the interaction of intense lasers with overdense plasma is analyzed. It is shown, that the surface waves lead to the emission of harmonics away from the optical axis. These off-axis harmonics violate the parity selection rules found from 1D models. Further, our investigations in the highly relativistic regime point towards the existence of a new SPW generation process.

physics.plasm-ph↗

Maximum-Likelihood Sequence Detector for Dynamic Mode High Density Probe Storage

There is an increasing need for high density data storage devices driven by the increased demand of consumer electronics. In this work, we consider a data storage system that operates by encoding information as topographic profiles on a polymer medium. A cantilever probe with a sharp tip (few nm radius) is used to create and sense the presence of topographic profiles, resulting in a density of few Tb per in.2. The prevalent mode of using the cantilever probe is the static mode that is harsh on the probe and the media. In this article, the high quality factor dynamic mode operation, that is less harsh on the media and the probe, is analyzed. The read operation is modeled as a communication channel which incorporates system memory due to inter-symbol interference and the cantilever state. We demonstrate an appropriate level of abstraction of this complex nanoscale system that obviates the need for an involved physical model. Next, a solution to the maximum likelihood sequence detection problem based on the Viterbi algorithm is devised. Experimental and simulation results demonstrate that the performance of this detector is several orders of magnitude better than the performance of other existing schemes.

cs.IT↗