SearcharxivSearch

arXiv subjects

Yonatan Piasetzky

Publications and source records attributed to Yonatan Piasetzky.

4 recordsLinked to original sources

High-speed Networking for Giga-Scale AI Factories

As distributed model training scales to span hundreds of thousands of GPUs, scale-out networks face unprecedented performance and efficiency demands. NVIDIA Spectrum-X Ethernet has been designed from the ground up to achieve predictable and stable network performance with high utilization and low latency. This paper presents the Spectrum-X multiplane architecture, which replaces hierarchical depth with topological parallelism, and introduces hardware-accelerated load balancing in NICs and switches as the key architectural approach to provide fast reaction to highly dynamic network conditions at the microsecond timescales that AI training workloads demand. We describe the motivation, design principles, evaluation methodology and performance on state-of-the-art benchmarks, as well as the lessons we learned from deploying and debugging Spectrum-X networks in large-scale systems. Our evaluation highlights production-grade AI infrastructure performance across three core dimensions: 98% of the theoretical line rate with low jitter-free latency; strong cross-tenant isolation for concurrent workloads; robust, capacity-proportional bisection bandwidth and 7% latency increase for 10% fabric link failures; and rapid reaction to host and fabric link flaps during LLM training workloads.

cs.NI

OptiReduce: Resilient and Tail-Optimal AllReduce for Distributed Deep Learning in the Cloud

We present OptiReduce, a new collective-communication system for the cloud with bounded, predictable completion times for deep-learning jobs in the presence of varying computation (stragglers) and communication (congestion and gradient drops) variabilities. OptiReduce exploits the inherent resiliency and the stochastic nature of distributed deep-learning (DDL) training and fine-tuning to work with approximated (or lost) gradients -- providing an efficient balance between (tail) performance and the resulting accuracy of the trained models. Exploiting this domain-specific characteristic of DDL, OptiReduce introduces (1) mechanisms (e.g., unreliable bounded transport with adaptive timeout) to improve the DDL jobs' tail execution time, and (2) strategies (e.g., Transpose AllReduce and Hadamard Transform) to mitigate the impact of gradient drops on model accuracy. Our evaluation shows that OptiReduce achieves 70% and 30% faster time-to-accuracy (TTA), on average, when operating in shared, cloud environments (e.g., CloudLab) compared to Gloo and NCCL, respectively.

cs.DC

Segmented Composite Design of Robust Single-Qubit Quantum Gates

Error mitigation schemes and error-correcting codes have been the center of much effort in quantum information processing research over the last few decades. While most of the successful proposed schemes for error mitigation are perturbative in the noise and assume deterministic systematic errors, studies of the problem considering the full noise and errors distribution are still scarce. In this work, we introduce an error mitigation scheme for robust single-qubit unitary gates based on composite segmented design, which accounts for the full distribution of the physical noise and errors in the system. We provide two optimization approaches to construct these robust segmented gates: perturbative and non-perturbative, that addresses all orders of errors. We demonstrate our scheme in the photonics realm for the dual-rail directional couplers realization. We show that the 3-segmented composite design for the fundamental single-qubits unitary operations reduces the error by an order of magnitude for a realistic distribution of errors, and that the two approaches are compatible for small errors. This is shown to significantly reduce the overhead of modern error correction codes. Our methods are rather general and can be applied to other realizations of quantum information processing units.

quant-ph

Fault-Tolerant Directional Couplers for State Manipulation in Silicon Photonic-Integrated Circuits

Photonic integrated circuits play a central role in current and future applications such as communications, sensing, ranging, and information processing. Photonic quantum computing will also likely require an integrated optics architecture for improved stability, scalability, and performance. Fault-tolerant quantum computing mandates very accurate and robust quantum gates. In this work, we demonstrate high-fidelity directional couplers for single-qubit gates in photonic integrated waveguides, utilizing a novel scheme of detuning-modulated composite segments. Specific designs for reduced sensitivity to wavelength variations and real-world geometrical fabrication errors in waveguides width and depth are presented. Enhanced wavelength tolerance is demonstrated experimentally. The concept shows great promise for scaling high fidelity gates as part of integrated quantum optics architectures.

physics.optics