SearcharxivSearch

arXiv subjects

Fangyu Wang

Publications and source records attributed to Fangyu Wang.

7 recordsLinked to original sources

Nonlinear Photonic Tripartite Phase

Anderson localization is usually understood as a transition between extended and localized phases, with criticality confined to a single mobility edge. Recent advances predict that quasiperiodic systems can instead host a finite critical window bounded by mobility edges, in which localized, critical and extended states coexist. Yet both the experimental realization of this regime and whether interactions can provide controlled access to it remain unknown. Here, we realize such a tripartite phase in a nonlinear quasiperiodic photonic lattice and show that Kerr nonlinearity, acting as an effective interaction, enables state-selective access to the critical window. By tracking wavepacket dynamics, we distinguish localized, critical and extended transport regimes and uncover a state-selective response: rather than simply reinforcing localization through self-trapping, weak nonlinearity drives low-energy localized states into the critical window, whereas stronger nonlinearity restores localization. By contrast, critical, extended and high-energy localized states evolve monotonically towards self-trapped behaviour. Our results reveal a state-selective mechanism by which interactions provide controlled access to a pre-existing critical window in quasiperiodic systems.

cond-mat.mes-hall

Massive coherent equipartition of light by the geometric phase of null space

Light source is a foundational to photonic science and technology. However, a significant challenge remains in generating and distributing coherent light from a single on-chip source with high phase stability across multiple channels. Integrated lasers typically operate independently, and conventional splitters (e.g., multi-mode interferometers) do not guarantee the phase coherence required for advanced applications. Here, we report a purely geometric scheme for achieving massive equipartition of coherent light on a photonic chip by leveraging the geometric phases of a null space spanned by degenerate states with zero eigenvalue. The evolution of the null space maps to real-space rotation described by the special orthogonal group SO(N), thus enabling precise and scalable control over light distribution by engineering the system parameters. We experimentally realize up to one-to-nine equipartition of light on a waveguide array fabricated on a glass-based photonic chip. The framework can be upscaled for one-to-N light distribution. This work establishes a versatile and scalable platform for integrated coherent light sources, paving the way for integrated photonic applications such as quantum photonics and optical computing.

physics.optics

Quantum boomerang effect of light

The quantum boomerang effect is a counterintuitive phenomenon where a wave packet, despite having an initial momentum, returns to its starting position in a disordered medium. However, up to now, the experimental exploration of this effect remains largely unexplored. Here, we report the experimental observation of the quantum boomerang effect of light. Our experiment is based on a one-dimensional disordered photonic lattice, which is composed of on-chip optical waveguides with engineered on-site random potential. We first characterize this optical disordered system by demonstrating the static Anderson localization of light beams. Next, through launching a kinetic light beam into the system, we observe that the light beam first moves away from its starting point, arrives at a maximum value, reverses its direction, and returns to its original position over time, confirming the observation of the quantum boomerang effect of light. Surprisingly, we find that optical loss, usually considered to be detrimental to optical experiments, can enhance the quantum boomerang effect by accelerating the light back to its original position. Our work provides new insights into the light-matter interactions in disordered medium and opens an avenue for future study of this phenomenon in nonlinear and many-photon contexts.

physics.optics

On the Performance and Memory Footprint of Distributed Training: An Empirical Study on Transformers

Transformer models have emerged as potent solutions to a wide array of multidisciplinary challenges. The deployment of Transformer architectures is significantly hindered by their extensive computational and memory requirements, necessitating the reliance on advanced efficient distributed training methodologies. Prior research has delved into the performance bottlenecks associated with distributed training, aiming to unravel these bottlenecks and suggest optimization directions. However, such analyses often overlook three aspects unique to Transformer models: the specialized architecture, the dependency on various distributed strategies, and the requirement to balance computational and memory overhead. This paper aims to bridge this gap by offering a comprehensive examination of the performance bottlenecks inherent in distributed training of Transformer models, leveraging both theoretical analysis and empirical investigation. We propose an analytical framework tailored to these unique aspects of Transformers, facilitating a holistic evaluation of model architectures, distributed strategies, and resource consumption. Based on this analytical framework, we conduct a comparative analysis of theoretical performances and further systematically explore how various distributed training strategies fare in real-world scenarios. Most of the experimental results can be well explained by the analytical outcomes derived from the analytical framework. Notably, our findings suggest an advantage of pipeline parallelism over data parallelism for Transformer models. Moreover, we shed light on some unexpected outcomes, such as the potential for increased total memory overhead due to suboptimal model partitioning within pipeline parallelism. Additionally, we underscore the significance of communication block size and waiting time to further enhance performance.

cs.DC

Holmes: Towards Distributed Training Across Clusters with Heterogeneous NIC Environment

Large language models (LLMs) such as GPT-3, OPT, and LLaMA have demonstrated remarkable accuracy in a wide range of tasks. However, training these models can incur significant expenses, often requiring tens of thousands of GPUs for months of continuous operation. Typically, this training is carried out in specialized GPU clusters equipped with homogeneous high-speed Remote Direct Memory Access (RDMA) network interface cards (NICs). The acquisition and maintenance of such dedicated clusters is challenging. Current LLM training frameworks, like Megatron-LM and Megatron-DeepSpeed, focus primarily on optimizing training within homogeneous cluster settings. In this paper, we introduce Holmes, a training framework for LLMs that employs thoughtfully crafted data and model parallelism strategies over the heterogeneous NIC environment. Our primary technical contribution lies in a novel scheduling method that intelligently allocates distinct computational tasklets in LLM training to specific groups of GPU devices based on the characteristics of their connected NICs. Furthermore, our proposed framework, utilizing pipeline parallel techniques, demonstrates scalability to multiple GPU clusters, even in scenarios without high-speed interconnects between nodes in distinct clusters. We conducted comprehensive experiments that involved various scenarios in the heterogeneous NIC environment. In most cases, our framework achieves performance levels close to those achievable with homogeneous RDMA-capable networks (InfiniBand or RoCE), significantly exceeding training efficiency within the pure Ethernet environment. Additionally, we verified that our framework outperforms other mainstream LLM frameworks under heterogeneous NIC environment in terms of training efficiency and can be seamlessly integrated with them.

cs.CL

FlattenQuant: Breaking Through the Inference Compute-bound for Large Language Models with Per-tensor Quantization

Large language models (LLMs) have demonstrated state-of-the-art performance across various tasks. However, the latency of inference and the large GPU memory consumption of LLMs restrict their deployment performance. Recently, there have been some efficient attempts to quantize LLMs, yet inference with large batch size or long sequence still has the issue of being compute-bound. Fine-grained quantization methods have showcased their proficiency in achieving low-bit quantization for LLMs, while requiring FP16 data type for linear layer computations, which is time-consuming when dealing with large batch size or long sequence. In this paper, we introduce a method called FlattenQuant, which significantly reduces the maximum value of the tensor by flattening the large channels in the tensor, to achieve low bit per-tensor quantization with minimal accuracy loss. Our experiments show that FlattenQuant can directly use 4 bits to achieve 48.29% of the linear layer calculation in LLMs, with the remaining layers using 8 bits. The 4-bit matrix multiplication introduced in the FlattenQuant method can effectively address the compute-bound caused by large matrix calculation. Our work achieves up to 2$\times$ speedup and 2.3$\times$ memory reduction for LLMs with negligible loss in accuracy.

cs.LG

Photonic Floquet skin-topological effect

Non-Hermitian skin effect and photonic topological edge states are of great interest in non-Hermitian physics and optics. However, the interplay between them is largly unexplored. Here, we propose and demonstrate experimentally the non-Hermitian skin effect that constructed from the nonreciprocal flow of Floquet topological edge states, which can be dubbed 'Floquet skin-topological effect'. We first show the non-Hermitian skin effect can be induced by pure loss when the one-dimensional (1D) system is periodically driven. Next, based on a two-dimensional (2D) Floquet topological photonic lattice with structured loss, we investigate the interaction between the non-Hermiticity and the topological edge states. We observe that all the one-way edge states are imposed onto specific corners, featuring both the non-Hermitian skin effect and topological edge states. Furthermore, a topological switch for the skin-topological effect is presented by utilizing the gap-closing mechanism. Our experiment paves the way of realizing non-Hermitian topological effects in nonlinear and quantum regimes.

cond-mat.mes-hall