Searcharxiv⌕ Search

arXiv subjects

Ning Sun

Publications and source records attributed to Ning Sun.

At least 37 records · Page 2Linked to original sources

Optimization-free Smooth Control Barrier Function for Polygonal Collision Avoidance

Polygonal collision avoidance (PCA) is short for the problem of collision avoidance between two polygons (i.e., polytopes in planar) that own their dynamic equations. This problem suffers the inherent difficulty in dealing with non-smooth boundaries and recently optimization-defined metrics, such as signed distance field (SDF) and its variants, have been proposed as control barrier functions (CBFs) to tackle PCA problems. In contrast, we propose an optimization-free smooth CBF method in this paper, which is computationally efficient and proved to be nonconservative. It is achieved by three main steps: a lower bound of SDF is expressed as a nested Boolean logic composition first, then its smooth approximation is established by applying the latest log-sum-exp method, after which a specified CBF-based safety filter is proposed to address this class of problems. To illustrate its wide applications, the optimization-free smooth CBF method is extended to solve distributed collision avoidance of two underactuated nonholonomic vehicles and drive an underactuated container crane to avoid a moving obstacle respectively, for which numerical simulations are also performed.

math.OC↗

Efimov Effect in Long-range Quantum Spin Chains

When two non-relativistic particles interact resonantly in three dimensions, an infinite tower of three-body bound states emerges, exhibiting a discrete scale invariance. This universal phenomenon, known as the Efimov effect, has garnered extensive attention across various fields, including atomic, nuclear, condensed matter, and particle physics. In this letter, we demonstrate that the Efimov effect also manifests in long-range quantum spin chains. The long-range coupling modifies the low-energy dispersion of magnons, enabling the emergence of continuous scale invariance for two-magnon states at resonance. This invariance is subsequently broken to discrete scale invariance upon imposing short-range boundary conditions for the three-magnon problem, leading to the celebrated Efimov bound states. Using effective field theory, we theoretically determine how the ratio of two successive binding energies depends on the interaction range, which agrees with the numerical solution of the bound-state problem. We further discuss generalizations to arbitrary spatial dimensions, where the traditional Efimov effect serves as a special case. Our results reveal universal physics in dilute quantum gases of magnons that can be experimentally tested in trapped-ion systems.

cond-mat.quant-gas↗

Scheme to Detect the Strong-to-weak Symmetry Breaking via Randomized Measurements

Symmetry breaking plays a central role in classifying the phases of quantum many-body systems. Recent developments have highlighted a novel symmetry-breaking pattern, in which the strong symmetry of a density matrix spontaneously breaks to the week symmetry. This strong-to-weak symmetry breaking is typically detected using multi-replica correlation functions, such as the Rényi-2 correlator. In this letter, we propose a practical protocol for detecting strong-to-weak symmetry breaking in experiments using the randomized measurement toolbox. Our scheme involves collecting the results of random Pauli measurements for (i) the original quantum state and (ii) the quantum state after evolution with the charged operators. Based on the measurement results, with a large number of samples, we can obtain the exact solution to the Rényi-2 correlator. With a small sample size, we can still provide an alternative approach to estimate the phase boundary to a decent accuracy. We perform numerical simulations of Ising chains with all-to-all decoherence as an exemplary demonstration. Our result opens the opportunity for the experimental studies of the novel quantum phases in mixed quantum states.

quant-ph↗

Towards Universal Performance Modeling for Machine Learning Training on Multi-GPU Platforms

Characterizing and predicting the training performance of modern machine learning (ML) workloads on compute systems with compute and communication spread between CPUs, GPUs, and network devices is not only the key to optimization and planning but also a complex goal to achieve. The primary challenges include the complexity of synchronization and load balancing between CPUs and GPUs, the variance in input data distribution, and the use of different communication devices and topologies (e.g., NVLink, PCIe, network cards) that connect multiple compute devices, coupled with the desire for flexible training configurations. Built on top of our prior work for single-GPU platforms, we address these challenges and enable multi-GPU performance modeling by incorporating (1) data-distribution-aware performance models for embedding table lookup, and (2) data movement prediction of communication collectives, into our upgraded performance modeling pipeline equipped with inter-and intra-rank synchronization for ML workloads trained on multi-GPU platforms. Beyond accurately predicting the per-iteration training time of DLRM models with random configurations with a geomean error of 5.21% on two multi-GPU platforms, our prediction pipeline generalizes well to other types of ML workloads, such as Transformer-based NLP models with a geomean error of 3.00%. Moreover, even without actually running ML workloads like DLRMs on the hardware, it is capable of generating insights such as quickly selecting the fastest embedding table sharding configuration (with a success rate of 85%).

cs.DC↗

Strong-to-weak Symmetry Breaking and Entanglement Transitions

When interacting with an environment, the entanglement within quantum many-body systems is rapidly transferred to the entanglement between the system and the bath. For systems with a large local Hilbert space dimension, this leads to a first-order entanglement transition for the reduced density matrix of the system. On the other hand, recent studies have introduced a new paradigm for classifying density matrices, with particular focus on scenarios where a strongly symmetric density matrix undergoes spontaneous symmetry breaking to a weak symmetry phase. This is typically characterized by a finite Rényi-2 correlator or a finite Wightman correlator. In this work, we study the entanglement transition from the perspective of strong-to-weak symmetry breaking, using solvable complex Brownian SYK models. We perform analytical calculations for both the early-time and late-time saddles. The results show that while the Rényi-2 correlator indicates a transition from symmetric to symmetry-broken phase, the Wightman correlator becomes finite even in the early-time saddle due to the single-replica limit, demonstrating that the first-order transition occurs between a near-symmetric phase and a deeply symmetry-broken phase in the sense of Wightman correlator. Our results provide a novel viewpoint on the entanglement transition under symmetry constraints and can be readily generalized to systems with repeated measurements.

quant-ph↗

ADFQ-ViT: Activation-Distribution-Friendly Post-Training Quantization for Vision Transformers

Vision Transformers (ViTs) have exhibited exceptional performance across diverse computer vision tasks, while their substantial parameter size incurs significantly increased memory and computational demands, impeding effective inference on resource-constrained devices. Quantization has emerged as a promising solution to mitigate these challenges, yet existing methods still suffer from significant accuracy loss at low-bit. We attribute this issue to the distinctive distributions of post-LayerNorm and post-GELU activations within ViTs, rendering conventional hardware-friendly quantizers ineffective, particularly in low-bit scenarios. To address this issue, we propose a novel framework called Activation-Distribution-Friendly post-training Quantization for Vision Transformers, ADFQ-ViT. Concretely, we introduce the Per-Patch Outlier-aware Quantizer to tackle irregular outliers in post-LayerNorm activations. This quantizer refines the granularity of the uniform quantizer to a per-patch level while retaining a minimal subset of values exceeding a threshold at full-precision. To handle the non-uniform distributions of post-GELU activations between positive and negative regions, we design the Shift-Log2 Quantizer, which shifts all elements to the positive region and then applies log2 quantization. Moreover, we present the Attention-score enhanced Module-wise Optimization which adjusts the parameters of each quantizer by reconstructing errors to further mitigate quantization error. Extensive experiments demonstrate ADFQ-ViT provides significant improvements over various baselines in image classification, object detection, and instance segmentation tasks at 4-bit. Specifically, when quantizing the ViT-B model to 4-bit, we achieve a 10.23% improvement in Top-1 accuracy on the ImageNet dataset.

cs.CV↗

Non-stationary BERT: Exploring Augmented IMU Data For Robust Human Activity Recognition

Human Activity Recognition (HAR) has gained great attention from researchers due to the popularity of mobile devices and the need to observe users' daily activity data for better human-computer interaction. In this work, we collect a human activity recognition dataset called OPPOHAR consisting of phone IMU data. To facilitate the employment of HAR system in mobile phone and to achieve user-specific activity recognition, we propose a novel light-weight network called Non-stationary BERT with a two-stage training method. We also propose a simple yet effective data augmentation method to explore the deeper relationship between the accelerator and gyroscope data from the IMU. The network achieves the state-of-the-art performance testing on various activity recognition datasets and the data augmentation method demonstrates its wide applicability.

cs.AI↗

Holmes: Towards Distributed Training Across Clusters with Heterogeneous NIC Environment

Large language models (LLMs) such as GPT-3, OPT, and LLaMA have demonstrated remarkable accuracy in a wide range of tasks. However, training these models can incur significant expenses, often requiring tens of thousands of GPUs for months of continuous operation. Typically, this training is carried out in specialized GPU clusters equipped with homogeneous high-speed Remote Direct Memory Access (RDMA) network interface cards (NICs). The acquisition and maintenance of such dedicated clusters is challenging. Current LLM training frameworks, like Megatron-LM and Megatron-DeepSpeed, focus primarily on optimizing training within homogeneous cluster settings. In this paper, we introduce Holmes, a training framework for LLMs that employs thoughtfully crafted data and model parallelism strategies over the heterogeneous NIC environment. Our primary technical contribution lies in a novel scheduling method that intelligently allocates distinct computational tasklets in LLM training to specific groups of GPU devices based on the characteristics of their connected NICs. Furthermore, our proposed framework, utilizing pipeline parallel techniques, demonstrates scalability to multiple GPU clusters, even in scenarios without high-speed interconnects between nodes in distinct clusters. We conducted comprehensive experiments that involved various scenarios in the heterogeneous NIC environment. In most cases, our framework achieves performance levels close to those achievable with homogeneous RDMA-capable networks (InfiniBand or RoCE), significantly exceeding training efficiency within the pure Ethernet environment. Additionally, we verified that our framework outperforms other mainstream LLM frameworks under heterogeneous NIC environment in terms of training efficiency and can be seamlessly integrated with them.

cs.CL↗

NLP-based detection of systematic anomalies among the narratives of consumer complaints

We develop an NLP-based procedure for detecting systematic nonmeritorious consumer complaints, simply called systematic anomalies, among complaint narratives. While classification algorithms are used to detect pronounced anomalies, in the case of smaller and frequent systematic anomalies, the algorithms may falter due to a variety of reasons, including technical ones as well as natural limitations of human analysts. Therefore, as the next step after classification, we convert the complaint narratives into quantitative data, which are then analyzed using an algorithm for detecting systematic anomalies. We illustrate the entire procedure using complaint narratives from the Consumer Complaint Database of the Consumer Financial Protection Bureau.

stat.ME↗

Flux ratios for effects of permanent charges on ionic flows with three ion species: Case study (II)

In this paper, we study effects of permanent charges on ion flows through membrane channels via a quasi-one-dimensional classical Poisson-Nernst-Planck system. This system includes three ion species, two cations with different valences and one anion, and permanent charges with a simple structure, zeros at the two end regions and a constant over the middle region. For small permanent charges, our main goal is to analyze the effects of permanent charges on ionic flows, interacting with the boundary conditions and channel structure. Continuing from a previous work, we investigate the problem for a new case toward a more comprehensive understanding about effects of permanent charges on ionic fluxes.

cond-mat.soft↗

Exploring Post-Training Quantization of Protein Language Models

Recent advancements in unsupervised protein language models (ProteinLMs), like ESM-1b and ESM-2, have shown promise in different protein prediction tasks. However, these models face challenges due to their high computational demands, significant memory needs, and latency, restricting their usage on devices with limited resources. To tackle this, we explore post-training quantization (PTQ) for ProteinLMs, focusing on ESMFold, a simplified version of AlphaFold based on ESM-2 ProteinLM. Our study is the first attempt to quantize all weights and activations of ProteinLMs. We observed that the typical uniform quantization method performs poorly on ESMFold, causing a significant drop in TM-Score when using 8-bit quantization. We conducted extensive quantization experiments, uncovering unique challenges associated with ESMFold, particularly highly asymmetric activation ranges before Layer Normalization, making representation difficult using low-bit fixed-point formats. To address these challenges, we propose a new PTQ method for ProteinLMs, utilizing piecewise linear quantization for asymmetric activation values to ensure accurate approximation. We demonstrated the effectiveness of our method in protein structure prediction tasks, demonstrating that ESMFold can be accurately quantized to low-bit widths without compromising accuracy. Additionally, we applied our method to the contact prediction task, showcasing its versatility. In summary, our study introduces an innovative PTQ method for ProteinLMs, addressing specific quantization challenges and potentially leading to the development of more efficient ProteinLMs with significant implications for various protein-related applications.

cs.LG↗

V2X-Seq: A Large-Scale Sequential Dataset for Vehicle-Infrastructure Cooperative Perception and Forecasting

Utilizing infrastructure and vehicle-side information to track and forecast the behaviors of surrounding traffic participants can significantly improve decision-making and safety in autonomous driving. However, the lack of real-world sequential datasets limits research in this area. To address this issue, we introduce V2X-Seq, the first large-scale sequential V2X dataset, which includes data frames, trajectories, vector maps, and traffic lights captured from natural scenery. V2X-Seq comprises two parts: the sequential perception dataset, which includes more than 15,000 frames captured from 95 scenarios, and the trajectory forecasting dataset, which contains about 80,000 infrastructure-view scenarios, 80,000 vehicle-view scenarios, and 50,000 cooperative-view scenarios captured from 28 intersections' areas, covering 672 hours of data. Based on V2X-Seq, we introduce three new tasks for vehicle-infrastructure cooperative (VIC) autonomous driving: VIC3D Tracking, Online-VIC Forecasting, and Offline-VIC Forecasting. We also provide benchmarks for the introduced tasks. Find data, code, and more up-to-date information at \href{https://github.com/AIR-THU/DAIR-V2X-Seq}{https://github.com/AIR-THU/DAIR-V2X-Seq}.

cs.CV↗

Tail maximal dependence in bivariate models: estimation and applications

Assessing dependence within co-movements of financial instruments has been of much interest in risk management. Typically, indices of tail dependence are used to quantify the strength of such dependence, although many of the indices underestimate the strength. Hence, we advocate the use of a statistical procedure designed to estimate the maximal strength of dependence that can possibly occur among the co-movements. We illustrate the procedure using simulated and real data-sets.

stat.ME↗

Detecting systematic anomalies affecting systems when inputs are stationary time series

We develop an anomaly-detection method when systematic anomalies, possibly statistically very similar to genuine inputs, are affecting control systems at the input and/or output stages. The method allows anomaly-free inputs (i.e., those before contamination) to originate from a wide class of random sequences, thus opening up possibilities for diverse applications. To illustrate how the method works on data, and how to interpret its results and make decisions, we analyze several actual time series, which are originally non-stationary but in the process of analysis are converted into stationary. As a further illustration, we provide a controlled experiment with anomaly-free inputs following an ARMA time series model under various contamination scenarios.

stat.ME↗

On the Equivalence between Spin and Charge Dynamics of the Fermi Hubbard Model

Utilizing the Fermi gas microscope, recently the MIT group has measured the spin transport of the Fermi Hubbard model starting from a spin-density-wave state, and the Princeton group has measured the charge transport of the Fermi Hubbard model starting from a charge-density-wave state. Motivated by these two experiments, we prove a theorem that shows under certain conditions, the spin and charge transports can be equivalent to each other. The proof makes use of the particle-hole transformation of the Fermi Hubbard model and a recently discovered symmetry protected dynamical symmetry. Our results can be directly verified in future cold atom experiment with the Fermi gas microscope.

cond-mat.quant-gas↗

Resonant Driving induced Ferromagnetism in the Fermi Hubbard Model

In this letter we consider quantum phases and the phase diagram of a Fermi Hubbard model under periodic driving that has been realized in recent cold atom experiments, in particular, when the driving frequency is resonant with the interaction energy. Due to the resonant driving, the effective Hamiltonian contains a correlated hopping term where the density occupation strongly modifies the hopping strength. Focusing on half filling, in addition to the charge and spin density wave phases, large regions of ferromagnetic phase and phase separation are discovered in the weakly interacting regime. The mechanism of this ferromagnetism is attributed to the correlated hopping because the hopping strength within a ferromagnetic domain is normalized to a larger value than the hopping strength across the domain. Thus, the kinetic energy favors a large ferromagnetic domain and consequently drives the system into a ferromagnetic phase. We note that this is a different mechanism in contrast to the well-known Stoner mechanism for ferromagnetism where the ferromagnetism is driven by interaction energy.

cond-mat.quant-gas↗

Deep Learning Topological Invariants of Band Insulators

In this work we design and train deep neural networks to predict topological invariants for one-dimensional four-band insulators in AIII class whose topological invariant is the winding number, and two-dimensional two-band insulators in A class whose topological invariant is the Chern number. Given Hamiltonians in the momentum space as the input, neural networks can predict topological invariants for both classes with accuracy close to or higher than 90%, even for Hamiltonians whose invariants are beyond the training data set. Despite the complexity of the neural network, we find that the output of certain intermediate hidden layers resembles either the winding angle for models in AIII class or the solid angle (Berry curvature) for models in A class, indicating that neural networks essentially capture the mathematical formula of topological invariants. Our work demonstrates the ability of neural networks to predict topological invariants for complicated models with local Hamiltonians as the only input, and offers an example that even a deep neural network is understandable.

cond-mat.str-el↗

Universal relations for spin-orbit coupled Fermi gas near an s-wave resonance

The synthetic spin-orbit coupled quantum gases is widely studied both experimentally and theoretically in recent years. As previous studies show, this modification of single-body dispersion will in general couple different partial waves and thus distort the wave-function of bound states which determines the short-distance behavior of many-body wave function. In this work, we focus on the two-component Fermi gas with one-dimensional or three-dimensional spin-orbit coupling near an s-wave resonance. Using the method of effective field theory and the operator product expansion, we derive universal relations for both systems, and obtain the momentum distribution matrix $\left<ψ^\dagger_a(\mathbf{q})ψ_b(\mathbf{q})\right>$ at large $\mathbf{q}$ ($a,b$ are spin index), which shows anisotropic features. We also discuss the experimental implication of these results depending on the realization of the spin-orbit coupling.

cond-mat.quant-gas↗