SearcharxivSearch

arXiv subjects

Gang Lu

Publications and source records attributed to Gang Lu.

At least 19 recordsLinked to original sources

Exact Entanglement-Depth Speed Frontier for Complete Quantum Charging

Complete quantum charging provides a sharp setting in which to ask how much multipartite entanglement is forced by speed itself. For a closed \(N\)-qubit battery evolving from \(\ket{\downarrow}^{\otimes N}\) to \(\ket{\uparrow}^{\otimes N}\) under a time-independent Hamiltonian, we exactly solve the pure-state depth-constrained speed problem. If the realized trajectory has entanglement depth at most \(k\), then the largest possible QSL-normalized rate \(\eta=\tau_{\rm QSL}/T\) is \(\eta_{\max}(k)=\lceil N/k\rceil^{-1/2}\). Conversely, an observed rate \(\eta\) certifies trajectory entanglement depth at least \(\bigl\lceil N/\lfloor \eta^{-2}\rfloor\bigr\rceil\). The mechanism is block orthogonalization: under a fixed product partition, complete charging forces all blocks to orthogonalize simultaneously, and the quantum speed limit converts this counting constraint into the speed bound. Balanced cluster-flip evolutions saturate the bound, establishing an exact integer staircase frontier. Thus fast complete charging cannot be explained by many small independently charging blocks; in particular, crossing the threshold \(\eta>1/\sqrt2\) certifies, for \(N>1\), the generation of genuine \(N\)-partite entanglement.

quant-ph

Magnetic and moir\'e Proximity Effects in WSe2/WSe2/CrI3 Trilayers

Integrating magnetic order to moir\'e superlattices is of significant scientific and technological interest. Based on first-principles calculations, we study the interplay of magnetic proximity and moir\'e proximity in WSe2/WSe2/CrI3 trilayers with different stackings and twist angles. Large valley splitting is observed due to redistribution of the exciton charge density across layers via a super-exchange-like mechanism, and its electric-field dependence bears similarity to electrically tunable and valley-selective Feshbach resonances. The valley splitting can be magnified in moir\'e superlattices owing to the superposition of Umklapp excitons folded from moir\'e minibands, yielding spatially modulated and enhanced magnetic proximity. The moir\'e proximity effect is demonstrated via an imprinted moir\'e potential on CrI3 layer and its feedback to the direct moir\'e potential on WSe2 bilayers is observed. The cooperation between the direct and imprinted moir\'e potentials is shown to yield novel topological and correlated states.

cond-mat.mtrl-sci

Adaptive Ensemble Learning with Gaussian Copula for Load Forecasting

Machine learning (ML) is capable of accurate Load Forecasting from complete data. However, there are many uncertainties that affect data collection, leading to sparsity. This article proposed a model called Adaptive Ensemble Learning with Gaussian Copula to deal with sparsity, which contains three modules: data complementation, ML construction, and adaptive ensemble. First, it applies Gaussian Copula to eliminate sparsity. Then, we utilise five ML models to make predictions individually. Finally, it employs adaptive ensemble to get final weighted-sum result. Experiments have demonstrated that our model are robust.

cs.LG

Robust Load Prediction of Power Network Clusters Based on Cloud-Model-Improved Transformer

Load data from power network clusters indicates economic development in each area, crucial for predicting regional trends and guiding power enterprise decisions. The Transformer model, a leading method for load prediction, faces challenges modeling historical data due to variables like weather, events, festivals, and data volatility. To tackle this, the cloud model's fuzzy feature is utilized to manage uncertainties effectively. Presenting an innovative approach, the Cloud Model Improved Transformer (CMIT) method integrates the Transformer model with the cloud model utilizing the particle swarm optimization algorithm, with the aim of achieving robust and precise power load predictions. Through comparative experiments conducted on 31 real datasets within a power network cluster, it is demonstrated that CMIT significantly surpasses the Transformer model in terms of prediction accuracy, thereby highlighting its effectiveness in enhancing forecasting capabilities within the power network cluster sector.

cs.LG

Pay Attention to Weak Ties: A Heterogeneous Multiplex Representation Learning Framework for Link Prediction

Graph neural networks (GNNs) can learn effective node representations that significantly improve link prediction accuracy. However, most GNN-based link prediction algorithms are incompetent to predict weak ties connecting different communities. Most link prediction algorithms are designed for networks with only one type of relation between nodes but neglect the fact that many complex systems, including transportation and social networks, consisting of multi-modalities of interactions that correspond to different nature of interactions and dynamics that can be modeled as multiplex network, where different types of relation are represented in different layers. This paper proposes a Multi-Relations-aware Graph Neural Network (MRGNN) framework to learn effective node representations for multiplex networks and make more accurate link predictions, especially for weak ties. Specifically, our model utilizes an intra-layer node-level feature propagation process and an inter-layer representation merge process, which applies a simple yet effective logistic or semantic attention voting mechanism to adaptively aggregate information from different layers. Extensive experiments on four diversified multiplex networks show that MRGNN outperforms the state-of-the-art multiplex link prediction algorithms on overall prediction accuracy, and works pretty well on forecasting weak ties

cs.SI

Enhancing Large-Scale AI Training Efficiency: The C4 Solution for Real-Time Anomaly Detection and Communication Optimization

The emergence of Large Language Models (LLMs) has necessitated the adoption of distributed training techniques, involving the deployment of thousands of GPUs to train a single model. Unfortunately, the efficiency of large-scale distributed training systems is often suboptimal due to the increased likelihood of hardware errors in high-end GPU products and the heightened risk of network traffic collisions. Moreover, any local hardware failure can disrupt training tasks, and the inability to swiftly identify faulty components leads to a significant waste of GPU resources. And, prolonged communication due to traffic collisions can substantially increase GPU waiting times. To address these challenges, we propose a communication-driven solution, namely the C4. The key insights of C4 are twofold. First, the load in distributed training exhibits homogeneous characteristics and is divided into iterations through periodic synchronization, therefore hardware anomalies would incur certain syndrome in collective communication. By leveraging this feature, C4 can rapidly identify the faulty components, swiftly isolate the anomaly, and restart the task, thereby avoiding resource wastage caused by delays in anomaly detection. Second, the predictable communication model of collective communication, involving a limited number of long-lived flows, allows C4 to efficiently execute traffic planning, substantially reducing bandwidth competition among these flows. The C4 has been extensively deployed across real-world production systems in a hyperscale cloud provider, yielding a significant improvement in system efficiency, from 30% to 45%. This enhancement is attributed to a 30% reduction in error-induced overhead and a 15% reduction in communication costs.

cs.DC

Square Moir\'e Superlattices in Twisted Two-Dimensional Halide Perovskites

Moir\'e superlattices have emerged as a new platform for studying strongly correlated quantum phenomena, but these systems have been largely limited to van der Waals layer two-dimensional (2D) materials. Here we introduce moir\'e superlattices leveraging ultra-thin, ligand-free halide perovskites, facilitated by ionic interactions. Square moir\'e superlattices with varying periodic lengths are clearly visualized through high-resolution transmission electron microscopy. Twist-angle-dependent transient photoluminescence microscopy and electrical characterizations indicate the emergence of localized bright excitons and trapped charge carriers near a twist angle of ~10{\deg}. The localized excitons are accompanied by enhanced exciton emission, attributed to an increased oscillator strength by a theoretically forecasted flat band. This work illustrates the potential of extended ionic interaction in realizing moir\'e physics at room temperature, broadening the horizon for future investigations.

cond-mat.mtrl-sci

Emergence of scaling in dockless bike-sharing systems

Fundamental laws of human mobility have been extensively studied, yet we are still lacking a comprehensive understanding of the mobility patterns of sharing conveyances. Since travellers would highly probably no longer possess their own conveyances in the near future, the interplay between travellers and sharing bikes is a central question for developing more sustainable transportation. Dockless bike-sharing systems that record detailed information of every trip provide us a unique opportunity for revealing the hidden patterns behind riding activities. By treating each bike as an individual entity, we reveal that distributions of mobility indicators of bikes are quite different from humans; and mobility patterns are even inconsistent across cities. All above discrepancies can be well explained by a choice model that is characterized by a universal scaling. Our model unveils that instead of choosing among the newest bikes, the distribution of rank values of selected bikes on usage condition manifests a truncated power-law and is quite stable across several cities despite various diversities. Our framework would have broad implications in sharing economy and contribute towards developing a greener, healthier, and more sustainable future city.

physics.soc-ph

Shedding Light on Moire Excitons: A First-Principles Perspective

Moire superlattices in van der Waals (vdW) heterostructures could trap strongly bonded and long lived interlayer excitons. Assumed to be localized, these moire excitons could form ordered quantum dot arrays, paving the way for novel optoelectronic and quantum information applications. Here we perform first principles simulations to shed light on moire excitons in twisted MoS2/WS2 heterostructures. We provide the direct evidence of localized interlayer moire excitons in vdW heterostructures. The moire potentials are mapped out based on spatial modulations of energy gaps. Nearly flat valence bands are observed in the heterostructures without magic angles. The dependence of spatial localization and binding energy of the moire excitons on the twist angle of the heterostructures is examined. We explore how electric field can be tuned to control the position, polarity, emission energy, and hybridization strength of the moire excitons. We predict that alternating electric fields could modulate the dipole moments of hybridized moire excitons and suppress their diffusion in Moire lattices.

physics.comp-ph

AIBench: An Agile Domain-specific Benchmarking Methodology and an AI Benchmark Suite

Domain-specific software and hardware co-design is encouraging as it is much easier to achieve efficiency for fewer tasks. Agile domain-specific benchmarking speeds up the process as it provides not only relevant design inputs but also relevant metrics, and tools. Unfortunately, modern workloads like Big data, AI, and Internet services dwarf the traditional one in terms of code size, deployment scale, and execution path, and hence raise serious benchmarking challenges. This paper proposes an agile domain-specific benchmarking methodology. Together with seventeen industry partners, we identify ten important end-to-end application scenarios, among which sixteen representative AI tasks are distilled as the AI component benchmarks. We propose the permutations of essential AI and non-AI component benchmarks as end-to-end benchmarks. An end-to-end benchmark is a distillation of the essential attributes of an industry-scale application. We design and implement a highly extensible, configurable, and flexible benchmark framework, on the basis of which, we propose the guideline for building end-to-end benchmarks, and present the first end-to-end Internet service AI benchmark. The preliminary evaluation shows the value of our benchmark suite---AIBench against MLPerf and TailBench for hardware and software designers, micro-architectural researchers, and code developers. The specifications, source code, testbed, and results are publicly available from the web site \url{http://www.benchcouncil.org/AIBench/index.html}.

cs.PF

Isolate First, Then Share: a New OS Architecture for Datacenter Computing

This paper presents the "isolate first, then share" OS model in which the processor cores, memory, and devices are divided up between disparate OS instances and a new abstraction, subOS, is proposed to encapsulate an OS instance that can be created, destroyed, and resized on-the-fly. The intuition is that this avoids shared kernel states between applications, which in turn reduces performance loss caused by contention. We decompose the OS into the supervisor and several subOSes running at the same privilege level: a subOS directly manages physical resources, while the supervisor can create, destroy, resize a subOS on-the-fly. The supervisor and subOSes have few state sharing, but fast inter-subOS communication mechanisms are provided on demand. We present the first implementation, RainForest, which supports unmodified Linux binaries. Our comprehensive evaluation shows RainForest outperforms Linux with four different kernels, LXC, and Xen in terms of worst-case and average performance most of time when running a large number of benchmarks. The source code is available soon.

cs.OS

Charge Transport in Hybrid Halide Perovskites

Charge transport is crucial to the performance of hybrid halide perovskite solar cells. A theoretical model based on large polarons is proposed to elucidate charge transport properties in the perovskites. Critical new physical insights are incorporated into the model, including the recognitions that acoustic phonons as opposed to optical phonons are responsible for the scattering of the polarons; these acoustic phonons are fully excited due to the "softness" of the perovskites, and the temperature-dependent dielectric function underlies the temperature dependence of charge carrier mobility. This work resolves key controversies in literature and forms a starting point for more rigorous first-principles predictions of charge transport.

cond-mat.mtrl-sci

10-millisecond Computing

Despite computation becomes much complex on data with an unprecedented scale, we argue computers or smart devices should and will consistently provide information and knowledge to human being in the order of a few tens milliseconds. We coin a new term 10-millisecond computing to call attention to this class of workloads. 10-millisecond computing raises many challenges for both software and hardware stacks. In this paper, using a typical workload-memcached on a 40-core server (a main-stream server in near future), we quantitatively measure 10-ms computing's challenges to conventional operating systems. For better communication, we propose a simple metric-outlier proportion to measure quality of service: for N completed requests or jobs, if M jobs or requests' latencies exceed the outlier threshold t, the outlier proportion is M/N . For a 1K-scale system running Linux (version 2.6.32), LXC (version 0.7.5) or XEN (version 4.0.0), respectively, we surprisingly find that so as to reduce the service outlier proportion to 10% (10% users will feel QoS degradation), the outlier proportion of a single server has to be reduced by 871X, 2372X, 2372X accordingly. Also, we discuss the possible design spaces of 10-ms computing systems from perspectives of datacenter architectures, networking, OS and scheduling, and benchmarking.

cs.PF

Comment on "Linear Scaling of the Exciton Binding Energy versus the Band Gap of Two-Dimensional Materials"

In a recent letter [J. -H. Choi et al. Phys. Rev. Lett. 115, 066403 (2015)], a universal linear relation between the binding energy Eb of exciton and the band gap Eg is found in different quasi-2D semiconductors. However, when one extrapolates the straight line to Eg=0, a positive Eb is obtained. This means that a stable exciton exists in a metal (Eg=0), or in other words, the excitation energy is negative. Therefore, for quasi-2D semiconductors with small band gap, Eb vs. Eg must deviate from the straight line suggested in the Letter. We performed similar GW+BSE calculation as the Letter for compressed phosphorene and stretched graphene. At larger band gap, our data fall in the same straight lien as the Letter. But for the quasi-2D semiconductors with small band gap, the Eb vs. Eg curve approaches to zero. A simple model based on local instantaneous screened interaction gives reasonable agreement with the more rigorous calculations. For the quasi-2D semiconductors with diminishing gap Eg<0.3eV, due to the nonlocal retarded effect induced by the strong screening in the small gap materials, one has to use 2-time transition amplitude (Bethe-Salpeter wave function) to describe exciton.

cond-mat.mtrl-sci

BigDataBench-MT: A Benchmark Tool for Generating Realistic Mixed Data Center Workloads

Long-running service workloads (e.g. web search engine) and short-term data analysis workloads (e.g. Hadoop MapReduce jobs) co-locate in today's data centers. Developing realistic benchmarks to reflect such practical scenario of mixed workload is a key problem to produce trustworthy results when evaluating and comparing data center systems. This requires using actual workloads as well as guaranteeing their submissions to follow patterns hidden in real-world traces. However, existing benchmarks either generate actual workloads based on probability models, or replay real-world workload traces using basic I/O operations. To fill this gap, we propose a benchmark tool that is a first step towards generating a mix of actual service and data analysis workloads on the basis of real workload traces. Our tool includes a combiner that enables the replaying of actual workloads according to the workload traces, and a multi-tenant generator that flexibly scales the workloads up and down according to users' requirements. Based on this, our demo illustrates the workload customization and generation process using a visual interface. The proposed tool, called BigDataBench-MT, is a multi-tenant version of our comprehensive benchmark suite BigDataBench and it is publicly available from http://prof.ict.ac.cn/BigDataBench/multi-tenancyversion/.

cs.DC