Searcharxiv⌕ Search

arXiv subjects

Long Cheng

Publications and source records attributed to Long Cheng.

At least 91 records · Page 5Linked to original sources

Multi-level Distillation of Semantic Knowledge for Pre-training Multilingual Language Model

Pre-trained multilingual language models play an important role in cross-lingual natural language understanding tasks. However, existing methods did not focus on learning the semantic structure of representation, and thus could not optimize their performance. In this paper, we propose Multi-level Multilingual Knowledge Distillation (MMKD), a novel method for improving multilingual language models. Specifically, we employ a teacher-student framework to adopt rich semantic representation knowledge in English BERT. We propose token-, word-, sentence-, and structure-level alignment objectives to encourage multiple levels of consistency between source-target pairs and correlation similarity between teacher and student models. We conduct experiments on cross-lingual evaluation benchmarks including XNLI, PAWS-X, and XQuAD. Experimental results show that MMKD outperforms other baseline models of similar size on XNLI and XQuAD and obtains comparable performance on PAWS-X. Especially, MMKD obtains significant performance gains on low-resource languages.

cs.CL↗

Statistical Modeling of Soft Error Influence on Neural Networks

Soft errors in large VLSI circuits pose dramatic influence on computing- and memory-intensive neural network (NN) processing. Understanding the influence of soft errors on NNs is critical to protect against soft errors for reliable NN processing. Prior work mainly rely on fault simulation to analyze the influence of soft errors on NN processing. They are accurate but usually specific to limited configurations of errors and NN models due to the prohibitively slow simulation speed especially for large NN models and datasets. With the observation that the influence of soft errors propagates across a large number of neurons and accumulates as well, we propose to characterize the soft error induced data disturbance on each neuron with normal distribution model according to central limit theorem and develop a series of statistical models to analyze the behavior of NN models under soft errors in general. The statistical models reveal not only the correlation between soft errors and NN model accuracy, but also how NN parameters such as quantization and architecture affect the reliability of NNs. The proposed models are compared with fault simulation and verified comprehensively. In addition, we observe that the statistical models that characterize the soft error influence can also be utilized to predict fault simulation results in many cases and we explore the use of the proposed statistical models to accelerate fault simulations of NNs. According to our experiments, the accelerated fault simulation shows almost two orders of magnitude speedup with negligible simulation accuracy loss over the baseline fault simulations.

cs.LG↗

Iterative Qubits Management for Quantum Index Searching in a Hybrid System

Recent advances in quantum computing systems attract tremendous attention. Commercial companies, such as IBM, Amazon, and IonQ, have started to provide access to noisy intermediate-scale quantum computers. Researchers and entrepreneurs attempt to deploy their applications that aim to achieve a quantum speedup. Grover's algorithm and quantum phase estimation are the foundations of many applications with the potential for such a speedup. While these algorithms, in theory, obtain marvelous performance, deploying them on existing quantum devices is a challenging task. For example, quantum phase estimation requires extra qubits and a large number of controlled operations, which are impractical due to low-qubit and noisy hardware. To fully utilize the limited onboard qubits, we propose IQuCS, which aims at index searching and counting in a quantum-classical hybrid system. IQuCS is based on Grover's algorithm. From the problem size perspective, it analyzes results and tries to filter out unlikely data points iteratively. A reduced data set is fed to the quantum computer in the next iteration. With a reduction in the problem size, IQuCS requires fewer qubits iteratively, which provides the potential for a shared computing environment. We implement IQuCS with Qiskit and conduct intensive experiments. The results demonstrate that it reduces qubits consumption by up to 66.2%.

quant-ph↗

Differentiate Quality of Experience Scheduling for Deep Learning Inferences with Docker Containers in the Cloud

With the prevalence of big-data-driven applications, such as face recognition on smartphones and tailored recommendations from Google Ads, we are on the road to a lifestyle with significantly more intelligence than ever before. Various neural network powered models are running at the back end of their intelligence to enable quick responses to users. Supporting those models requires lots of cloud-based computational resources, e.g., CPUs and GPUs. The cloud providers charge their clients by the amount of resources that they occupy. Clients have to balance the budget and quality of experiences (e.g., response time). The budget leans on individual business owners, and the required Quality of Experience (QoE) depends on usage scenarios of different applications. For instance, an autonomous vehicle requires an real-time response, but unlocking your smartphone can tolerate delays. However, cloud providers fail to offer a QoE-based option to their clients. In this paper, we propose DQoES, differentiated quality of experience scheduler for deep learning inferences. DQoES accepts clients' specifications on targeted QoEs, and dynamically adjusts resources to approach their targets. Through the extensive cloud-based experiments, DQoES demonstrates that it can schedule multiple concurrent jobs with respect to various QoEs and achieve up to 8x times more satisfied models when compared to the existing system

cs.DC↗

The Sustainable Response Strategy to COVID-19: Pandemic Urban Zoning Based on Multimodal Transport Data

Since the outbreak of COVID-19, it has rapidly evolved into a sudden and major public health emergency globally. With the variants of COVID-19, the difficulty of pandemic control continues to increase, which has brought significant costs to the society. The existing pandemic control zoning method ignores the impact on residents'lives. In this study, we propose a refined and low-cost pandemic control method by scientifically delineating zoning areas. First, a spatial interaction network is built up based on the multimodal transport travel data in Nanjing, China, and an improved Leiden community detection method based on the gravity model is used to obtain a preliminary zoning scheme. Then, we use spatial constraints to correct the results with the discrete spatial distribution. Finally, reasonable zones for pandemic control are obtained. The modularity of the algorithm results is 0.4185, proving that the proposed method is suitable for pandemic control zoning. The proposed method is also demonstrated to be able to minimize traffic flows between pandemic control areas and only 24.8% of travel connections are cut off, thus reducing the impact of pandemic control on residents'daily life and reducing the cost of pandemic control. The findings can help to inform sustainable strategies and suggestions for the pandemic control.

stat.AP↗

Low-frequency Phonon at Perovskite Oxide Interface Studied by Surface-specific Nonlinear Terahertz Spectroscopy

The low-frequency collective excitations, which often occur in the terahertz or multi-terahertz spectral region, play an essential role in many novel emergent phenomena. Despite numerous studies in the bulk, detection of such excitations at interfaces remains challenging owing to the lack of feasible experimental techniques. Here, we show that interfacial low-frequency modes can be characterized using surface-specific nonlinear terahertz spectroscopy. This technique uses intra-pulse difference frequency mixing (DFM) process that can extend the second-order optical spectroscopy to the terahertz range. As a demonstration, the surface phonon of SrTiO3(001) at 2.8 THz was successfully measured. This surface polarization originates from the excess of oxygen vacancies or charge transfer at the interface. We have also developed an analytical procedure for remote measurement of the interfacial potential of complex oxides in a practical environment. Our method offers new opportunities for in situ studies of the low-frequency excitations at interfaces in broad disciplines.

cond-mat.mtrl-sci↗

Revealing the CO2 emission reduction of ridesplitting and its determinants based on real-world data

Ridesplitting, which is a form of pooled ridesourcing service, has great potential to alleviate the negative impacts of ridesourcing on the environment. However, most existing studies only explored its theoretical environmental benefits based on optimization models and simulations. By contrast, this study aims to reveal the real-world emission reduction of ridesplitting and its determinants based on the observed data of ridesourcing in Chengdu, China. Integrating the trip data with the COPERT model, this study calculates the CO2 emissions of shared rides (ridesplitting) and their substituted single rides (regular ridesourcing) to estimate the CO2 emission reduction of each ridesplitting trip. The results show that not all ridesplitting trips reduce emissions from ridesourcing in the real world. The CO2 emission reduction rate of ridesplitting varies from trip to trip, averaging at 43.15g/km. Then, interpretable machine learning models, gradient boosting machines, are applied to explore the relationship between the CO2 emission reduction rate of ridesplitting and its determinants. Based on the SHapley Additive exPlanations (SHAP) method, the overlap rate and detour rate of shared rides are identified to be the most important factors that determine the CO2 emission reduction rate of ridesplitting. Increasing the overlap rate, the number of shared rides, average speed, and ride distance ratio while decreasing the detour rate, actual trip distance, and ride distance gap can increase the CO2 emission reduction rate of ridesplitting. In addition, nonlinear effects and interactions of the determinants are examined through the partial dependence plots. To sum up, this study provides a scientific method for the government and ridesourcing companies to better assess and optimize the environmental benefits of ridesplitting.

cs.LG↗

Understanding and Measuring Robustness of Multimodal Learning

The modern digital world is increasingly becoming multimodal. Although multimodal learning has recently revolutionized the state-of-the-art performance in multimodal tasks, relatively little is known about the robustness of multimodal learning in an adversarial setting. In this paper, we introduce a comprehensive measurement of the adversarial robustness of multimodal learning by focusing on the fusion of input modalities in multimodal models, via a framework called MUROAN (MUltimodal RObustness ANalyzer). We first present a unified view of multimodal models in MUROAN and identify the fusion mechanism of multimodal models as a key vulnerability. We then introduce a new type of multimodal adversarial attacks called decoupling attack in MUROAN that aims to compromise multimodal models by decoupling their fused modalities. We leverage the decoupling attack of MUROAN to measure several state-of-the-art multimodal models and find that the multimodal fusion mechanism in all these models is vulnerable to decoupling attacks. We especially demonstrate that, in the worst case, the decoupling attack of MUROAN achieves an attack success rate of 100% by decoupling just 1.16% of the input space. Finally, we show that traditional adversarial training is insufficient to improve the robustness of multimodal models with respect to decoupling attacks. We hope our findings encourage researchers to pursue improving the robustness of multimodal learning.

cs.LG↗

Asteria: Deep Learning-based AST-Encoding for Cross-platform Binary Code Similarity Detection

Binary code similarity detection is a fundamental technique for many security applications such as vulnerability search, patch analysis, and malware detection. There is an increasing need to detect similar code for vulnerability search across architectures with the increase of critical vulnerabilities in IoT devices. The variety of IoT hardware architectures and software platforms requires to capture semantic equivalence of code fragments in the similarity detection. However, existing approaches are insufficient in capturing the semantic similarity. We notice that the abstract syntax tree (AST) of a function contains rich semantic information. Inspired by successful applications of natural language processing technologies in sentence semantic understanding, we propose a deep learning-based AST-encoding method, named ASTERIA, to measure the semantic equivalence of functions in different platforms. Our method leverages the Tree-LSTM network to learn the semantic representation of a function from its AST. Then the similarity detection can be conducted efficiently and accurately by measuring the similarity between two representation vectors. We have implemented an open-source prototype of ASTERIA. The Tree-LSTM model is trained on a dataset with 1,022,616 function pairs and evaluated on a dataset with 95,078 function pairs. Evaluation results show that our method outperforms the AST-based tool Diaphora and the-state-of-art method Gemini by large margins with respect to the binary similarity detection. And our method is several orders of magnitude faster than Diaphora and Gemini for the similarity calculation. In the application of vulnerability search, our tool successfully identified 75 vulnerable functions in 5,979 IoT firmware images.

cs.CR↗

Quantum percolation and magnetic nano-dropletstates in electronically phase-separated manganite nanowires

One-dimensional (1D) confinement has been revealed to effectively tune the properties of materials in homogeneous states. The 1D physics can be further enriched by electronic inhomogeneity, which unfortunately remains largely unknown. Here we demonstrate the ultra-high sensitivity to magnetic fluctuations and the tunability of phase stability in the electronic transport properties of self-assembled electronically phase-separated manganite nanowires with extreme aspect ratio. The onset of magnetic nano-droplet state, a precursor to the ferromagnetic metallic state, is unambiguously revealed, which is attributed to the small lateral size of the nanowires that is comparable to the droplet size. Moreover, the quasi-1D anisotropy stabilizes thin insulating domains to form intrinsic tunneling junctions in the low temperature range, which is robust even under magnetic field up to 14 T, and thus essentially modifies the classic 1D percolation picture to stabilize a novel quantum percolation state. A new phase diagram is therefore established for the manganite system under quasi-1D confinement for the first time. Our findings offer new insight to understand and manipulate the colorful properties of the electronically phase-separated systems via dimensionality engineering.

cond-mat.str-el↗

R2F: A Remote Retraining Framework for AIoT Processors with Computing Errors

AIoT processors fabricated with newer technology nodes suffer rising soft errors due to the shrinking transistor sizes and lower power supply. Soft errors on the AIoT processors particularly the deep learning accelerators (DLAs) with massive computing may cause substantial computing errors. These computing errors are difficult to be captured by the conventional training on general purposed processors like CPUs and GPUs in a server. Applying the offline trained neural network models to the edge accelerators with errors directly may lead to considerable prediction accuracy loss. To address the problem, we propose a remote retraining framework (R2F) for remote AIoT processors with computing errors. It takes the remote AIoT processor with soft errors in the training loop such that the on-site computing errors can be learned with the application data on the server and the retrained models can be resilient to the soft errors. Meanwhile, we propose an optimized partial TMR strategy to enhance the retraining. According to our experiments, R2F enables elastic design trade-offs between the model accuracy and the performance penalty. The top-5 model accuracy can be improved by 1.93%-13.73% with 0%-200% performance penalty at high fault error rate. In addition, we notice that the retraining requires massive data transmission and even dominates the training time, and propose a sparse increment compression approach for the data transmission optimization, which reduces the retraining time by 38%-88% on average with negligible accuracy loss over a straightforward remote retraining.

cs.AR↗

Phonon-related monochromatic THz radiation and its magneto-modulation in 2D ferromagnetic Cr2Ge2Te6

Searching multiple types of terahertz (THz) irradiation source is crucial for the THz technology. Here, by utilizing a two-dimensional (2D) ferromagnetic Cr2Ge2Te6 crystal, we firstly demonstrate a magneto-tunable monochromatic THz irradiation source. With a low-photonic-energy broadband THz pump, a strong THz irradiation with frequency ~0.9 THz and bandwidth ~0.25 THz can be generated and its conversion efficiency could even reach 2.1% at 160 K. Moreover, it is intriguing to find that such monochromatic THz irradiation can be efficiently modulated by the magnetic field below 160 K. According to both experimental and theoretical analyses, the emergent THz irradiation is identified as the emission from the phonon-polariton and its temperature and magnetic field dependent behaviors confirmed the large spin-lattice coupling in this 2D ferromagnetic crystal. These observations provide a new route for the creation of tunable monochromatic THz source which may have great practical interests in future applications in photonic and spintronic devices.

physics.optics↗

Deep Learning-Based Anomaly Detection in Cyber-Physical Systems: Progress and Opportunities

Anomaly detection is crucial to ensure the security of cyber-physical systems (CPS). However, due to the increasing complexity of CPSs and more sophisticated attacks, conventional anomaly detection methods, which face the growing volume of data and need domain-specific knowledge, cannot be directly applied to address these challenges. To this end, deep learning-based anomaly detection (DLAD) methods have been proposed. In this paper, we review state-of-the-art DLAD methods in CPSs. We propose a taxonomy in terms of the type of anomalies, strategies, implementation, and evaluation metrics to understand the essential properties of current methods. Further, we utilize this taxonomy to identify and highlight new characteristics and designs in each CPS domain. Also, we discuss the limitations and open problems of these methods. Moreover, to give users insights into choosing proper DLAD methods in practice, we experimentally explore the characteristics of typical neural models, the workflow of DLAD methods, and the running performance of DL models. Finally, we discuss the deficiencies of DL approaches, our findings, and possible directions to improve DLAD methods and motivate future research.

cs.CR↗

MPC-CSAS: Multi-Party Computation for Real-time Privacy-preserving Speed Advisory Systems

As a part of Advanced Driver Assistance Systems (ADASs), Consensus-based Speed Advisory Systems (CSAS) have been proposed to recommend a common speed to a group of vehicles for specific application purposes, such as emission control and energy management. With Vehicle-to-Vehicle (V2V), Vehicle-to-Infrastructure (V2I) technologies and advanced control theories in place, state-of-the-art CSAS can be designed to get an optimal speed in a privacy-preserving and decentralized manner. However, the current method only works for specific cost functions of vehicles, and its execution usually involves many algorithm iterations leading long convergence time. Therefore, the state-of-the-art design method is not applicable to a CSAS design which requires real-time decision making. In this paper, we address the problem by introducing MPC-CSAS, a Multi-Party Computation (MPC) based design approach for privacy-preserving CSAS. Our proposed method is simple to implement and applicable to all types of cost functions of vehicles. Moreover, our simulation results show that the proposed MPC-CSAS can achieve very promising system performance in just one algorithm iteration without using extra infrastructure for a typical CSAS.

eess.SY↗

Should bike sharing continue operating during the COVID-19 pandemic? Empirical findings from Nanjing, China

Coronavirus disease 2019 (COVID-19) has triggered a worldwide outbreak of pandemic, and transportation services have played a key role in coronavirus transmission. Although not crowded in a confined space like a bus or a metro car, bike sharing users will be exposed to the bike surface and take the transmission risk. During the COVID-19 pandemic, how to meet user demand and avoid virus spreading has become an important issue for bike sharing. Based on the trip data of bike sharing in Nanjing, China, this study analyzes the travel demand and operation management before and after the pandemic outbreak from the perspective of stations, users, and bikes. Semi-logarithmic difference-in-differences model, visualization methods, and statistic indexes are applied to explore the transportation service and risk prevention of bike sharing during the pandemic. The results show that pandemic control strategies sharply reduced user demand, and commuting trips decreased more significantly. Some stations around health and religious places become more important. Men and older adults are more dependent on bike sharing systems. Besides, the trip decrease reduces user contact and increases idle bikes. And a new concept of user distancing is proposed to avoid transmission risk and activate idle bikes. This study evaluates the role of shared micro-mobility during the COVID-19 pandemic, and also inspires the blocking of viral transmission within the city.

physics.soc-ph↗

Understanding High-Field Electron Transport Properties of Monolayer Transition Metal Dichalcogenides and Strain Effects

Monolayer transition metal dichalcogenides (MX2) are promising candidates for future electronics. Although the transport properties (e.g. mobility) at low electric field have been widely studied, there are limited studies on high-field properties, which are important for many applications. Particularly, there is lack of understanding of the physical origins underlying the property differences across different MX2. Here by combining first-principles calculations with Monte Carlo simulations, we study the high-field electron transport in defects-free unstrained and tensilely strained MX2 (M=Mo, W and X=S, Se). We find that WS2 has the highest peak velocity (due to its smallest effective mass) that can be reached at the lowest electric field (owing to its highest mobility). Strain can increase the peak velocity by increasing the scattering energy. After reaching the peak velocity, most MX2 demonstrates negative differential mobility (NDM). WS2 shows the largest NDM among unstrained MX2 due to the strongest effect of electron transfer from the low-energy small-mass valley to the high-energy large-mass valley. The tensile strain increases the valley separation, which on one hand suppresses the electron transfer in WS2, on the other hand allows the electrons to access the non-parabolic band region of the low-energy valley. The latter effect leads to an NDM for electrons in the low-energy valley, which can significantly increase the overall NDM at moderate strain. The valley-separation induced NDM in the low-energy valley is found to be a general phenomenon. Our work unveils the physical factors underlying the differences in high-field transport properties of different MX2, and also identifies the most promising candidate as well as effective approach for further improvement.

cond-mat.mtrl-sci↗

Why Two-Dimensional Semiconductors Generally Have Low Electron Mobility

Atomically thin (two-dimensional, 2D) semiconductors have shown great potential as the fundamental building blocks for next-generation electronics. However, all the 2D semiconductors that have been experimentally made so far have room-temperature electron mobility lower than that of bulk silicon, which is not understood. Here, by using first-principles calculations and reformulating the transport equations to isolate and quantify contributions of different mobility-determining factors, we show that the universally low mobility of 2D semiconductors originates from the high 'density of scatterings,' which is intrinsic to the 2D material with a parabolic electron band. The density of scatterings characterizes the density of phonons that can interact with the electrons and can be fully determined from the electron and phonon band structures without knowledge of electron-phonon coupling strength. Our work reveals the underlying physics limiting the electron mobility of 2D semiconductors and offers a descriptor to quickly assess the mobility.

cond-mat.mtrl-sci↗

Speculative Container Scheduling for Deep Learning Applications in a Kubernetes Cluster

In the past decade, we have witnessed a dramatically increasing volume of data collected from varied sources. The explosion of data has transformed the world as more information is available for collection and analysis than ever before. To maximize the utilization, various machine and deep learning models have been developed, e.g. CNN [1] and RNN [2], to study data and extract valuable information from different perspectives. While data-driven applications improve countless products, training models for hyperparameter tuning is still a time-consuming and resource-intensive process. Cloud computing provides infrastructure support for the training of deep learning applications. The cloud service providers, such as Amazon Web Services [3], create an isolated virtual environment (virtual machines and containers) for clients, who share physical resources, e.g., CPU and memory. On the cloud, resource management schemes are implemented to enable better sharing among users and boost the system-wide performance. However, general scheduling approaches, such as spread priority and balanced resource schedulers, do not work well with deep learning workloads. In this project, we propose SpeCon, a novel container scheduler that is optimized for shortlived deep learning applications. Based on virtualized containers, such as Kubernetes [4] and Docker [5], SpeCon analyzes the common characteristics of training processes. We design a suite of algorithms to monitor the progress of the training and speculatively migrate the slow-growing models to release resources for fast-growing ones. Specifically, the extensive experiments demonstrate that SpeCon improves the completion time of an individual job by up to 41.5%, 14.8% system-wide and 24.7% in terms of makespan.

cs.DC↗