SearcharxivSearch

arXiv subjects

Haiyang Yu

Publications and source records attributed to Haiyang Yu.

140 records · Page 8Linked to original sources

Alexa Teacher Model: Pretraining and Distilling Multi-Billion-Parameter Encoders for Natural Language Understanding Systems

We present results from a large-scale experiment on pretraining encoders with non-embedding parameter counts ranging from 700M to 9.3B, their subsequent distillation into smaller models ranging from 17M-170M parameters, and their application to the Natural Language Understanding (NLU) component of a virtual assistant system. Though we train using 70% spoken-form data, our teacher models perform comparably to XLM-R and mT5 when evaluated on the written-form Cross-lingual Natural Language Inference (XNLI) corpus. We perform a second stage of pretraining on our teacher models using in-domain data from our system, improving error rates by 3.86% relative for intent classification and 7.01% relative for slot filling. We find that even a 170M-parameter model distilled from our Stage 2 teacher model has 2.88% better intent classification and 7.69% better slot filling error rates when compared to the 2.3B-parameter teacher trained only on public data (Stage 1), emphasizing the importance of in-domain data for pretraining. When evaluated offline using labeled NLU data, our 17M-parameter Stage 2 distilled model outperforms both XLM-R Base (85M params) and DistillBERT (42M params) by 4.23% to 6.14%, respectively. Finally, we present results from a full virtual assistant experimentation platform, where we find that models trained using our pretraining and distillation pipeline outperform models distilled from 85M-parameter teachers by 3.74%-4.91% on an automatic measurement of full-system user dissatisfaction.

cs.CL

Extragalactic HI survey with FAST : First look of the pilot survey results

As first data release of a pilot extragalactic HI survey with Five-hundred-meter Aperture Spherical radio Telescope (FAST),we extracted 544 extragalaxies from three-dimensional(3D) spectral data to perform interactive searching and computing, yielding global parameters for these detections, extending redshift ranges of HI 21cm line up to z = 0.04 ,which covers part of the sky region in right ascension(R.A. or $α$) and declination(Dec or $δ$) range $00^{\rm h} 47^{\rm m}< \rm R.A.(J2000)<23^{\rm h}22^{\rm m}$ and $+24^{\circ}<\rm Dec.(J2000) <+43^{\circ}$ . The S/N of 544 HI detections are greater than 5 flagged with code 1 to 4 based on baseline qualities or RFI contamination. Besides, we find 16 of which without any counterparts in the existing galaxy catalogs. The catalog can give a guidence for the future HI observation with FAST.

astro-ph.GA

Text Gestalt: Stroke-Aware Scene Text Image Super-Resolution

In the last decade, the blossom of deep learning has witnessed the rapid development of scene text recognition. However, the recognition of low-resolution scene text images remains a challenge. Even though some super-resolution methods have been proposed to tackle this problem, they usually treat text images as general images while ignoring the fact that the visual quality of strokes (the atomic unit of text) plays an essential role for text recognition. According to Gestalt Psychology, humans are capable of composing parts of details into the most similar objects guided by prior knowledge. Likewise, when humans observe a low-resolution text image, they will inherently use partial stroke-level details to recover the appearance of holistic characters. Inspired by Gestalt Psychology, we put forward a Stroke-Aware Scene Text Image Super-Resolution method containing a Stroke-Focused Module (SFM) to concentrate on stroke-level internal structures of characters in text images. Specifically, we attempt to design rules for decomposing English characters and digits at stroke-level, then pre-train a text recognizer to provide stroke-level attention maps as positional clues with the purpose of controlling the consistency between the generated super-resolution image and high-resolution ground truth. The extensive experimental results validate that the proposed method can indeed generate more distinguishable images on TextZoom and manually constructed Chinese character dataset Degraded-IC13. Furthermore, since the proposed SFM is only used to provide stroke-level guidance when training, it will not bring any time overhead during the test phase. Code is available at https://github.com/FudanVI/FudanOCR/tree/main/text-gestalt.

cs.CV

DIG: A Turnkey Library for Diving into Graph Deep Learning Research

Although there exist several libraries for deep learning on graphs, they are aiming at implementing basic operations for graph deep learning. In the research community, implementing and benchmarking various advanced tasks are still painful and time-consuming with existing libraries. To facilitate graph deep learning research, we introduce DIG: Dive into Graphs, a turnkey library that provides a unified testbed for higher level, research-oriented graph deep learning tasks. Currently, we consider graph generation, self-supervised learning on graphs, explainability of graph neural networks, and deep learning on 3D graphs. For each direction, we provide unified implementations of data interfaces, common algorithms, and evaluation metrics. Altogether, DIG is an extensible, open-source, and turnkey library for researchers to develop new methods and effortlessly compare with common baselines using widely used datasets and evaluation metrics. Source code is available at https://github.com/divelab/DIG.

cs.LG

On Explainability of Graph Neural Networks via Subgraph Explorations

We consider the problem of explaining the predictions of graph neural networks (GNNs), which otherwise are considered as black boxes. Existing methods invariably focus on explaining the importance of graph nodes or edges but ignore the substructures of graphs, which are more intuitive and human-intelligible. In this work, we propose a novel method, known as SubgraphX, to explain GNNs by identifying important subgraphs. Given a trained GNN model and an input graph, our SubgraphX explains its predictions by efficiently exploring different subgraphs with Monte Carlo tree search. To make the tree search more effective, we propose to use Shapley values as a measure of subgraph importance, which can also capture the interactions among different subgraphs. To expedite computations, we propose efficient approximation schemes to compute Shapley values for graph data. Our work represents the first attempt to explain GNNs via identifying subgraphs explicitly and directly. Experimental results show that our SubgraphX achieves significantly improved explanations, while keeping computations at a reasonable level.

cs.LG

Interventional Aspect-Based Sentiment Analysis

Recent neural-based aspect-based sentiment analysis approaches, though achieving promising improvement on benchmark datasets, have reported suffering from poor robustness when encountering confounder such as non-target aspects. In this paper, we take a causal view to addressing this issue. We propose a simple yet effective method, namely, Sentiment Adjustment (SENTA), by applying a backdoor adjustment to disentangle those confounding factors. Experimental results on the Aspect Robustness Test Set (ARTS) dataset demonstrate that our approach improves the performance while maintaining accuracy in the original test set.

cs.CL

Bridging Text and Knowledge with Multi-Prototype Embedding for Few-Shot Relational Triple Extraction

Current supervised relational triple extraction approaches require huge amounts of labeled data and thus suffer from poor performance in few-shot settings. However, people can grasp new knowledge by learning a few instances. To this end, we take the first step to study the few-shot relational triple extraction, which has not been well understood. Unlike previous single-task few-shot problems, relational triple extraction is more challenging as the entities and relations have implicit correlations. In this paper, We propose a novel multi-prototype embedding network model to jointly extract the composition of relational triples, namely, entity pairs and corresponding relations. To be specific, we design a hybrid prototypical learning mechanism that bridges text and knowledge concerning both entities and relations. Thus, implicit correlations between entities and relations are injected. Additionally, we propose a prototype-aware regularization to learn more representative prototypes. Experimental results demonstrate that the proposed method can improve the performance of the few-shot triple extraction.

cs.CL

Mining Truck Platooning Patterns Through Massive Trajectory Data

Truck platooning refers to a series of trucks driving in close proximity via communication technologies, and it is considered one of the most implementable systems of connected and automated vehicles, bringing huge energy savings and safety improvements. Properly planning platoons and evaluating the potential of truck platooning are crucial to trucking companies and transportation authorities. This study proposes a series of data mining approaches to learn spontaneous truck platooning patterns from massive trajectories. An enhanced map matching algorithm is developed to identify truck headings by using digital map data, followed by an adaptive spatial clustering algorithm to detect instantaneous co-moving truck sets. These sets are then aggregated to find the network-wide maximum platoon duration and size through frequent itemset mining for computational efficiency. We leverage real GPS data collected from truck fleeting systems in Liaoning Province, China, to evaluate platooning performance and successfully extract spatiotemporal platooning patterns. Results show that approximately 36% spontaneous truck platoons can be coordinated by speed adjustment without changing routes and schedules. The average platooning distance and duration ratios for these platooned trucks are 9.6% and 9.9%, respectively, leading to a 2.8% reduction in total fuel consumption. We also distinguish the optimal platooning periods and space headways for national freeways and trunk roads, and prioritize the road segments with high possibilities of truck platooning. The derived results are reproducible, providing useful policy implications and operational strategies for large-scale truck platoon planning and roadside infrastructure construction.

cs.LG

The Devil is the Classifier: Investigating Long Tail Relation Classification with Decoupling Analysis

Long-tailed relation classification is a challenging problem as the head classes may dominate the training phase, thereby leading to the deterioration of the tail performance. Existing solutions usually address this issue via class-balancing strategies, e.g., data re-sampling and loss re-weighting, but all these methods adhere to the schema of entangling learning of the representation and classifier. In this study, we conduct an in-depth empirical investigation into the long-tailed problem and found that pre-trained models with instance-balanced sampling already capture the well-learned representations for all classes; moreover, it is possible to achieve better long-tailed classification ability at low cost by only adjusting the classifier. Inspired by this observation, we propose a robust classifier with attentive relation routing, which assigns soft weights by automatically aggregating the relations. Extensive experiments on two datasets demonstrate the effectiveness of our proposed approach. Code and datasets are available in https://github.com/zjunlp/deepke.

cs.LG

Orientation dependence of the nano-indentation behaviour of pure tungsten

Coupling of nano-indentation and crystal plasticity finite element (CPFE) simulations is widely used to quantitatively probe the small-scale mechanical behaviour of materials. Earlier studies showed that CPFE can successfully reproduce the load-displacement curves and surface morphology for different crystal orientations. Here, we report the orientation dependence of residual lattice strain patterns and dislocation structures in tungsten. For orientations with one or more Burgers vectors close to parallel to the sample surface, dislocation movement and residual lattice strains are confined to long, narrow channels. CPFE is unable to reproduce this behaviour, and our analysis reveals the responsible underlying mechanisms.

cond-mat.mtrl-sci

A Survey on Complex Question Answering over Knowledge Base: Recent Advances and Challenges

Question Answering (QA) over Knowledge Base (KB) aims to automatically answer natural language questions via well-structured relation information between entities stored in knowledge bases. In order to make KBQA more applicable in actual scenarios, researchers have shifted their attention from simple questions to complex questions, which require more KB triples and constraint inference. In this paper, we introduce the recent advances in complex QA. Besides traditional methods relying on templates and rules, the research is categorized into a taxonomy that contains two main branches, namely Information Retrieval-based and Neural Semantic Parsing-based. After describing the methods of these branches, we analyze directions for future research and introduce the models proposed by the Alime team.

cs.CL

The influence of hydrogen core force shielding on dislocation junctions in iron

The influence of hydrogen on dislocation junctions was analysed by incorporating a hydrogen dependent core force into nodal based discrete dislocation dynamics. Hydrogen reduces the core energy of dislocations, which reduces the magnitude of the dislocation core force. We refer to this as hydrogen core force shielding, as it is analogous to hydrogen elastic shielding but occurs at much lower hydrogen concentrations. The dislocation core energy change due to hydrogen was calibrated at the atomic scale accounting for the nonlinear inter-atomic interactions at the dislocation core, giving the model a sound physical basis. Hydrogen was found to strengthen binary junctions and promote the nucleation of dislocations from triple junctions. Simulations of microcantilever bend tests with hydrogen core force shielding showed an increase in the junction density and subsequent hardening. These simulations were performed at a small hydrogen concentration realistic for bcc iron.

cond-mat.mtrl-sci

Discrete dislocation plasticity HELPs understand hydrogen effects in bcc materials

In an attempt to bridge the gap between atomistic and continuum plasticity simulations of hydrogen in iron, we present three dimensional discrete dislocation plasticity simulations incorporating the hydrogen elastic stress and a hydrogen dependent dislocation mobility law. The hydrogen induced stress is incorporated following the formulation derived by Gu and El-Awady (2018) which here we extend to a finite boundary value problem, a microcantilever beam, via the superposition principle. The hydrogen dependent mobility law is based on first principle calculations by Katzarov et al. (2017) and was found to promote dislocation generation and enhance slip planarity at a bulk hydrogen concentration of 0.1 appm; which is typical for bcc materials. The hydrogen elastic stress produced the same behaviour, but only when the bulk concentration was extremely high. In a microcantilever, hydrogen was found to promote dislocation activity which lowered the flow stress and generated more pronounced slip steps on the free surfaces. These observations are consistent with the hydrogen enhanced localized plasticity (HELP) mechanism, and it is concluded that both the hydrogen elastic stress and hydrogen increased dislocation mobility are viable explanations for HELP. However it is the latter that dominates at the low concentrations typically found in bcc metals.

cond-mat.mtrl-sci

Spatiotemporal Recurrent Convolutional Networks for Traffic Prediction in Transportation Networks

Predicting large-scale transportation network traffic has become an important and challenging topic in recent decades. Inspired by the domain knowledge of motion prediction, in which the future motion of an object can be predicted based on previous scenes, we propose a network grid representation method that can retain the fine-scale structure of a transportation network. Network-wide traffic speeds are converted into a series of static images and input into a novel deep architecture, namely, spatiotemporal recurrent convolutional networks (SRCNs), for traffic forecasting. The proposed SRCNs inherit the advantages of deep convolutional neural networks (DCNNs) and long short-term memory (LSTM) neural networks. The spatial dependencies of network-wide traffic can be captured by DCNNs, and the temporal dynamics can be learned by LSTMs. An experiment on a Beijing transportation network with 278 links demonstrates that SRCNs outperform other deep learning-based algorithms in both short-term and long-term traffic prediction.

cs.LG