Searcharxiv⌕ Search

arXiv subjects

Yan Song

Publications and source records attributed to Yan Song.

At least 127 records · Page 7Linked to original sources

XLST: Cross-lingual Self-training to Learn Multilingual Representation for Low Resource Speech Recognition

In this paper, we propose a weakly supervised multilingual representation learning framework, called cross-lingual self-training (XLST). XLST is able to utilize a small amount of annotated data from high-resource languages to improve the representation learning on multilingual un-annotated data. Specifically, XLST uses a supervised trained model to produce initial representations and another model to learn from them, by maximizing the similarity between output embeddings of these two models. Furthermore, the moving average mechanism and multi-view data augmentation are employed, which are experimentally shown to be crucial to XLST. Comprehensive experiments have been conducted on the CommonVoice corpus to evaluate the effectiveness of XLST. Results on 5 downstream low-resource ASR tasks shows that our multilingual pretrained model achieves relatively 18.6% PER reduction over the state-of-the-art self-supervised method, with leveraging additional 100 hours of annotated English data.

eess.AS↗

Supertagging Combinatory Categorial Grammar with Attentive Graph Convolutional Networks

Supertagging is conventionally regarded as an important task for combinatory categorial grammar (CCG) parsing, where effective modeling of contextual information is highly important to this task. However, existing studies have made limited efforts to leverage contextual features except for applying powerful encoders (e.g., bi-LSTM). In this paper, we propose attentive graph convolutional networks to enhance neural CCG supertagging through a novel solution of leveraging contextual information. Specifically, we build the graph from chunks (n-grams) extracted from a lexicon and apply attention over the graph, so that different word pairs from the contexts within and across chunks are weighted in the model and facilitate the supertagging accordingly. The experiments performed on the CCGbank demonstrate that our approach outperforms all previous studies in terms of both supertagging and parsing. Further analyses illustrate the effectiveness of each component in our approach to discriminatively learn from word pairs to enhance CCG supertagging.

cs.CL↗

Named Entity Recognition for Social Media Texts with Semantic Augmentation

Existing approaches for named entity recognition suffer from data sparsity problems when conducted on short and informal texts, especially user-generated social media content. Semantic augmentation is a potential way to alleviate this problem. Given that rich semantic information is implicitly preserved in pre-trained word embeddings, they are potential ideal resources for semantic augmentation. In this paper, we propose a neural-based approach to NER for social media texts where both local (from running text) and augmented semantics are taken into account. In particular, we obtain the augmented semantic information from a large-scale corpus, and propose an attentive semantic augmentation module and a gate module to encode and aggregate such information, respectively. Extensive experiments are performed on three benchmark datasets collected from English and Chinese social media platforms, where the results demonstrate the superiority of our approach to previous studies across all three datasets.

cs.CL↗

Improving Named Entity Recognition with Attentive Ensemble of Syntactic Information

Named entity recognition (NER) is highly sensitive to sentential syntactic and semantic properties where entities may be extracted according to how they are used and placed in the running text. To model such properties, one could rely on existing resources to providing helpful knowledge to the NER task; some existing studies proved the effectiveness of doing so, and yet are limited in appropriately leveraging the knowledge such as distinguishing the important ones for particular context. In this paper, we improve NER by leveraging different types of syntactic information through attentive ensemble, which functionalizes by the proposed key-value memory networks, syntax attention, and the gate mechanism for encoding, weighting and aggregating such syntactic information, respectively. Experimental results on six English and Chinese benchmark datasets suggest the effectiveness of the proposed model and show that it outperforms previous studies on all experiment datasets.

cs.CL↗

Improving Constituency Parsing with Span Attention

Constituency parsing is a fundamental and important task for natural language understanding, where a good representation of contextual information can help this task. N-grams, which is a conventional type of feature for contextual information, have been demonstrated to be useful in many tasks, and thus could also be beneficial for constituency parsing if they are appropriately modeled. In this paper, we propose span attention for neural chart-based constituency parsing to leverage n-gram information. Considering that current chart-based parsers with Transformer-based encoder represent spans by subtraction of the hidden states at the span boundaries, which may cause information loss especially for long spans, we incorporate n-grams into span representations by weighting them according to their contributions to the parsing process. Moreover, we propose categorical span attention to further enhance the model by weighting n-grams within different length categories, and thus benefit long-sentence parsing. Experimental results on three widely used benchmark datasets demonstrate the effectiveness of our approach in parsing Arabic, Chinese, and English, where state-of-the-art performance is obtained by our approach on all of them.

cs.CL↗

Holographic flows with scalar self-interaction toward the Kasner universe

Considering a thermal state of the dual CFT with a uniform deformation by a scalar operator, we study a holographic renormalization group flow at nonzero temperature in the bulk described by the Einstein-scalar field theory with the self-interaction term $λϕ^4$ in asymptotic anti-de Sitter spacetime. We show that the holographic flow with the self-interaction term could run smoothly through the event horizon of a black hole and deform the Schwarzschild singularity to a Kasner universe at late times. Furthermore, we also study the effect of the scalar self-interaction on the deformed near-singularity Kasner exponents and the relationship between entanglement velocity and Kasner singularity exponents at late times.

hep-th↗

Exploring Unknown States with Action Balance

Exploration is a key problem in reinforcement learning. Recently bonus-based methods have achieved considerable successes in environments where exploration is difficult such as Montezuma's Revenge, which assign additional bonuses (e.g., intrinsic rewards) to guide the agent to rarely visited states. Since the bonus is calculated according to the novelty of the next state after performing an action, we call such methods as the next-state bonus methods. However, the next-state bonus methods force the agent to pay overmuch attention in exploring known states and ignore finding unknown states since the exploration is driven by the next state already visited, which may slow the pace of finding reward in some environments. In this paper, we focus on improving the effectiveness of finding unknown states and propose action balance exploration, which balances the frequency of selecting each action at a given state and can be treated as an extension of upper confidence bound (UCB) to deep reinforcement learning. Moreover, we propose action balance RND that combines the next-state bonus methods (e.g., random network distillation exploration, RND) and our action balance exploration to take advantage of both sides. The experiments on the grid world and Atari games demonstrate action balance exploration has a better capability in finding unknown states and can improve the performance of RND in some hard exploration environments respectively.

cs.AI↗

From Measuring Land Use Mix to Measuring Land Use Pattern -- New Methods for Assessing Land Use

Land use mix is one of the central concepts in the urban planning field, though its measure has been found to have many fallacies. In this study, we propose multiple alternative methods to the Conventional Shannon Entropy land use mix index, which is generally employed to measure land use diversity. The study attempted to measure land use mix from two new perspectives and test the performance of the new measures with economic, transportation, and social outcomes. We first offered an improved single measure, entropy-based land use mix index weighted by adding surrounding land use attributes and the regional land use context to mitigate the modifiable area unit problem (MAUP) and equal composition issue. Second, we designed a multi-measure method to describe land use patterns via clustering. We found that the new single measure method is more effective than the existing Conventional Shannon Entropy index in accurately delivering the vision of land use mix that Jane Jacobs originally proposed. We also found the multi-measure clustering method performed much better in the regression model on various effect variables, compared with both conventional and other new index. Upon further review, we found that land use pattern complexity's relationship with various outcomes makes a single definition measure very difficult to represent true land use fully. Therefore, we recommend that future research measure whether or not land use is mixed to then assess land use pattern.

physics.soc-ph↗

Weak cosmic censorship with self-interacting scalar and bound on charge to mass ratio

We study the model of Einstein-Maxwell theory minimally coupling to a massive charged self-interacting scalar field, parameterized by the quartic and hexic coupling, labelled by $λ$ and $β$, respectively. In the absence of scalar field, there is a class of counterexamples to cosmic censorship. Moveover, we investigate the properties of full nonlinear solution with nonzero scalar field, and argue that, by assuming massive charged self-interacting scalar field with sufficiently large charge above one certain bound, these counterexamples can be removed. In particular, this bound on charge for self-interacting scalar field is no longer equal to the weak gravity bound for free scalar case. In the quartic case, the bounds are below free scalar case for $λ<0$, while above free scalar case for $λ>0$. Meanwhile, in the hexic case, the bounds are above free scalar case for both $β>0$ and $β<0$.

hep-th↗

Self-interacting multistate boson stars

In this paper, we consider rotating multistate boson stars with quartic self-interactions. In contrast to the nodeless quartic-boson stars in \cite{Herdeiro:2015tia}, the self-interacting multistate boson stars (SIMBSs) have two types of nodes, including the $^1S^2S$ and $^1S^2P$ states. We show the mass $M$ of SIMBSs as a function of the synchronized frequency $ω$, and the nonsynchronized frequency $ω_2$ for three different cases. Moreover, for the case of two coexisting states with self-interacting potential, we study the mass $M$ of SIMBSs versus the angular momentum $J$ for the synchronized frequency $ω$ and the nonsynchronized frequency $ω_2$. Furthermore, for three different cases, we analyze the coexisting phase with both the ground and first excited states for SIMBSs. We also calculate the maximum value of coupling parameter $Λ$, and find the coupling parameter $Λ$ exists the finite range.

gr-qc↗

Conditional Augmentation for Aspect Term Extraction via Masked Sequence-to-Sequence Generation

Aspect term extraction aims to extract aspect terms from review texts as opinion targets for sentiment analysis. One of the big challenges with this task is the lack of sufficient annotated data. While data augmentation is potentially an effective technique to address the above issue, it is uncontrollable as it may change aspect words and aspect labels unexpectedly. In this paper, we formulate the data augmentation as a conditional generation task: generating a new sentence while preserving the original opinion targets and labels. We propose a masked sequence-to-sequence method for conditional augmentation of aspect term extraction. Unlike existing augmentation approaches, ours is controllable and allows us to generate more diversified sentences. Experimental results confirm that our method alleviates the data scarcity problem significantly. It also effectively boosts the performances of several current models for aspect term extraction.

cs.CL↗

Rotating multistate boson stars

In this paper, we construct rotating boson stars composed of the coexisting states of two scalar fields, including the ground and first excited states. We show the coexisting phase with both the ground and first excited states for rotating multistate boson stars. In contrast to the solutions of the nodeless boson stars, the rotating boson stars with two states have two types of nodes, including the $^1S^2S$ state and the $^1S^2P$ state. Moreover, we explore the properties of the mass $M$ of rotating boson stars with two states as a function of the synchronized frequency $ω$, as well as the nonsynchronized frequency $ω_2$. Finally, we also study the dependence of the mass $M$ of rotating boson stars with two states on angular momentum for both the synchronized frequency $ω$ and the nonsynchronized frequency $ω_2$.

gr-qc↗

Coordinated Reasoning for Cross-Lingual Knowledge Graph Alignment

Existing entity alignment methods mainly vary on the choices of encoding the knowledge graph, but they typically use the same decoding method, which independently chooses the local optimal match for each source entity. This decoding method may not only cause the "many-to-one" problem but also neglect the coordinated nature of this task, that is, each alignment decision may highly correlate to the other decisions. In this paper, we introduce two coordinated reasoning methods, i.e., the Easy-to-Hard decoding strategy and joint entity alignment algorithm. Specifically, the Easy-to-Hard strategy first retrieves the model-confident alignments from the predicted results and then incorporates them as additional knowledge to resolve the remaining model-uncertain alignments. To achieve this, we further propose an enhanced alignment model that is built on the current state-of-the-art baseline. In addition, to address the many-to-one problem, we propose to jointly predict entity alignments so that the one-to-one constraint can be naturally incorporated into the alignment prediction. Experimental results show that our model achieves the state-of-the-art performance and our reasoning methods can also significantly improve existing baselines.

cs.CL↗

Multiplex Word Embeddings for Selectional Preference Acquisition

Conventional word embeddings represent words with fixed vectors, which are usually trained based on co-occurrence patterns among words. In doing so, however, the power of such representations is limited, where the same word might be functionalized separately under different syntactic relations. To address this limitation, one solution is to incorporate relational dependencies of different words into their embeddings. Therefore, in this paper, we propose a multiplex word embedding model, which can be easily extended according to various relations among words. As a result, each word has a center embedding to represent its overall semantics, and several relational embeddings to represent its relational dependencies. Compared to existing models, our model can effectively distinguish words with respect to different relations without introducing unnecessary sparseness. Moreover, to accommodate various relations, we use a small dimension for relational embeddings and our model is able to keep their effectiveness. Experiments on selectional preference acquisition and word similarity demonstrate the effectiveness of the proposed model, and a further study of scalability also proves that our embeddings only need 1/20 of the original embedding size to achieve better performance.

cs.CL↗

ZEN: Pre-training Chinese Text Encoder Enhanced by N-gram Representations

The pre-training of text encoders normally processes text as a sequence of tokens corresponding to small text units, such as word pieces in English and characters in Chinese. It omits information carried by larger text granularity, and thus the encoders cannot easily adapt to certain combinations of characters. This leads to a loss of important semantic information, which is especially problematic for Chinese because the language does not have explicit word boundaries. In this paper, we propose ZEN, a BERT-based Chinese (Z) text encoder Enhanced by N-gram representations, where different combinations of characters are considered during training. As a result, potential word or phase boundaries are explicitly pre-trained and fine-tuned with the character encoder (BERT). Therefore ZEN incorporates the comprehensive information of both the character sequence and words or phrases it contains. Experimental results illustrated the effectiveness of ZEN on a series of Chinese NLP tasks. We show that ZEN, using less resource than other published encoders, can achieve state-of-the-art performance on most tasks. Moreover, it is shown that reasonable performance can be obtained when ZEN is trained on a small corpus, which is important for applying pre-training techniques to scenarios with limited data. The code and pre-trained models of ZEN are available at https://github.com/sinovation/zen.

cs.CL↗

Electrical Control of Large Rashba Effect in Oxide Heterostructures

Large Rashba effect efficiently tuned by an external electric field is highly desired for spintronic devices. Using first-principles calculations, we demonstrate that large Rashba splitting is locked at conduction band minimum in ferroelectric Bi(Sc/Y/La/Al/Ga/In)O3/PbTiO3 heterostructures where the position of Fermi level is precisely controlled via its stoichiometry. Fully reversible Rashba spin texture and drastic change of Rashba splitting strength with ferroelectric polarization switching are realized in the symmetric and asymmetric heterostructures, respectively. By artificially tuning the local ferroelectric displacement and the orbital hybridization, the synergetic effect of local potential gradient and orbital overlap on the dramatic change of splitting strength is confirmed. These results improve the feasibility of utilizing Rashba spin-orbit coupling in spintronic devices.

cond-mat.mtrl-sci↗

What You See is What You Get: Visual Pronoun Coreference Resolution in Dialogues

Grounding a pronoun to a visual object it refers to requires complex reasoning from various information sources, especially in conversational scenarios. For example, when people in a conversation talk about something all speakers can see, they often directly use pronouns (e.g., it) to refer to it without previous introduction. This fact brings a huge challenge for modern natural language understanding systems, particularly conventional context-based pronoun coreference models. To tackle this challenge, in this paper, we formally define the task of visual-aware pronoun coreference resolution (PCR) and introduce VisPro, a large-scale dialogue PCR dataset, to investigate whether and how the visual information can help resolve pronouns in dialogues. We then propose a novel visual-aware PCR model, VisCoref, for this task and conduct comprehensive experiments and case studies on our dataset. Results demonstrate the importance of the visual information in this PCR case and show the effectiveness of the proposed model.

cs.CL↗

Cross-lingual Knowledge Graph Alignment via Graph Matching Neural Network

Previous cross-lingual knowledge graph (KG) alignment studies rely on entity embeddings derived only from monolingual KG structural information, which may fail at matching entities that have different facts in two KGs. In this paper, we introduce the topic entity graph, a local sub-graph of an entity, to represent entities with their contextual information in KG. From this view, the KB-alignment task can be formulated as a graph matching problem; and we further propose a graph-attention based solution, which first matches all entities in two topic entity graphs, and then jointly model the local matching information to derive a graph-level matching vector. Experiments show that our model outperforms previous state-of-the-art methods by a large margin.

cs.LG↗