Searcharxiv⌕ Search

arXiv subjects

Xiaoyan Gao

Publications and source records attributed to Xiaoyan Gao.

8 recordsLinked to original sources

Beyond Exact Match: Semantically Reassessing Event Extraction by Large Language Models

Event extraction has gained extensive research attention due to its broad range of applications. However, the current mainstream evaluation method for event extraction relies on token-level exact match, which misjudges numerous semantic-level correct cases. This reliance leads to a significant discrepancy between the evaluated performance of models under exact match criteria and their real performance. To address this problem, we propose a reliable and semantic evaluation framework for event extraction, named RAEE, which accurately assesses extraction results at semantic-level instead of token-level. Specifically, RAEE leverages large language models (LLMs) as evaluation agents, incorporating an adaptive mechanism to achieve adaptive evaluations for precision and recall of triggers and arguments. Extensive experiments demonstrate that: (1) RAEE achieves a very strong correlation with human judgments; (2) after reassessing 14 models, including advanced LLMs, on 10 datasets, there is a significant performance gap between exact match and RAEE. The exact match evaluation significantly underestimates the performance of existing event extraction models, and in particular underestimates the capabilities of LLMs; (3) fine-grained analysis under RAEE evaluation reveals insightful phenomena worth further exploration. The evaluation toolkit of our proposed RAEE is publicly released.

cs.CL↗

Ultra-Fast, Low-Storage, Highly Effective Coarse-grained Selection in Retrieval-based Chatbot by Using Deep Semantic Hashing

We study the coarse-grained selection module in retrieval-based chatbot. Coarse-grained selection is a basic module in a retrieval-based chatbot, which constructs a rough candidate set from the whole database to speed up the interaction with customers. So far, there are two kinds of approaches for coarse-grained selection module: (1) sparse representation; (2) dense representation. To the best of our knowledge, there is no systematic comparison between these two approaches in retrieval-based chatbots, and which kind of method is better in real scenarios is still an open question. In this paper, we first systematically compare these two methods from four aspects: (1) effectiveness; (2) index stoarge; (3) search time cost; (4) human evaluation. Extensive experiment results demonstrate that dense representation method significantly outperforms the sparse representation, but costs more time and storage occupation. In order to overcome these fatal weaknesses of dense representation method, we propose an ultra-fast, low-storage, and highly effective Deep Semantic Hashing Coarse-grained selection method, called DSHC model. Specifically, in our proposed DSHC model, a hashing optimizing module that consists of two autoencoder models is stacked on a trained dense representation model, and three loss functions are designed to optimize it. The hash codes provided by hashing optimizing module effectively preserve the rich semantic and similarity information in dense vectors. Extensive experiment results prove that, our proposed DSHC model can achieve much faster speed and lower storage than sparse representation, with limited performance loss compared with dense representation. Besides, our source codes have been publicly released for future research.

cs.CL↗

PONE: A Novel Automatic Evaluation Metric for Open-Domain Generative Dialogue Systems

Open-domain generative dialogue systems have attracted considerable attention over the past few years. Currently, how to automatically evaluate them, is still a big challenge problem. As far as we know, there are three kinds of automatic methods to evaluate the open-domain generative dialogue systems: (1) Word-overlap-based metrics; (2) Embedding-based metrics; (3) Learning-based metrics. Due to the lack of systematic comparison, it is not clear which kind of metrics are more effective. In this paper, we will first measure systematically all kinds of automatic evaluation metrics over the same experimental setting to check which kind is best. Through extensive experiments, the learning-based metrics are demonstrated that they are the most effective evaluation metrics for open-domain generative dialogue systems. Moreover, we observe that nearly all learning-based metrics depend on the negative sampling mechanism, which obtains an extremely imbalanced and low-quality dataset to train a score model. In order to address this issue, we propose a novel and feasible learning-based metric that can significantly improve the correlation with human judgments by using augmented POsitive samples and valuable NEgative samples, called PONE. Extensive experiments demonstrate that our proposed evaluation method significantly outperforms the state-of-the-art learning-based evaluation methods, with an average correlation improvement of 13.18%. In addition, we have publicly released the codes of our proposed method and state-of-the-art baselines.

cs.CL↗

Hierarchical Attention Network for Visually-aware Food Recommendation

Food recommender systems play an important role in assisting users to identify the desired food to eat. Deciding what food to eat is a complex and multi-faceted process, which is influenced by many factors such as the ingredients, appearance of the recipe, the user's personal preference on food, and various contexts like what had been eaten in the past meals. In this work, we formulate the food recommendation problem as predicting user preference on recipes based on three key factors that determine a user's choice on food, namely, 1) the user's (and other users') history; 2) the ingredients of a recipe; and 3) the descriptive image of a recipe. To address this challenging problem, we develop a dedicated neural network based solution Hierarchical Attention based Food Recommendation (HAFR) which is capable of: 1) capturing the collaborative filtering effect like what similar users tend to eat; 2) inferring a user's preference at the ingredient level; and 3) learning user preference from the recipe's visual images. To evaluate our proposed method, we construct a large-scale dataset consisting of millions of ratings from AllRecipes.com. Extensive experiments show that our method outperforms several competing recommender solutions like Factorization Machine and Visual Bayesian Personalized Ranking with an average improvement of 12%, offering promising results in predicting user preference for food. Codes and dataset will be released upon acceptance.

cs.IR↗

Joint Source-Channel Coding for Real-Time Video Transmission to Multi-homed Mobile Terminals

This study focuses on the mobile video delivery from a video server to a multi-homed client with a network of heterogeneous wireless. Joint Source-Channel Coding is effectively used to transmit video over bandwidth-limited, noisy wireless networks. But most existing JSCC methods only consider single path video transmission of the server and the client network. The problem will become more complicated when consider multi-path video transmission, because involving low-bandwidth, high-drop-rate or high-latency wireless network will only reduce the video quality. To solve this critical problem, we propose a novel Path Adaption JSCC (PA-JSCC) method that contain below characters: (1) path adaption, and (2) dynamic rate allocation. We use Exata to evaluate the performance of PA-JSCC and Experiment show that PA-JSCC has a good results in terms of PSNR (Peak Signal-to-Noise Ratio).

cs.NI↗

A New Method to Study the Origin of the EGB and the First Application on AT20G

In this letter, we introduce a new method of image stacking to directly study the undetected but possible gamma-ray point sources. Applying the method to the Australia Telescope 20 GHz Survey (AT20G) sources which have not been detected by LAT on Fermi, we find that the sources contribute (10.5+/-1.1)% and (4.3+/-0.9)% of the extragalactic gamma-ray background (EGB) and have a very soft spectrum with the photon indexes of 3.09+/-0.23 and 2.61+/-0.26, in the 1-3 and 3-300GeV energy ranges. In the 0.1-1GeV range, they probably contribute more large faction to the EGB, but it is not quite sure. It maybe not appropriate to assume that the undetected sources have the similar property to the detected sources.

astro-ph.HE↗

The new model of fitting the spectral energy distributions of Mkn 421 and Mkn 501

The spectral energy distribution (SED) of TeV blazars has a double-humped shape that is usually interpreted as Synchrotron Self Compton (SSC) model. The one zone SSC model is used broadly but cannot fit the high energy tail of SED very well. It need bulk Lorentz factor which is conflict with the observation. Furthermore one zone SSC model can not explain the entire spectrum. In the paper, we propose a new model that the high energy emission is produced by the accelerated protons in the blob with a small size and high magnetic field, the low energy radiation comes from the electrons in the expanded blob. Because the high and low energy photons are not produced at the same time, the requirement of large Doppler factor from pair production is relaxed. We present the fitting results of the SEDs for Mkn 501 during April 1997 and Mkn 421 during March 2001 respectively.

astro-ph.HE↗

Implications of Bulk Velocity Structures in AGN Jets

The synchrotron self-Compton (SSC) models and External Compton (EC) models of AGN jets with continually longitudinal and transverse bulk velocity structures are constructed. The observed spectra show complex and interesting patterns in different velocity structures and viewing angles. These models are used to calculate the synchrotron and inverse Compton spectra of two typical BL Lac objects (BLO) (Mrk 421 and 0716+714) and one Flat Spectrum Radio Quasars (FSRQs) (3c 279), and to discuss the implications of jet bulk velocity structures in unification of the BLO and FR I radio galaxies (FRI). By calculating the synchrotron spectra and SSC spectra of BL Lac object jets with continually bulk velocity structures, we find that the spectra are much different from ones in jets with uniform velocity structure under the increase of viewing angles. The unification of BLO and FRI is less constrained by viewing angles and would be imprinted by velocity structures intrinsic to the jet themselves. By considering the jets with bulk velocity structures constrained by apparent speed, we discuss the velocity structures imprinted on the observed spectra for different viewing angles. We find that the spectra are greatly impacted by longitudinal velocity structures, becasue the volume elements are compressed or expanded. Finally, we present the EC spectra of FSRQs and FR II radio galaxies (FRII) and find that they are weakly affected by velocity structures compared to synchrotron and SSC spectra.

astro-ph.GA↗