Searcharxiv⌕ Search

arXiv subjects

Ting Han

Publications and source records attributed to Ting Han.

27 records · Page 2Linked to original sources

A spin-torque nano-oscillator based on interlayer-coupled meron-skyrmion pairs with a fixed orbit

In recent years, magnetic skyrmion-based spin-torque nano-oscillators (STNOs) attract considerable interest for their prospect in future-generation communication and spintronic technologies. However, some critical issues, which hamper their practical applications, e.g., the long start-up time and variable skyrmion gyration orbit, remain to be resolved. Here, we numerically demonstrate a realization of a fixed-orbit STNO, which is based on an interlayer-coupled meron-skyrmion (MS) pair other than a magnetic skyrmion. In this STNO, the MS pair possesses a structurally defined, fixed orbit within a broad range of driving current, even in the presence of random defects. The output frequency range of the STNO based on an MS pair far exceeds that of the STNO typically based on a single skyrmion. Moreover, the output frequency of this STNO can be further elevated if more MS pairs are incorporated. Our results reveal the nontrivial dynamics of the interlayer-coupled MS pair, opening perspectives for the design and optimization of fundamental spintronic devices.

cond-mat.mes-hall↗

Adaptive loose optimization for robust question answering

Question answering methods are well-known for leveraging data bias, such as the language prior in visual question answering and the position bias in machine reading comprehension (extractive question answering). Current debiasing methods often come at the cost of significant in-distribution performance to achieve favorable out-of-distribution generalizability, while non-debiasing methods sacrifice a considerable amount of out-of-distribution performance in order to obtain high in-distribution performance. Therefore, it is challenging for them to deal with the complicated changing real-world situations. In this paper, we propose a simple yet effective novel loss function with adaptive loose optimization, which seeks to make the best of both worlds for question answering. Our main technical contribution is to reduce the loss adaptively according to the ratio between the previous and current optimization state on mini-batch training data. This loose optimization can be used to prevent non-debiasing methods from overlearning data bias while enabling debiasing methods to maintain slight bias learning. Experiments on the visual question answering datasets, including VQA v2, VQA-CP v1, VQA-CP v2, GQA-OOD, and the extractive question answering dataset SQuAD demonstrate that our approach enables QA methods to obtain state-of-the-art in- and out-of-distribution performance in most cases. The source code has been released publicly in \url{https://github.com/reml-group/ALO}.

cs.CL↗

Fengshenbang 1.0: Being the Foundation of Chinese Cognitive Intelligence

Nowadays, foundation models become one of fundamental infrastructures in artificial intelligence, paving ways to the general intelligence. However, the reality presents two urgent challenges: existing foundation models are dominated by the English-language community; users are often given limited resources and thus cannot always use foundation models. To support the development of the Chinese-language community, we introduce an open-source project, called Fengshenbang, which leads by the research center for Cognitive Computing and Natural Language (CCNL). Our project has comprehensive capabilities, including large pre-trained models, user-friendly APIs, benchmarks, datasets, and others. We wrap all these in three sub-projects: the Fengshenbang Model, the Fengshen Framework, and the Fengshen Benchmark. An open-source roadmap, Fengshenbang, aims to re-evaluate the open-source community of Chinese pre-trained large-scale models, prompting the development of the entire Chinese large-scale model community. We also want to build a user-centered open-source ecosystem to allow individuals to access the desired models to match their computing resources. Furthermore, we invite companies, colleges, and research institutions to collaborate with us to build the large-scale open-source model-based ecosystem. We hope that this project will be the foundation of Chinese cognitive intelligence.

cs.CL↗

TCBERT: A Technical Report for Chinese Topic Classification BERT

Bidirectional Encoder Representations from Transformers or BERT~\cite{devlin-etal-2019-bert} has been one of the base models for various NLP tasks due to its remarkable performance. Variants customized for different languages and tasks are proposed to further improve the performance. In this work, we investigate supervised continued pre-training~\cite{gururangan-etal-2020-dont} on BERT for Chinese topic classification task. Specifically, we incorporate prompt-based learning and contrastive learning into the pre-training. To adapt to the task of Chinese topic classification, we collect around 2.1M Chinese data spanning various topics. The pre-trained Chinese Topic Classification BERTs (TCBERTs) with different parameter sizes are open-sourced at \url{https://huggingface.co/IDEA-CCNL}.

cs.CL↗

Coreference Augmentation for Multi-Domain Task-Oriented Dialogue State Tracking

Dialogue State Tracking (DST), which is the process of inferring user goals by estimating belief states given the dialogue history, plays a critical role in task-oriented dialogue systems. A coreference phenomenon observed in multi-turn conversations is not addressed by existing DST models, leading to sub-optimal performances. In this paper, we propose Coreference Dialogue State Tracker (CDST) that explicitly models the coreference feature. In particular, at each turn, the proposed model jointly predicts the coreferred domain-slot pair and extracts the coreference values from the dialogue context. Experimental results on MultiWOZ 2.1 dataset show that the proposed model achieves the state-of-the-art joint goal accuracy of 56.47%.

cs.CL↗

MultiWOZ 2.3: A multi-domain task-oriented dialogue dataset enhanced with annotation corrections and co-reference annotation

Task-oriented dialogue systems have made unprecedented progress with multiple state-of-the-art (SOTA) models underpinned by a number of publicly available MultiWOZ datasets. Dialogue state annotations are error-prone, leading to sub-optimal performance. Various efforts have been put in rectifying the annotation errors presented in the original MultiWOZ dataset. In this paper, we introduce MultiWOZ 2.3, in which we differentiate incorrect annotations in dialogue acts from dialogue states, identifying a lack of co-reference when publishing the updated dataset. To ensure consistency between dialogue acts and dialogue states, we implement co-reference features and unify annotations of dialogue acts and dialogue states. We update the state of the art performance of natural language understanding and dialogue state tracking on MultiWOZ 2.3, where the results show significant improvements than on previous versions of MultiWOZ datasets (2.0-2.2).

cs.CL↗

Enabling Robots to Draw and Tell: Towards Visually Grounded Multimodal Description Generation

Socially competent robots should be equipped with the ability to perceive the world that surrounds them and communicate about it in a human-like manner. Representative skills that exhibit such ability include generating image descriptions and visually grounded referring expressions. In the NLG community, these generation tasks are largely investigated in non-interactive and language-only settings. However, in face-to-face interaction, humans often deploy multiple modalities to communicate, forming seamless integration of natural language, hand gestures and other modalities like sketches. To enable robots to describe what they perceive with speech and sketches/gestures, we propose to model the task of generating natural language together with free-hand sketches/hand gestures to describe visual scenes and real life objects, namely, visually-grounded multimodal description generation. In this paper, we discuss the challenges and evaluation metrics of the task, and how the task can benefit from progress recently made in the natural language processing and computer vision realms, where related topics such as visually grounded NLG, distributional semantics, and photo-based sketch generation have been extensively studied.

cs.RO↗

Four-Dimensional Usability Investigation of Image CAPTCHA

Image CAPTCHA, aiming at effectively distinguishing human users from malicious script attacks, has been an important mechanism to protect online systems from spams and abuses. Despite the increasing interests in developing and deploying image CAPTCHAs, the usability aspect of those CAPTCHAs has hardly been explored systematically. In this paper, the universal design factors of image CAPTCHAs, such as image layouts, quantities, sizes, tilting angles and colors were experimentally evaluated through the following four dimensions: eye-tracking, efficiency, effectiveness and satisfaction. The cognitive processes revealed by eye-tracking indicate that the distribution of eye gaze is equally assigned to each candidate image and irrelevant to the variation of image contents. In addition, the gazing plot suggests that more than 70% of the participants inspected CAPTCHA images row-by-row, which is more efficient than scanning randomly. Those four-dimensional evaluations essentially suggest that square and horizontal rectangle are the preferred layout; image quantities may not exceed 16 while the image color is insignificant. Meanwhile, the image size and tilting angle are suggested to be larger than 55 pixels x 55 pixels and within -45~45 degrees, respectively. Basing on those usability experiment results, we proposed a design guideline that is expected to be useful for developing more usable image CAPTCHAs.

cs.HC↗

Usability Investigation on the Localization of Text CAPTCHAs: Take Chinese Characters as a Case Study

Text CAPTCHA has been an effective means to protect online systems from spams and abuses caused by automatic scripts which pretend to be human beings. However, nearly all the Text CAPTCHA designs in nowadays are based on English characters, which may not be the most user-friendly option for non-English speakers. Therefore, under the background of globalization, there is an increasing interest in designing local-language CAPTCHA, which is expected to be more usable for native speakers. However, systematic studies on the usability of localized CAPTCHAs are rare, and a general procedure for the design of usable localized CAPTCHA is still unavailable. Here, we comprehensively explored the design of CAPTCHAs based on Chinese characters from a usability perspective: cognitive processes of solving alphanumeric and Chinese CAPTCHAs are analyzed, followed by a usability comparison of those two types of CAPTCHAs and the evaluation of intrinsic design factors of Chinese CAPTCHAs. It was found that Chinese CAPTCHAs could be equally usable comparing with alphanumeric ones. Meanwhile, guidelines for the design of usable Chinese CAPTCHAs were also presented. Moreover, those design practices were also summarized as a general procedure which is expected to be applicable for the design of CAPTCHAs based on other languages.

cs.HC↗