Searcharxiv⌕ Search

arXiv subjects

Yu Meng

Publications and source records attributed to Yu Meng.

At least 91 records · Page 5Linked to original sources

PIEClass: Weakly-Supervised Text Classification with Prompting and Noise-Robust Iterative Ensemble Training

Weakly-supervised text classification trains a classifier using the label name of each target class as the only supervision, which largely reduces human annotation efforts. Most existing methods first use the label names as static keyword-based features to generate pseudo labels, which are then used for final classifier training. While reasonable, such a commonly adopted framework suffers from two limitations: (1) keywords can have different meanings in different contexts and some text may not have any keyword, so keyword matching can induce noisy and inadequate pseudo labels; (2) the errors made in the pseudo label generation stage will directly propagate to the classifier training stage without a chance of being corrected. In this paper, we propose a new method, PIEClass, consisting of two modules: (1) a pseudo label acquisition module that uses zero-shot prompting of pre-trained language models (PLM) to get pseudo labels based on contextualized text understanding beyond static keyword matching, and (2) a noise-robust iterative ensemble training module that iteratively trains classifiers and updates pseudo labels by utilizing two PLM fine-tuning methods that regularize each other. Extensive experiments show that PIEClass achieves overall better performance than existing strong baselines on seven benchmark datasets and even achieves similar performance to fully-supervised classifiers on sentiment classification tasks.

cs.CL↗

Large Language Model as Attributed Training Data Generator: A Tale of Diversity and Bias

Large language models (LLMs) have been recently leveraged as training data generators for various natural language processing (NLP) tasks. While previous research has explored different approaches to training models using generated data, they generally rely on simple class-conditional prompts, which may limit the diversity of the generated data and inherit systematic biases of LLM. Thus, we investigate training data generation with diversely attributed prompts (e.g., specifying attributes like length and style), which have the potential to yield diverse and attributed generated data. Our investigation focuses on datasets with high cardinality and diverse domains, wherein we demonstrate that attributed prompts outperform simple class-conditional prompts in terms of the resulting model's performance. Additionally, we present a comprehensive empirical study on data generation encompassing vital aspects like bias, diversity, and efficiency, and highlight three key observations: firstly, synthetic datasets generated by simple prompts exhibit significant biases, such as regional bias; secondly, attribute diversity plays a pivotal role in enhancing model performance; lastly, attributed prompts achieve the performance of simple class-conditional prompts while utilizing only 5\% of the querying cost of ChatGPT associated with the latter. The data and code are available on \url{https://github.com/yueyu1030/AttrPrompt}.

cs.CL↗

Impossible ecologies: Interaction networks and stability of coexistence in ecological communities

Does an ecological community allow stable coexistence? Identifying the general principles that determine the answer to this question is a central problem of theoretical ecology. Random matrix theory approaches have uncovered the general trends of the effect of competitive, mutualistic, and predator-prey interactions between species on stability of coexistence. However, an ecological community is determined not only by the counts of these different interaction types, but also by their network arrangement. This cannot be accounted for in a direct statistical description that would enable random matrix theory approaches. Here, we therefore develop a different approach, of exhaustive analysis of small ecological communities, to show that this arrangement of interactions can influence stability of coexistence more than these general trends. We analyse all interaction networks of $N\leqslant 5$ species with Lotka-Volterra dynamics by combining exact results for $N\leqslant 3$ species and numerical exploration. Surprisingly, we find that a very small subset of these networks are "impossible ecologies", in which stable coexistence is non-trivially impossible. We prove that the possibility of stable coexistence in general ecologies is determined by similarly rare "irreducible ecologies". By random sampling of interaction strengths, we then show that the probability of stable coexistence varies over many orders of magnitude even in ecologies that differ only in the network arrangement of identical ecological interactions. Finally, we demonstrate that our approach can reveal the effect of evolutionary or environmental perturbations of the interaction network. Overall, this work reveals the importance of the full structure of the network of interactions for stability of coexistence in ecological communities.

q-bio.PE↗

Lattice QCD calculation of the invisible decay $J/ψ\rightarrow γν\barν$

In this work, we present the first lattice QCD study on the invisible decay $J/ψ\rightarrow γν\barν$. The calculation is accomplished using $N_f=2$ twisted mass fermion ensembles. The excited-state effects are observed and eliminated using a multi-state fit. The impact of finite-volume effects is also examined and confirmed to be well-controlled. After a continuous extrapolation under three lattice spacings, we obtain the branching fraction as $\operatorname{Br}[J/ψ\rightarrow γν\barν]=1.00(9)(7)\times 10^{-10}$, where the first error is the statistical error and the second is an estimate of the systematics. The exact theoretical prediction can be used to remove the only invisible contamination from the standard model background in searching for the possible dark matter by the channel $J/ψ\rightarrow γ+\textrm{invisible}$

hep-lat↗

First-principle calculation of $η_c\rightarrow 2γ$ decay width from lattice QCD

We perform a lattice QCD calculation of the $η_c\to2γ$ decay width using a model-independent method that requires no momentum extrapolation of the off-shell form factors. This method also provides a straightforward and simple way to examine the finite-volume effects. The calculation is accomplished using $N_f=2$ twisted mass fermion ensembles. The statistically significant excited-state effects are observed and eliminated using a multi-state fit.The impact of fine-tuning the charm quark mass is also examined and confirmed to be well-controlled. Finally, using three lattice spacings for the continuum extrapolation, we obtain the decay width $Γ_{η_cγγ}=6.67(16)_{\mathrm{stat}}(6)_{\mathrm{syst}}$ keV, which differs significantly from the Particle Data Group's reported value of $Γ_{η_cγγ}=5.4(4)$ keV (2.9~$σ$ tension). We provide insight into the comparison between our findings, previous theoretical predictions, and experimental measurements.

hep-lat↗

Weakly Supervised Multi-Label Classification of Full-Text Scientific Papers

Instead of relying on human-annotated training samples to build a classifier, weakly supervised scientific paper classification aims to classify papers only using category descriptions (e.g., category names, category-indicative keywords). Existing studies on weakly supervised paper classification are less concerned with two challenges: (1) Papers should be classified into not only coarse-grained research topics but also fine-grained themes, and potentially into multiple themes, given a large and fine-grained label space; and (2) full text should be utilized to complement the paper title and abstract for classification. Moreover, instead of viewing the entire paper as a long linear sequence, one should exploit the structural information such as citation links across papers and the hierarchy of sections and paragraphs in each paper. To tackle these challenges, in this study, we propose FUTEX, a framework that uses the cross-paper network structure and the in-paper hierarchy structure to classify full-text scientific papers under weak supervision. A network-aware contrastive fine-tuning module and a hierarchy-aware aggregation module are designed to leverage the two types of structural signals, respectively. Experiments on two benchmark datasets demonstrate that FUTEX significantly outperforms competitive baselines and is on par with fully supervised classifiers that use 1,000 to 60,000 ground-truth training samples.

cs.CL↗

Coherent control of an ultrabright single spin in hexagonal boron nitride at room temperature

Hexagonal boron nitride (hBN) is a remarkable two-dimensional (2D) material that hosts solid-state spins and has great potential to be used in quantum information applications, including quantum networks. However, in this application, both the optical and spin properties are crucial for single spins but have not yet been discovered simultaneously for hBN spins. Here, we realize an efficient method for arraying and isolating the single defects of hBN and use this method to discover a new spin defect with a high probability of 85%. This single defect exhibits outstanding optical properties and an optically controllable spin, as indicated by the observed significant Rabi oscillation and Hahn echo experiments at room temperature. First principles calculations indicate that complexes of carbon and oxygen dopants may be the origin of the single spin defects. This provides a possibility for further addressing spins that can be optically controlled.

physics.optics↗

Patton: Language Model Pretraining on Text-Rich Networks

A real-world text corpus sometimes comprises not only text documents but also semantic links between them (e.g., academic papers in a bibliographic network are linked by citations and co-authorships). Text documents and semantic connections form a text-rich network, which empowers a wide range of downstream tasks such as classification and retrieval. However, pretraining methods for such structures are still lacking, making it difficult to build one generic model that can be adapted to various tasks on text-rich networks. Current pretraining objectives, such as masked language modeling, purely model texts and do not take inter-document structure information into consideration. To this end, we propose our PretrAining on TexT-Rich NetwOrk framework Patton. Patton includes two pretraining strategies: network-contextualized masked language modeling and masked node prediction, to capture the inherent dependency between textual attributes and network structure. We conduct experiments on four downstream tasks in five datasets from both academic and e-commerce domains, where Patton outperforms baselines significantly and consistently.

cs.CL↗

ReGen: Zero-Shot Text Classification via Training Data Generation with Progressive Dense Retrieval

With the development of large language models (LLMs), zero-shot learning has attracted much attention for various NLP tasks. Different from prior works that generate training data with billion-scale natural language generation (NLG) models, we propose a retrieval-enhanced framework to create training data from a general-domain unlabeled corpus. To realize this, we first conduct contrastive pretraining to learn an unsupervised dense retriever for extracting the most relevant documents using class-descriptive verbalizers. We then further propose two simple strategies, namely Verbalizer Augmentation with Demonstrations and Self-consistency Guided Filtering to improve the topic coverage of the dataset while removing noisy examples. Experiments on nine datasets demonstrate that REGEN achieves 4.3% gain over the strongest baselines and saves around 70% of the time compared to baselines using large NLG models. Besides, REGEN can be naturally integrated with recently proposed large language models to boost performance.

cs.CL↗

Tuning Language Models as Training Data Generators for Augmentation-Enhanced Few-Shot Learning

Recent studies have revealed the intriguing few-shot learning ability of pretrained language models (PLMs): They can quickly adapt to a new task when fine-tuned on a small amount of labeled data formulated as prompts, without requiring abundant task-specific annotations. Despite their promising performance, most existing few-shot approaches that only learn from the small training set still underperform fully supervised training by nontrivial margins. In this work, we study few-shot learning with PLMs from a different perspective: We first tune an autoregressive PLM on the few-shot samples and then use it as a generator to synthesize a large amount of novel training samples which augment the original training set. To encourage the generator to produce label-discriminative samples, we train it via weighted maximum likelihood where the weight of each token is automatically adjusted based on a discriminative meta-learning objective. A classification PLM can then be fine-tuned on both the few-shot and the synthetic samples with regularization for better generalization and stability. Our approach FewGen achieves an overall better result across seven classification tasks of the GLUE benchmark than existing few-shot learning methods, improving no-augmentation methods by 5+ average points, and outperforming augmentation methods by 3+ average points.

cs.CL↗

Realization of algorithmic identification of cause and effect in quantum correlations

Causal inference revealing causal dependencies between variables from empirical data has found applications in multiple sub-fields of scientific research. A quantum perspective of correlations holds the promise of overcoming the limitation by Reichenbach's principle and enabling causal inference with only the observational data. However, it is still not clear how quantum causal inference can provide operational advantages in general cases. Here, we have devised a photonic setup and experimentally realized an algorithm capable of identifying any two-qubit statistical correlations generated by the two basic causal structures under an observational scenario, thus revealing a universal quantum advantage in causal inference over its classical counterpart. We further demonstrate the explainability and stability of our causal discovery method which is widely sought in data processing algorithms. Employing a fully observational approach, our result paves the way for studying quantum causality in general settings.

quant-ph↗

Laser Direct Writing of Visible Spin Defects in Hexagonal Boron Nitride for Applications in Spin-Based Technologies

Optically addressable spins in two-dimensional hexagonal boron nitride (hBN) attract widespread attention for their potential advantage in on-chip quantum devices, such as quantum sensors and quantum network. A variety of spin defects have been found in hBN, but no convenient and deterministic generation methods have been reported for other defects except negatively charged boron vacancy ($V_B^-$). Here we report that by using femtosecond laser direct writing technology, we can deterministically create spin defect ensembles with spectra range from 550 nm to 800 nm on nanoscale hBN flakes. Positive single-peak optically detected magnetic resonance (ODMR) signals are detected in the presence of magnetic field perpendicular to the substrate, and the contrast can reach 0.8%. With the appropriate thickness of hBN flakes, substrate and femtosecond laser pulse energy, we can deterministically and efficiently generate bright spin defect array. Our results provide a convenient deterministic method to create spin defects in hBN, which will motivate more endeavors for future researches and applications of spin-based technologies such as quantum magnetometer array.

physics.optics↗

Edgeformers: Graph-Empowered Transformers for Representation Learning on Textual-Edge Networks

Edges in many real-world social/information networks are associated with rich text information (e.g., user-user communications or user-product reviews). However, mainstream network representation learning models focus on propagating and aggregating node attributes, lacking specific designs to utilize text semantics on edges. While there exist edge-aware graph neural networks, they directly initialize edge attributes as a feature vector, which cannot fully capture the contextualized text semantics of edges. In this paper, we propose Edgeformers, a framework built upon graph-enhanced Transformers, to perform edge and node representation learning by modeling texts on edges in a contextualized way. Specifically, in edge representation learning, we inject network information into each Transformer layer when encoding edge texts; in node representation learning, we aggregate edge representations through an attention mechanism within each node's ego-graph. On five public datasets from three different domains, Edgeformers consistently outperform state-of-the-art baselines in edge classification and link prediction, demonstrating the efficacy in learning edge and node representations, respectively.

cs.LG↗

The Effect of Metadata on Scientific Literature Tagging: A Cross-Field Cross-Model Study

Due to the exponential growth of scientific publications on the Web, there is a pressing need to tag each paper with fine-grained topics so that researchers can track their interested fields of study rather than drowning in the whole literature. Scientific literature tagging is beyond a pure multi-label text classification task because papers on the Web are prevalently accompanied by metadata information such as venues, authors, and references, which may serve as additional signals to infer relevant tags. Although there have been studies making use of metadata in academic paper classification, their focus is often restricted to one or two scientific fields (e.g., computer science and biomedicine) and to one specific model. In this work, we systematically study the effect of metadata on scientific literature tagging across 19 fields. We select three representative multi-label classifiers (i.e., a bag-of-words model, a sequence-based model, and a pre-trained language model) and explore their performance change in scientific literature tagging when metadata are fed to the classifiers as additional features. We observe some ubiquitous patterns of metadata's effects across all fields (e.g., venues are consistently beneficial to paper tagging in almost all cases), as well as some unique patterns in fields other than computer science and biomedicine, which are not explored in previous studies.

cs.DL↗

Effective Seed-Guided Topic Discovery by Integrating Multiple Types of Contexts

Instead of mining coherent topics from a given text corpus in a completely unsupervised manner, seed-guided topic discovery methods leverage user-provided seed words to extract distinctive and coherent topics so that the mined topics can better cater to the user's interest. To model the semantic correlation between words and seeds for discovering topic-indicative terms, existing seed-guided approaches utilize different types of context signals, such as document-level word co-occurrences, sliding window-based local contexts, and generic linguistic knowledge brought by pre-trained language models. In this work, we analyze and show empirically that each type of context information has its value and limitation in modeling word semantics under seed guidance, but combining three types of contexts (i.e., word embeddings learned from local contexts, pre-trained language model representations obtained from general-domain training, and topic-indicative sentences retrieved based on seed information) allows them to complement each other for discovering quality topics. We propose an iterative framework, SeedTopicMine, which jointly learns from the three types of contexts and gradually fuses their context signals via an ensemble ranking process. Under various sets of seeds and on multiple datasets, SeedTopicMine consistently yields more coherent and accurate topics than existing seed-guided topic discovery approaches.

cs.CL↗

Generating Training Data with Language Models: Towards Zero-Shot Language Understanding

Pretrained language models (PLMs) have demonstrated remarkable performance in various natural language processing tasks: Unidirectional PLMs (e.g., GPT) are well known for their superior text generation capabilities; bidirectional PLMs (e.g., BERT) have been the prominent choice for natural language understanding (NLU) tasks. While both types of models have achieved promising few-shot learning performance, their potential for zero-shot learning has been underexplored. In this paper, we present a simple approach that uses both types of PLMs for fully zero-shot learning of NLU tasks without requiring any task-specific data: A unidirectional PLM generates class-conditioned texts guided by prompts, which are used as the training data for fine-tuning a bidirectional PLM. With quality training data selected based on the generation probability and regularization techniques (label smoothing and temporal ensembling) applied to the fine-tuning stage for better generalization and stability, our approach demonstrates strong performance across seven classification tasks of the GLUE benchmark (e.g., 72.3/73.8 on MNLI-m/mm and 92.8 on SST-2), significantly outperforming zero-shot prompting methods and achieving even comparable results to strong few-shot approaches using 32 training samples per class.

cs.CL↗

Reflective Dielectric Cavity Enhanced Emission from Hexagonal Boron Nitride Spin Defect Arrays

Among the various kinds of spin defects in hBN, the negatively charged boron vacancy ($\rm V_B^-$) spin defect that can be deterministically generated is undoubtedly a potential candidate for quantum sensing, but its low quantum efficiency restricts its %use in practical applications. Here, we demonstrate a robust enhancement structure with advantages including easy on-chip integration, convenient processing, low cost and suitable broad-spectrum enhancement for $\rm V_B^-$ defects. %Improved photoluminescence (PL) intensity and optically detected magnetic resonance (ODMR) contrast of $\rm V_B^-$ defect arrays. In the experiment, we used a metal reflective layer under the hBN flakes, filled with a transition dielectric layer in the middle, and adjusted the thickness of the dielectric layer to achieve the best coupling between the reflective dielectric cavity and the hBN spin defect. Using a reflective dielectric cavity, we achieved a PL enhancement of approximately 7-fold, and the corresponding ODMR contrast achieved 18\%. Additionally, the oxide layer of the reflective dielectric cavity can be used as an integrated material for micro-nano photonic devices for secondary processing, which means that it can be combined with other enhancement structures to achieve stronger enhancement. This work has guiding significance for realizing the on-chip integration of spin defects in two-dimensional materials.

quant-ph↗

Few-Shot Fine-Grained Entity Typing with Automatic Label Interpretation and Instance Generation

We study the problem of few-shot Fine-grained Entity Typing (FET), where only a few annotated entity mentions with contexts are given for each entity type. Recently, prompt-based tuning has demonstrated superior performance to standard fine-tuning in few-shot scenarios by formulating the entity type classification task as a ''fill-in-the-blank'' problem. This allows effective utilization of the strong language modeling capability of Pre-trained Language Models (PLMs). Despite the success of current prompt-based tuning approaches, two major challenges remain: (1) the verbalizer in prompts is either manually designed or constructed from external knowledge bases, without considering the target corpus and label hierarchy information, and (2) current approaches mainly utilize the representation power of PLMs, but have not explored their generation power acquired through extensive general-domain pre-training. In this work, we propose a novel framework for few-shot FET consisting of two modules: (1) an entity type label interpretation module automatically learns to relate type labels to the vocabulary by jointly leveraging few-shot instances and the label hierarchy, and (2) a type-based contextualized instance generator produces new instances based on given instances to enlarge the training set for better generalization. On three benchmark datasets, our model outperforms existing methods by significant margins. Code can be found at https://github.com/teapot123/Fine-Grained-Entity-Typing.

cs.CL↗