arXiv · 2405.09335
Prompting-based Synthetic Data Generation for Few-Shot Question Answering
Abstract
Although language models (LMs) have boosted the performance of Question Answering, they still need plenty of data. Data annotation, in contrast, is a time-consuming process. This especially applies to Question Answering, where possibly large documents have to be parsed and annotated with questions and their corresponding answers. Furthermore, Question Answering models often only work well for the domain they were trained on. Since annotation is costly, we argue that domain-agnostic knowledge from LMs, such as linguistic understanding, is sufficient to create a well-curated dataset. With this motivation, we show that using large language models can improve Question Answering performance on various datasets in the few-shot setting compared to state-of-the-art approaches. For this, we perform data generation leveraging the Prompting framework, suggesting that language models contain valuable task-agnostic knowledge that can be used beyond the common pre-training/fine-tuning scheme. As a result, we consistently outperform previous approaches on few-shot Question Answering.
Explore related subjects
Keep this discovery
Maximilian Schmidt, Andrea Bartezzaghi, Ngoc Thang Vu. 2024-05-15. Prompting-based Synthetic Data Generation for Few-Shot Question Answering. https://arxiv.org/abs/2405.09335
Cite the original work for its findings. Save a collection to share your selection of sources.