arXiv · 1909.05424
VizSeq: A Visual Analysis Toolkit for Text Generation Tasks
Abstract
Automatic evaluation of text generation tasks (e.g. machine translation, text summarization, image captioning and video description) usually relies heavily on task-specific metrics, such as BLEU and ROUGE. They, however, are abstract numbers and are not perfectly aligned with human assessment. This suggests inspecting detailed examples as a complement to identify system error patterns. In this paper, we present VizSeq, a visual analysis toolkit for instance-level and corpus-level system evaluation on a wide variety of text generation tasks. It supports multimodal sources and multiple text references, providing visualization in Jupyter notebook or a web app interface. It can be used locally or deployed onto public servers for centralized data hosting and benchmarking. It covers most common n-gram based metrics accelerated with multiprocessing, and also provides latest embedding-based metrics such as BERTScore.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Changhan Wang, Anirudh Jain, Danlu Chen, Jiatao Gu. 2019-09-12. VizSeq: A Visual Analysis Toolkit for Text Generation Tasks. https://doi.org/10.18653/v1%2Fd19-3043
Cite the original work for its findings. Save a collection to share your selection of sources.