arXiv · 2305.05432
WikiWeb2M: A Page-Level Multimodal Wikipedia Dataset
Abstract
Webpages have been a rich resource for language and vision-language tasks. Yet only pieces of webpages are kept: image-caption pairs, long text articles, or raw HTML, never all in one place. Webpage tasks have resultingly received little attention and structured image-text data underused. To study multimodal webpage understanding, we introduce the Wikipedia Webpage 2M (WikiWeb2M) suite; the first to retain the full set of images, text, and structure data available in a page. WikiWeb2M can be used for tasks like page description generation, section summarization, and contextual image captioning.
Explore related subjects
Keep this discovery
Andrea Burns, Krishna Srinivasan, Joshua Ainslie, Geoff Brown, Bryan A. Plummer, Kate Saenko, Jianmo Ni, Mandy Guo. 2023-05-09. WikiWeb2M: A Page-Level Multimodal Wikipedia Dataset. https://arxiv.org/abs/2305.05432
Cite the original work for its findings. Save a collection to share your selection of sources.