arXiv · 2012.11691
Alleviating Noisy Data in Image Captioning with Cooperative Distillation
Abstract
Image captioning systems have made substantial progress, largely due to the availability of curated datasets like Microsoft COCO or Vizwiz that have accurate descriptions of their corresponding images. Unfortunately, scarce availability of such cleanly labeled data results in trained algorithms producing captions that can be terse and idiosyncratically specific to details in the image. We propose a new technique, cooperative distillation that combines clean curated datasets with the web-scale automatically extracted captions of the Google Conceptual Captions dataset (GCC), which can have poor descriptions of images, but is abundant in size and therefore provides a rich vocabulary resulting in more expressive captions.
Explore related subjects
Keep this discovery
Pierre Dognin, Igor Melnyk, Youssef Mroueh, Inkit Padhi, Mattia Rigotti, Jarret Ross, Yair Schiff. 2020-12-21. Alleviating Noisy Data in Image Captioning with Cooperative Distillation. https://arxiv.org/abs/2012.11691
Cite the original work for its findings. Save a collection to share your selection of sources.