arXiv · 2109.03529
RefineCap: Concept-Aware Refinement for Image Captioning
Abstract
Automatically translating images to texts involves image scene understanding and language modeling. In this paper, we propose a novel model, termed RefineCap, that refines the output vocabulary of the language decoder using decoder-guided visual semantics, and implicitly learns the mapping between visual tag words and images. The proposed Visual-Concept Refinement method can allow the generator to attend to semantic details in the image, thereby generating more semantically descriptive captions. Our model achieves superior performance on the MS-COCO dataset in comparison with previous visual-concept based models.
Explore related subjects
Keep this discovery
Yekun Chai, Shuo Jin, Junliang Xing. 2021-09-08. RefineCap: Concept-Aware Refinement for Image Captioning. https://arxiv.org/abs/2109.03529
Cite the original work for its findings. Save a collection to share your selection of sources.