arXiv · 2412.00095
OPCap:Object-aware Prompting Captioning
Abstract
In the field of image captioning, the phenomenon where missing or nonexistent objects are used to explain an image is referred to as object bias (or hallucination). To mitigate this issue, we propose a target-aware prompting strategy. This method first extracts object labels and their spatial information from the image using an object detector. Then, an attribute predictor further refines the semantic features of the objects. These refined features are subsequently integrated and fed into the decoder, enhancing the model's understanding of the image context. Experimental results on the COCO and nocaps datasets demonstrate that OPCap effectively mitigates hallucination and significantly improves the quality of generated captions.
Explore related subjects
Keep this discovery
Feiyang Huang. 2024-11-27. OPCap:Object-aware Prompting Captioning. https://arxiv.org/abs/2412.00095
Cite the original work for its findings. Save a collection to share your selection of sources.