arXiv · 2311.10998
Learning Scene Context Without Images
Abstract
Teaching machines of scene contextual knowledge would enable them to interact more effectively with the environment and to anticipate or predict objects that may not be immediately apparent in their perceptual field. In this paper, we introduce a novel transformer-based approach called $LMOD$ ( Label-based Missing Object Detection) to teach scene contextual knowledge to machines using an attention mechanism. A distinctive aspect of the proposed approach is its reliance solely on labels from image datasets to teach scene context, entirely eliminating the need for the actual image itself. We show how scene-wide relationships among different objects can be learned using a self-attention mechanism. We further show that the contextual knowledge gained from label based learning can enhance performance of other visual based object detection algorithm.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Amirreza Rouhi, David Han. 2023-11-18. Learning Scene Context Without Images. https://arxiv.org/abs/2311.10998
Cite the original work for its findings. Save a collection to share your selection of sources.