arXiv · 2609.16565
Vision And Text Transformer For Predicting Answerability On Visual Question Answering
Abstract
Answerability on Visual Question Answering is a novel and attractive task to predict answerable scores between images and questions in multi-modal data. Existing works often utilize a binary mapping from visual question answering systems into Answerability. It does not reflect the essence of this problem. Together with our consideration of Answerability in a regression task, we propose VT-Transformer, which exploits visual and textual features through Transformer architecture. Experimental results on VizWiz 2020 dataset show the effectiveness and robustness of VT-Transformer for Answerability on Visual Question Answering when comparing with competitive baselines.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Tung Le, Huy Tien Nguyen, Le Minh Nguyen. 2026-09-15. Vision And Text Transformer For Predicting Answerability On Visual Question Answering. https://doi.org/10.1109/icip42928.2021.9506796
Cite the original work for its findings. Save a collection to share your selection of sources.