Vision And Text Transformer For Predicting Answerability On Visual Question Answering
Answerability on Visual Question Answering is a novel and attractive task to predict answerable scores between images and questions in multi-modal data. Existing works often utilize a binary mapping from visual question answering systems into Answerability. It does not reflect the essence of this problem. Together with our consideration of Answerability in a regression task, we propose VT-Transformer, which exploits visual and textual features through Transformer architecture. Experimental results on VizWiz 2020 dataset show the effectiveness and robustness of VT-Transformer for Answerability on Visual Question Answering when comparing with competitive baselines.
Code (0)
등록된 구현이 없습니다.
Tasks
Visual Question AnsweringSimilar Papers 제목 키워드 기반
Answerability Fields: Answerable Location Estimation via Diffusion Models
In an era characterized by advancements in artificial intelligence and robotics, enabling machines to interact with and understand their environment is a critical research endeavor. In this paper, we propose Answerabilit…
Question AnsweringScene UnderstandingSelectively Answering Visual Questions
Recently, large multi-modal models (LMMs) have emerged with the capacity to perform vision tasks such as captioning and visual question answering (VQA) with unprecedented accuracy. Applications such as helping the blind …
AvgIn-Context LearningQuestion AnsweringVisual Question Answering+1VIVID-Med: LLM-Supervised Structured Pretraining for Deployable Medical ViTs
Vision-language pretraining has driven significant progress in medical image analysis. However, current methods typically supervise visual encoders using one-hot labels or free-form text, neither of which effectively cap…
Lung Nodule ClassificationStructured PredictionKnowing When Not to Answer: Evaluating Abstention in Multimodal Reasoning Systems
Effective abstention (EA), recognizing evidence insufficiency and refraining from answering, is critical for reliable multimodal systems. Yet existing evaluation paradigms for vision-language models (VLMs) and multi-agen…
Multimodal ReasoningChallenges in Information-Seeking QA: Unanswerable Questions and Paragraph Retrieval
Recent pretrained language models "solved" many reading comprehension benchmarks, where questions are written with access to the evidence document. However, datasets containing information-seeking queries where evidence …
answerability predictionLanguage ModellingNatural QuestionsQuestion Answering+2