paper-with-me

Papers

Vision And Text Transformer For Predicting Answerability On Visual Question Answering

2026-09-15 · Tung Le, Huy Tien Nguyen, Le Minh Nguyen arxiv

Answerability on Visual Question Answering is a novel and attractive task to predict answerable scores between images and questions in multi-modal data. Existing works often utilize a binary mapping from visual question answering systems into Answerability. It does not reflect the essence of this problem. Together with our consideration of Answerability in a regression task, we propose VT-Transformer, which exploits visual and textual features through Transformer architecture. Experimental results on VizWiz 2020 dataset show the effectiveness and robustness of VT-Transformer for Answerability on Visual Question Answering when comparing with competitive baselines.

📄 PDF Abstract BibTeX arXiv:2609.16565

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question Answering

Similar Papers 제목 키워드 기반

Answerability Fields: Answerable Location Estimation via Diffusion Models

2024-07-26 · Daichi Azuma, Taiki Miyanishi, Shuhei Kurita, Koya Sakamoto 외

In an era characterized by advancements in artificial intelligence and robotics, enabling machines to interact with and understand their environment is a critical research endeavor. In this paper, we propose Answerabilit…

Question AnsweringScene Understanding

Selectively Answering Visual Questions

2024-06-03 · Julian Martin Eisenschlos, Hernán Maina, Guido Ivetta, Luciana Benotti

Recently, large multi-modal models (LMMs) have emerged with the capacity to perform vision tasks such as captioning and visual question answering (VQA) with unprecedented accuracy. Applications such as helping the blind …

AvgIn-Context LearningQuestion AnsweringVisual Question Answering+1

VIVID-Med: LLM-Supervised Structured Pretraining for Deployable Medical ViTs

2026-03-10 · Xiyao Wang, Xiaoyu Tan, Yang Dai, Yuxuan Fu 외 arxiv

Vision-language pretraining has driven significant progress in medical image analysis. However, current methods typically supervise visual encoders using one-hot labels or free-form text, neither of which effectively cap…

Lung Nodule ClassificationStructured Prediction

Knowing When Not to Answer: Evaluating Abstention in Multimodal Reasoning Systems

2026-04-16 · Nishanth Madhusudhan, Vikas Yadav, Alexandre Lacoste arxiv

Effective abstention (EA), recognizing evidence insufficiency and refraining from answering, is critical for reliable multimodal systems. Yet existing evaluation paradigms for vision-language models (VLMs) and multi-agen…

Multimodal Reasoning

Challenges in Information-Seeking QA: Unanswerable Questions and Paragraph Retrieval

2020-10-22 · ACL 2021 5 · Akari Asai, Eunsol Choi

Recent pretrained language models "solved" many reading comprehension benchmarks, where questions are written with access to the evidence document. However, datasets containing information-seeking queries where evidence …

answerability predictionLanguage ModellingNatural QuestionsQuestion Answering+2