Speech-Based Visual Question Answering
This paper introduces speech-based visual question answering (VQA), the task of generating an answer given an image and a spoken question. Two methods are studied: an end-to-end, deep neural network that directly uses audio waveforms as input versus a pipelined approach that performs ASR (Automatic Speech Recognition) on the question, followed by text-based visual question answering. Furthermore, we investigate the robustness of both methods by injecting various levels of noise into the spoken question and find both methods to be tolerate noise at similar levels.
Code (1)
Tasks
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Question Answeringspeech-recognitionSpeech RecognitionVisual Question AnsweringVisual Question Answering (VQA)Similar Papers 제목 키워드 기반
SViQA: A Unified Speech-Vision Multimodal Model for Textless Visual Question Answering
Multimodal models integrating speech and vision hold significant potential for advancing human-computer interaction, particularly in Speech-Based Visual Question Answering (SBVQA) where spoken questions about images requ…
cross-modal alignmentQuestion AnsweringVisual Question AnsweringSpoken question answering for visual queries
Question answering (QA) systems are designed to answer natural language questions. Visual QA (VQA) and Spoken QA (SQA) systems extend the textual QA system to accept visual and spoken input respectively. This work aims t…
Question AnsweringVisual Question Answering (VQA)SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization
Visual Language Models have demonstrated remarkable capabilities across tasks, including visual question answering and image captioning. However, most models rely on text-based instructions, limiting their effectiveness …
Image CaptioningMultimodal ReasoningObject LocalizationObject Recognition+2Worldly Wise (WoW) - Cross-Lingual Knowledge Fusion for Fact-based Visual Spoken-Question Answering
Although Question-Answering has long been of research interest, its accessibility to users through a speech interface and its support to multiple languages have not been addressed in prior studies. Towards these ends, we…
Knowledge GraphsQuestion AnsweringSpeech-to-TextSpoken Language Understanding+1TM-PATHVQA:90000+ Textless Multilingual Questions for Medical Visual Question Answering
In healthcare and medical diagnostics, Visual Question Answering (VQA) mayemergeasapivotal tool in scenarios where analysis of intricate medical images becomes critical for accurate diagnoses. Current text-based VQA syst…
Medical Visual Question AnsweringQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)