paper-with-me

Papers

Speech-Based Visual Question Answering

2017-05-01 · Ted Zhang, Dengxin Dai, Tinne Tuytelaars, Marie-Francine Moens, Luc van Gool

This paper introduces speech-based visual question answering (VQA), the task of generating an answer given an image and a spoken question. Two methods are studied: an end-to-end, deep neural network that directly uses audio waveforms as input versus a pipelined approach that performs ASR (Automatic Speech Recognition) on the question, followed by text-based visual question answering. Furthermore, we investigate the robustness of both methods by injecting various levels of noise into the spoken question and find both methods to be tolerate noise at similar levels.

📄 PDF Abstract BibTeX arXiv:1705.00464

Code (1)

zted/sbvqa 공식 구현

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Question Answeringspeech-recognitionSpeech RecognitionVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

SViQA: A Unified Speech-Vision Multimodal Model for Textless Visual Question Answering

2025-04-01 · Bingxin Li

Multimodal models integrating speech and vision hold significant potential for advancing human-computer interaction, particularly in Speech-Based Visual Question Answering (SBVQA) where spoken questions about images requ…

cross-modal alignmentQuestion AnsweringVisual Question Answering

Spoken question answering for visual queries

2025-05-29 · Nimrod Shabtay, Zvi Kons, Avihu Dekel, Hagai Aronowitz 외

Question answering (QA) systems are designed to answer natural language questions. Visual QA (VQA) and Spoken QA (SQA) systems extend the textual QA system to accept visual and spoken input respectively. This work aims t…

Question AnsweringVisual Question Answering (VQA)

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization

2024-12-21 · Tan-Hanh Pham, Hoang-Nam Le, Phu-Vinh Nguyen, Chris Ngo 외

Visual Language Models have demonstrated remarkable capabilities across tasks, including visual question answering and image captioning. However, most models rely on text-based instructions, limiting their effectiveness …

Image CaptioningMultimodal ReasoningObject LocalizationObject Recognition+2

Worldly Wise (WoW) - Cross-Lingual Knowledge Fusion for Fact-based Visual Spoken-Question Answering

2021-06-01 · NAACL 2021 4 · Kiran Ramnath, Leda Sari, Mark Hasegawa-Johnson, Chang Yoo

Although Question-Answering has long been of research interest, its accessibility to users through a speech interface and its support to multiple languages have not been addressed in prior studies. Towards these ends, we…

Knowledge GraphsQuestion AnsweringSpeech-to-TextSpoken Language Understanding+1

TM-PATHVQA:90000+ Textless Multilingual Questions for Medical Visual Question Answering

2024-07-16 · Tonmoy Rajkhowa, Amartya Roy Chowdhury, Sankalp Nagaonkar, Achyut Mani Tripathi

In healthcare and medical diagnostics, Visual Question Answering (VQA) mayemergeasapivotal tool in scenarios where analysis of intricate medical images becomes critical for accurate diagnoses. Current text-based VQA syst…

Medical Visual Question AnsweringQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)