paper-with-me

홈 › Papers

DUAL: Textless Spoken Question Answering with Speech Discrete Unit Adaptive Learning

2022-01-16 · ACL ARR January 2022 1 · Anonymous

Spoken Question Answering (SQA) has gained research attention and made remarkable progress in recent years. However, existing SQA methods rely on Automatic Speech Recognition (ASR) transcriptions, which is time and cost-prohibitive to collect. This work proposes an ASR transcription-free SQA framework named Discrete Unit Adaptive Learning (DUAL), which leverages unlabeled data for pre-training and is fine-tuned by the SQA downstream task. DAUL can directly predict the time interval of the spoken answer from the spoken document. We also release a new SQA benchmark corpus Natural Multi-speaker Spoken Question Answering (NMSQA) for testing SQA in realistic scenarios. The experimental results show that DUAL performs competitively with the cascade approach (ASR + text QA), and DUAL is robust to real-world speech. We will open-source our code and model to inspire more SQA innovations from the community.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Question Answeringspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

DUAL: Discrete Spoken Unit Adaptive Learning for Textless Spoken Question Answering

2022-03-09 · Guan-Ting Lin, Yung-Sung Chuang, Ho-Lam Chung, Shu-wen Yang 외

Spoken Question Answering (SQA) is to find the answer from a spoken document given a question, which is crucial for personal assistants when replying to the queries from the users. Existing SQA methods all rely on Automa…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Question Answeringspeech-recognition+1

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text

2025-09-09 · Peijin Xie, Shun Qian, Bingquan Liu, Dexin Wang 외 arxiv

Document images encapsulate a wealth of knowledge, while the portability of spoken queries enables broader and flexible application scenarios. Yet, no prior work has explored knowledge base question answering over visual…

Knowledge Base Question Answering

TM-PATHVQA:90000+ Textless Multilingual Questions for Medical Visual Question Answering

2024-07-16 · Tonmoy Rajkhowa, Amartya Roy Chowdhury, Sankalp Nagaonkar, Achyut Mani Tripathi

In healthcare and medical diagnostics, Visual Question Answering (VQA) mayemergeasapivotal tool in scenarios where analysis of intricate medical images becomes critical for accurate diagnoses. Current text-based VQA syst…

Medical Visual Question AnsweringQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

SViQA: A Unified Speech-Vision Multimodal Model for Textless Visual Question Answering

2025-04-01 · Bingxin Li

Multimodal models integrating speech and vision hold significant potential for advancing human-computer interaction, particularly in Speech-Based Visual Question Answering (SBVQA) where spoken questions about images requ…

cross-modal alignmentQuestion AnsweringVisual Question Answering

textless-lib: a Library for Textless Spoken Language Processing

2022-02-15 · NAACL (ACL) 2022 7 · Eugene Kharitonov, Jade Copet, Kushal Lakhotia, Tu Anh Nguyen 외

Textless spoken language processing research aims to extend the applicability of standard NLP toolset onto spoken language and languages with few or no textual resources. In this paper, we introduce textless-lib, a PyTor…

Resynthesis