paper-with-me

홈 › Papers

End-to-end Contrastive Language-Speech Pretraining Model For Long-form Spoken Question Answering

2025-11-12 · Jiliang Hu, Zuchao Li, Baoyuan Qi, Liu Guoming, Ping Wang arxiv

Significant progress has been made in spoken question answering (SQA) in recent years. However, many existing methods, including large audio language models, struggle with processing long audio. Follow the success of retrieval augmented generation, a speech-related retriever shows promising in help preprocessing long-form speech. But the performance of existing speech-related retrievers is lacking. To address this challenge, we propose CLSR, an end-to-end contrastive language-speech retriever that efficiently extracts question-relevant segments from long audio recordings for downstream SQA task. Unlike conventional speech-text contrastive models, CLSR incorporates an intermediate step that converts acoustic features into text-like representations prior to alignment, thereby more effectively bridging the gap between modalities. Experimental results across four cross-modal retrieval datasets demonstrate that CLSR surpasses both end-to-end speech related retrievers and pipeline approaches combining speech recognition with text retrieval, providing a robust foundation for advancing practical long-form SQA applications.

📄 PDF Abstract BibTeX arXiv:2511.09282

Code (0)

등록된 구현이 없습니다.

Tasks

Cross-Modal RetrievalSpeech RecognitionQuestion AnsweringText Retrieval

Similar Papers 제목 키워드 기반

Contrastive Learning for Task-Independent SpeechLLM-Pretraining

2024-12-20 · Maike Züfle, Jan Niehues

Large language models (LLMs) excel in natural language processing but adapting these LLMs to speech processing tasks efficiently is not straightforward. Direct task-specific fine-tuning is limited by overfitting risks, d…

Contrastive LearningQuestion Answering

Injecting Text in Self-Supervised Speech Pretraining

2021-08-27 · Zhehuai Chen, Yu Zhang, Andrew Rosenberg, Bhuvana Ramabhadran 외

Self-supervised pretraining for Automated Speech Recognition (ASR) has shown varied degrees of success. In this paper, we propose to jointly learn representations during pretraining from two different modalities: speech …

Contrastive LearningLanguage Modellingspeech-recognitionSpeech Recognition

GEmo-CLAP: Gender-Attribute-Enhanced Contrastive Language-Audio Pretraining for Accurate Speech Emotion Recognition

2023-06-13 · Yu Pan, Yanni Hu, Yuguang Yang, Wen Fei 외

Contrastive cross-modality pretraining has recently exhibited impressive success in diverse fields, whereas there is limited research on their merits in speech emotion recognition (SER). In this paper, we propose GEmo-CL…

AttributeContrastive LearningEmotion RecognitionMulti-Task Learning+2

CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval

2024-12-17 · Mohammad Mahdi Abootorabi, Ehsaneddin Asgari

This study introduces CLASP (Contrastive Language-Speech Pretraining), a multilingual, multimodal representation tailored for audio-text information retrieval. CLASP leverages the synergy between spoken content and textu…

Contrastive LearningInformation RetrievalMultimodal Deep LearningRetrieval+2

Unsupervised pretraining transfers well across languages

2020-02-07 · Morgane Rivière, Armand Joulin, Pierre-Emmanuel Mazaré, Emmanuel Dupoux

Cross-lingual and multi-lingual training of Automatic Speech Recognition (ASR) has been extensively investigated in the supervised setting. This assumes the existence of a parallel corpus of speech and orthographic trans…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition