paper-with-me

홈 › Papers

Dual-Encoder Architecture with Encoder Selection for Joint Close-Talk and Far-Talk Speech Recognition

2021-09-17 · Felix Weninger, Marco Gaudesi, Ralf Leibold, Roberto Gemello, Puming Zhan

In this paper, we propose a dual-encoder ASR architecture for joint modeling of close-talk (CT) and far-talk (FT) speech, in order to combine the advantages of CT and FT devices for better accuracy. The key idea is to add an encoder selection network to choose the optimal input source (CT or FT) and the corresponding encoder. We use a single-channel encoder for CT speech and a multi-channel encoder with Spatial Filtering neural beamforming for FT speech, which are jointly trained with the encoder selection. We validate our approach on both attention-based and RNN Transducer end-to-end ASR systems. The experiments are done with conversational speech from a medical use case, which is recorded simultaneously with a CT device and a microphone array. Our results show that the proposed dual-encoder architecture obtains up to 9% relative WER reduction when using both CT and FT input, compared to the best single-encoder system trained and tested in matched condition.

📄 PDF Abstract BibTeX arXiv:2109.08744

Code (0)

등록된 구현이 없습니다.

Tasks

speech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

LoopITR: Combining Dual and Cross Encoder Architectures for Image-Text Retrieval

2022-03-10 · Jie Lei, Xinlei Chen, Ning Zhang, Mengjiao Wang 외

Dual encoders and cross encoders have been widely used for image-text retrieval. Between the two, the dual encoder encodes the image and text independently followed by a dot product, while the cross encoder jointly feeds…

Image-text RetrievalRetrievalText Retrieval

Peeling Back the Layers: An In-Depth Evaluation of Encoder Architectures in Neural News Recommenders

2024-10-02 · Andreea Iana, Goran Glavaš, Heiko Paulheim

Encoder architectures play a pivotal role in neural news recommenders by embedding the semantic and contextual information of news and users. Thus, research has heavily focused on enhancing the representational capabilit…

Model SelectionNews RecommendationRecommendation Systems

Contrastive Pretraining with Dual Visual Encoders for Gloss-Free Sign Language Translation

2025-07-14 · Ozge Mercanoglu Sincan, Richard Bowden arxiv

Sign Language Translation (SLT) aims to convert sign language videos into spoken or written text. While early systems relied on gloss annotations as an intermediate supervision, such annotations are costly to obtain and …

Sign Language Translation

Ensembles and Encoders for Task-Free Continual Learning

2021-09-29 · Murray Shanahan, Christos Kaplanis, Jovana Mitrović

We present an architecture that is effective for continual learning in an especially demanding setting, where task boundaries do not exist or are unknown, and where classes have to be learned online (with each presented …

Continual LearningSelf-Supervised Learning

Encoders and Ensembles for Task-Free Continual Learning

2021-05-27 · Murray Shanahan, Christos Kaplanis, Jovana Mitrović

We present an architecture that is effective for continual learning in an especially demanding setting, where task boundaries do not exist or are unknown, and where classes have to be learned online (with each example pr…

Continual LearningImage ClassificationSelf-Supervised Learning