paper-with-me

홈 › Papers

BERTraffic: BERT-based Joint Speaker Role and Speaker Change Detection for Air Traffic Control Communications

2021-10-12 · Juan Zuluaga-Gomez, Seyyed Saeed Sarfjoo, Amrutha Prasad, Iuliia Nigmatulina, Petr Motlicek, Karel Ondrej, Oliver Ohneiser, Hartmut Helmke

Automatic speech recognition (ASR) allows transcribing the communications between air traffic controllers (ATCOs) and aircraft pilots. The transcriptions are used later to extract ATC named entities, e.g., aircraft callsigns. One common challenge is speech activity detection (SAD) and speaker diarization (SD). In the failure condition, two or more segments remain in the same recording, jeopardizing the overall performance. We propose a system that combines SAD and a BERT model to perform speaker change detection and speaker role detection (SRD) by chunking ASR transcripts, i.e., SD with a defined number of speakers together with SRD. The proposed model is evaluated on real-life public ATC databases. Our BERT SD model baseline reaches up to 10% and 20% token-based Jaccard error rate (JER) in public and private ATC databases. We also achieved relative improvements of 32% and 7.7% in JERs and SD error rate (DER), respectively, compared to VBx, a well-known SD system.

📄 PDF Abstract BibTeX arXiv:2110.05781

Code (2)

idiap/bert-text-diarization-atc 공식 구현
idiap/atco2-corpus

Tasks

Action DetectionActivity DetectionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Change DetectionChunkingspeaker-diarizationSpeaker Diarizationspeech-recognitionSpeech Recognition

Methods 이 논문이 사용한 방법론

Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
WordPiece 설명 없음
Weight Decay 설명 없음
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Attention Dropout Attention Dropout is a type of dropout used in attention-based architectures, where elements are randomly dropped out of the…

Similar Papers 제목 키워드 기반

Joint ASR and Speaker Role Tagging with Serialized Output Training

2025-06-12 · Anfeng Xu, Tiantian Feng, Shrikanth Narayanan

Automatic Speech Recognition systems have made significant progress with large-scale pre-trained models. However, most current systems focus solely on transcribing the speech without identifying speaker roles, a function…

Automatic Speech Recognitionspeech-recognitionSpeech Recognition

Joint Training or Not: An Exploration of Pre-trained Speech Models in Audio-Visual Speaker Diarization

2023-12-07 · Huan Zhao, Li Zhang, Yue Li, Yannan Wang 외

The scarcity of labeled audio-visual datasets is a constraint for training superior audio-visual speaker diarization systems. To improve the performance of audio-visual speaker diarization, we leverage pre-trained superv…

Decoderspeaker-diarizationSpeaker Diarization

Addressee and Response Selection in Multi-Party Conversations with Speaker Interaction RNNs

2017-09-12 · Rui Zhang, Honglak Lee, Lazaros Polymenakos, Dragomir Radev

In this paper, we study the problem of addressee and response selection in multi-party conversations. Understanding multi-party conversations is challenging because of complex speaker interactions: multiple speakers exch…

Conversational Response Selection

Speaker-Guided Encoder-Decoder Framework for Emotion Recognition in Conversation

2022-06-07 · Yinan Bao, Qianwen Ma, Lingwei Wei, Wei Zhou 외

The emotion recognition in conversation (ERC) task aims to predict the emotion label of an utterance in a conversation. Since the dependencies between speakers are complex and dynamic, which consist of intra- and inter-s…

DecoderEmotion RecognitionEmotion Recognition in Conversation

Expressive Voice Conversion: A Joint Framework for Speaker Identity and Emotional Style Transfer

2021-07-08 · Zongyang Du, Berrak Sisman, Kun Zhou, Haizhou Li

Traditional voice conversion(VC) has been focused on speaker identity conversion for speech with a neutral expression. We note that emotional expression plays an essential role in daily communication, and the emotional s…

Emotion RecognitionSpeech Emotion RecognitionStyle TransferVoice Conversion