Identifying Speakers in Dialogue Transcripts: A Text-based Approach Using Pretrained Language Models
We introduce an approach to identifying speaker names in dialogue transcripts, a crucial task for enhancing content accessibility and searchability in digital media archives. Despite the advancements in speech recognition, the task of text-based speaker identification (SpeakerID) has received limited attention, lacking large-scale, diverse datasets for effective model training. Addressing these gaps, we present a novel, large-scale dataset derived from the MediaSum corpus, encompassing transcripts from a wide range of media sources. We propose novel transformer-based models tailored for SpeakerID, leveraging contextual cues within dialogues to accurately attribute speaker names. Through extensive experiments, our best model achieves a great precision of 80.3\%, setting a new benchmark for SpeakerID. The data and code are publicly available here: \url{https://github.com/adobe-research/speaker-identification}
Code (1)
Tasks
AttributeSpeaker Identificationspeech-recognitionSpeech RecognitionSimilar Papers 제목 키워드 기반
Modeling Real-Time Interactive Conversations as Timed Diarized Transcripts
Chatbots built upon language models have exploded in popularity, but they have largely been limited to synchronous, turn-by-turn dialogues. In this paper we present a simple yet general method to simulate real-time inter…
Korean Drama Scene Transcript Dataset for Emotion Recognition in Conversations
Understanding emotions in conversation is a challenging task, as the sentences often have an implied meaning that is not generally understood in isolation. Efficient use of contextual information is essential for emotion…
Emotion RecognitionEmotion Recognition in ConversationIdentifying Speakers and Listeners of Quoted Speech in Literary Works
We present the first study that evaluates both speaker and listener identification for direct speech in literary texts. Our approach consists of two steps: identification of speakers and listeners near the quotes, and di…
SegmentationSpeaker IdentificationFriendsQA: Open-Domain Question Answering on TV Show Transcripts
This paper presents FriendsQA, a challenging question answering dataset that contains 1,222 dialogues and 10,610 open-domain questions, to tackle machine comprehension on everyday conversations. Each dialogue, involving …
Open-Domain Question AnsweringQuestion AnsweringReading ComprehensionA Hierarchical Network for Abstractive Meeting Summarization with Cross-Domain Pretraining
With the abundance of automatic meeting transcripts, meeting summarization is of great interest to both participants and other parties. Traditional methods of summarizing meetings depend on complex multi-step pipelines t…
Abstractive Text SummarizationArticlesMeeting SummarizationText Summarization