paper-with-me

홈 › Papers

Identifying Speakers in Dialogue Transcripts: A Text-based Approach Using Pretrained Language Models

2024-07-16 · Minh Nguyen, Franck Dernoncourt, Seunghyun Yoon, Hanieh Deilamsalehy, Hao Tan, Ryan Rossi, Quan Hung Tran, Trung Bui, Thien Huu Nguyen

We introduce an approach to identifying speaker names in dialogue transcripts, a crucial task for enhancing content accessibility and searchability in digital media archives. Despite the advancements in speech recognition, the task of text-based speaker identification (SpeakerID) has received limited attention, lacking large-scale, diverse datasets for effective model training. Addressing these gaps, we present a novel, large-scale dataset derived from the MediaSum corpus, encompassing transcripts from a wide range of media sources. We propose novel transformer-based models tailored for SpeakerID, leveraging contextual cues within dialogues to accurately attribute speaker names. Through extensive experiments, our best model achieves a great precision of 80.3\%, setting a new benchmark for SpeakerID. The data and code are publicly available here: \url{https://github.com/adobe-research/speaker-identification}

📄 PDF Abstract BibTeX arXiv:2407.12094

Code (1)

adobe-research/speaker-identification 공식 구현 pytorch

Tasks

AttributeSpeaker Identificationspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Modeling Real-Time Interactive Conversations as Timed Diarized Transcripts

2024-05-21 · Garrett Tanzer, Gustaf Ahdritz, Luke Melas-Kyriazi

Chatbots built upon language models have exploded in popularity, but they have largely been limited to synchronous, turn-by-turn dialogues. In this paper we present a simple yet general method to simulate real-time inter…

Korean Drama Scene Transcript Dataset for Emotion Recognition in Conversations

2022-11-11 · IEEE Access 2022 11 · Sudarshan Pant, Eunchae Lim, Hyung-Jeong Yang, Guee-Sang Lee 외

Understanding emotions in conversation is a challenging task, as the sentences often have an implied meaning that is not generally understood in isolation. Efficient use of contextual information is essential for emotion…

Emotion RecognitionEmotion Recognition in Conversation

Identifying Speakers and Listeners of Quoted Speech in Literary Works

2017-11-01 · IJCNLP 2017 11 · Chak Yan Yeung, John Lee

We present the first study that evaluates both speaker and listener identification for direct speech in literary texts. Our approach consists of two steps: identification of speakers and listeners near the quotes, and di…

SegmentationSpeaker Identification

FriendsQA: Open-Domain Question Answering on TV Show Transcripts

2019-09-01 · WS 2019 9 · Zhengzhe Yang, Jinho D. Choi

This paper presents FriendsQA, a challenging question answering dataset that contains 1,222 dialogues and 10,610 open-domain questions, to tackle machine comprehension on everyday conversations. Each dialogue, involving …

Open-Domain Question AnsweringQuestion AnsweringReading Comprehension

A Hierarchical Network for Abstractive Meeting Summarization with Cross-Domain Pretraining

2020-04-04 · Findings of the Association for Computational Linguistics 2020 · Chenguang Zhu, Ruochen Xu, Michael Zeng, Xuedong Huang

With the abundance of automatic meeting transcripts, meeting summarization is of great interest to both participants and other parties. Traditional methods of summarizing meetings depend on complex multi-step pipelines t…

Abstractive Text SummarizationArticlesMeeting SummarizationText Summarization