Evaluating the Effects of Embedding with Speaker Identity Information in Dialogue Summarization
Automatic dialogue summarization is a task used to succinctly summarize a dialogue transcript while correctly linking the speakers and their speech, which distinguishes this task from a conventional document summarization. To address this issue and reduce the “who said what”-related errors in a summary, we propose embedding the speaker identity information in the input embedding into the dialogue transcript encoder. Unlike the speaker embedding proposed by Gu et al. (2020), our proposal takes into account the informativeness of position embedding. By experimentally comparing several embedding methods, we confirmed that the scores of ROUGE and a human evaluation of the generated summaries were substantially increased by embedding speaker information at the less informative part of the fixed position embedding with sinusoidal functions.
Code (0)
등록된 구현이 없습니다.
Tasks
Document SummarizationInformativenessPositionSimilar Papers 제목 키워드 기반
Evaluating Identity Leakage in Speaker De-Identification Systems
Speaker de-identification aims to conceal a speaker's identity while preserving intelligibility of the underlying speech. We introduce a benchmark that quantifies residual identity leakage with three complementary error …
An analysis on the effects of speaker embedding choice in non auto-regressive TTS
In this paper we introduce a first attempt on understanding how a non-autoregressive factorised multi-speaker speech synthesis architecture exploits the information present in different speaker embedding sets. We analyse…
Representation LearningSpeech SynthesisSeparating Content from Speaker Identity in Speech for the Assessment of Cognitive Impairments
Deep speaker embeddings have been shown effective for assessing cognitive impairments aside from their original purpose of speaker verification. However, the research found that speaker embeddings encode speaker identity…
Speaker VerificationVoice ConversionTowards Low-Latency Tracking of Multiple Speakers With Short-Context Speaker Embeddings
Speaker embeddings are promising identity-related features that can enhance the identity assignment performance of a tracking system by leveraging its spatial predictions, i.e, by performing identity reassignment. Common…
Knowledge DistillationMIRNet: Learning multiple identities representations in overlapped speech
Many approaches can derive information about a single speaker's identity from the speech by learning to recognize consistent characteristics of acoustic parameters. However, it is challenging to determine identity inform…
Rgb-T TrackingSpeaker VerificationSpeech Separation