Cross-Lingual Speaker Identification Using Distant Supervision
Speaker identification, determining which character said each utterance in literary text, benefits many downstream tasks. Most existing approaches use expert-defined rules or rule-based features to directly approach this task, but these approaches come with significant drawbacks, such as lack of contextual reasoning and poor cross-lingual generalization. In this work, we propose a speaker identification framework that addresses these issues. We first extract large-scale distant supervision signals in English via general-purpose tools and heuristics, and then apply these weakly-labeled instances with a focus on encouraging contextual reasoning to train a cross-lingual language model. We show that the resulting model outperforms previous state-of-the-art methods on two English speaker identification benchmarks by up to 9% in accuracy and 5% with only distant supervision, as well as two Chinese speaker identification datasets by up to 4.7%.
Code (1)
Tasks
Language ModelingLanguage ModellingSpeaker IdentificationSimilar Papers 제목 키워드 기반
Cross-Lingual Speaker Identification from Weak Local Evidence
Speaker identification, determining which character said each utterance in text, benefits many downstream tasks. Most existing approaches use expert-defined rules or rule-based features to directly approach this task, bu…
Language ModelingLanguage ModellingSpeaker IdentificationJoint Representation Learning of Cross-lingual Words and Entities via Attentive Distant Supervision
Joint representation learning of words and entities benefits many NLP tasks, but has not been well explored in cross-lingual settings. In this paper, we propose a novel method for joint representation learning of cross-l…
Cross-Lingual Entity LinkingEntity LinkingRepresentation LearningTranslation+1Low-resource Cross-lingual Event Type Detection via Distant Supervision with Minimal Effort
The RELX Dataset and Matching the Multilingual Blanks for Cross-Lingual Relation Classification
Relation classification is one of the key topics in information extraction, which can be used to construct knowledge bases or to provide useful information for question answering. Current approaches for relation classifi…
ClassificationGeneral ClassificationQuestion AnsweringRelation+1The RELX Dataset and Matching the Multilingual Blanks for Cross-Lingual Relation Classification
Relation classification is one of the key topics in information extraction, which can be used to construct knowledge bases or to provide useful information for question answering. Current approaches for relation classifi…
ClassificationQuestion AnsweringRelationRelation Classification