What does the sea say to the shore? A BERT based DST style approach for speaker to dialogue attribution in novels
We present a complete pipeline to extract characters in a novel and link them to their direct-speech utterances. Our model is divided into three independent components: extracting direct-speech, compiling a list of characters, and attributing those characters to their utterances. Although we find that existing systems can perform the first two tasks accurately, attributing characters to direct speech is a challenging problem due to the narrator’s lack of explicit character mentions, and the frequent use of nominal and pronominal coreference when such explicit mentions are made. We adapt the progress made on Dialogue State Tracking to tackle a new problem: attributing speakers to dialogues. This is the first application of deep learning to speaker attribution, and it shows that is possible to overcome the need for the hand-crafted features and rules used in the past. Our full pipeline improves the performance of state-of-the-art models by a relative 50% in F1-score.
Code (0)
등록된 구현이 없습니다.
Tasks
Dialogue State TrackingSimilar Papers 제목 키워드 기반
Style Transfer for Co-Speech Gesture Animation: A Multi-Speaker Conditional-Mixture Approach
How can we teach robots or virtual assistants to gesture naturally? Can we go further and adapt the gesturing style to follow a specific speaker? Gestures that are naturally timed with corresponding speech during human c…
Gesture GenerationStyle TransferCross-speaker Style Transfer with Prosody Bottleneck in Neural Speech Synthesis
Cross-speaker style transfer is crucial to the applications of multi-style and expressive speech synthesis at scale. It does not require the target speakers to be experts in expressing all styles and to collect correspon…
Expressive Speech SynthesisSpeech SynthesisStyle Transfertext-to-speech+1Probing BERT’s priors with serial reproduction chains
We can learn as much about language models from what they say as we learn from their performance on targeted benchmarks. Sampling is a promising bottom-up method for probing, but generating samples from successful models…
Language ModelingLanguage ModellingMasked Language ModelingBERTphone: Phonetically-Aware Encoder Representations for Utterance-Level Speaker and Language Recognition
We introduce BERTphone, a Transformer encoder trained on large speech corpora that outputs phonetically-aware contextual representation vectors that can be used for both speaker and language recognition. This is accompli…
AvgRepresentation LearningSpeaker RecognitionSpeech RecognitionMany-to-Many Voice Conversion with Out-of-Dataset Speaker Support
We present a Cycle-GAN based many-to-many voice conversion method that can convert between speakers that are not in the training set. This property is enabled through speaker embeddings generated by a neural network that…
Speaker IdentificationVoice Conversion