Speaker-adapted neural-network-based fusion for multimodal reference resolution
Humans use a variety of approaches to reference objects in the external world, including verbal descriptions, hand and head gestures, eye gaze or any combination of them. The amount of useful information from each modality, however, may vary depending on the specific person and on several other factors. For this reason, it is important to learn the correct combination of inputs for inferring the best-fitting reference. In this paper, we investigate appropriate speaker-dependent and independent fusion strategies in a multimodal reference resolution task. We show that without any change in the modality models, only through an optimized fusion technique, it is possible to reduce the error rate of the system on a reference resolution task by more than 50{\%}.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Adapted End-to-End Coreference Resolution System for Anaphoric Identities in Dialogues
We present an effective system adapted from the end-to-end neural coreference resolution model, targeting on the task of anaphora resolution in dialogues. Three aspects are specifically addressed in our approach, includi…
coreference-resolutionCoreference ResolutionTransfer LearningStyleFusion TTS: Multimodal Style-control and Enhanced Feature Fusion for Zero-shot Text-to-speech Synthesis
We introduce StyleFusion-TTS, a prompt and/or audio referenced, style and speaker-controllable, zero-shot text-to-speech (TTS) synthesis system designed to enhance the editability and naturalness of current research lite…
Speech Synthesistext-to-speechText to SpeechText-To-Speech SynthesisAMR: Adaptive Modality Routing for Multimodal Polyglot Speaker Identification
Multimodal speaker identification systems face two key challenges in real-world deployment: missing modalities and language mismatch between training and testing conditions. In practical scenarios, background multi-speak…
Speaker IdentificationReference Resolution and Context Change in Multimodal Situated Dialogue for Exploring Data Visualizations
Reference resolution, which aims to identify entities being referred to by a speaker, is more complex in real world settings: new referents may be created by processes the agents engage in and/or be salient only because …
Transfer LearningOnline Coreference Resolution for Dialogue Processing: Improving Mention-Linking on Real-Time Conversations
This paper suggests a direction of coreference resolution for online decoding on actively generated input such as dialogue, where the model accepts an utterance and its past context, then finds mentions in the current ut…
coreference-resolutionCoreference Resolution