paper-with-me

홈 › Papers

Speaker-adapted neural-network-based fusion for multimodal reference resolution

2019-09-01 · WS 2019 9 · Diana Kleingarn, Nima Nabizadeh, Martin Heckmann, Dorothea Kolossa

Humans use a variety of approaches to reference objects in the external world, including verbal descriptions, hand and head gestures, eye gaze or any combination of them. The amount of useful information from each modality, however, may vary depending on the specific person and on several other factors. For this reason, it is important to learn the correct combination of inputs for inferring the best-fitting reference. In this paper, we investigate appropriate speaker-dependent and independent fusion strategies in a multimodal reference resolution task. We show that without any change in the modality models, only through an optimized fusion technique, it is possible to reduce the error rate of the system on a reference resolution task by more than 50{\%}.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Adapted End-to-End Coreference Resolution System for Anaphoric Identities in Dialogues

2021-09-01 · ACL (CODI, CRAC) 2021 11 · Liyan Xu, Jinho D. Choi

We present an effective system adapted from the end-to-end neural coreference resolution model, targeting on the task of anaphora resolution in dialogues. Three aspects are specifically addressed in our approach, includi…

coreference-resolutionCoreference ResolutionTransfer Learning

StyleFusion TTS: Multimodal Style-control and Enhanced Feature Fusion for Zero-shot Text-to-speech Synthesis

2024-09-24 · Zhiyong Chen, Xinnuo Li, Zhiqi Ai, Shugong Xu

We introduce StyleFusion-TTS, a prompt and/or audio referenced, style and speaker-controllable, zero-shot text-to-speech (TTS) synthesis system designed to enhance the editability and naturalness of current research lite…

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

AMR: Adaptive Modality Routing for Multimodal Polyglot Speaker Identification

2026-06-28 · Chuxiao Zuo, Yao Zhu, Minqiang Xu, Manhong Wang 외 arxiv

Multimodal speaker identification systems face two key challenges in real-world deployment: missing modalities and language mismatch between training and testing conditions. In practical scenarios, background multi-speak…

Speaker Identification

Reference Resolution and Context Change in Multimodal Situated Dialogue for Exploring Data Visualizations

2022-09-06 · Abhinav Kumar, Barbara Di Eugenio, Abari Bhattacharya, Jillian Aurisano 외

Reference resolution, which aims to identify entities being referred to by a speaker, is more complex in real world settings: new referents may be created by processes the agents engage in and/or be salient only because …

Transfer Learning

Online Coreference Resolution for Dialogue Processing: Improving Mention-Linking on Real-Time Conversations

2022-05-21 · *SEM (NAACL) 2022 7 · Liyan Xu, Jinho D. Choi

This paper suggests a direction of coreference resolution for online decoding on actively generated input such as dialogue, where the model accepts an utterance and its past context, then finds mentions in the current ut…

coreference-resolutionCoreference Resolution