paper-with-me

홈 › Papers

Feature-level Incongruence Reduction for Multimodal Translation

2021-06-01 · NAACL (ALVR) 2021 6 · Zhifeng Li, Yu Hong, Yuchen Pan, Jian Tang, Jianmin Yao, Guodong Zhou

Caption translation aims to translate image annotations (captions for short). Recently, Multimodal Neural Machine Translation (MNMT) has been explored as the essential solution. Besides of linguistic features in captions, MNMT allows visual(image) features to be used. The integration of multimodal features reinforces the semantic representation and considerably improves translation performance. However, MNMT suffers from the incongruence between visual and linguistic features. To overcome the problem, we propose to extend MNMT architecture with a harmonization network, which harmonizes multimodal features(linguistic and visual features)by unidirectional modal space conversion. It enables multimodal translation to be carried out in a seemingly monomodal translation pipeline. We experiment on the golden Multi30k-16 and 17. Experimental results show that, compared to the baseline,the proposed method yields the improvements of 2.2% BLEU for the scenario of translating English captions into German (En→De) at best,7.6% for the case of English-to-French translation(En→Fr) and 1.5% for English-to-Czech(En→Cz). The utilization of harmonization network leads to the competitive performance to the-state-of-the-art.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationTranslation

Similar Papers 제목 키워드 기반

Transformer-based Cascaded Multimodal Speech Translation

2019-10-29 · EMNLP (IWSLT) 2019 11 · Zixiu Wu, Ozan Caglayan, Julia Ive, Josiah Wang 외

This paper describes the cascaded multimodal speech translation systems developed by Imperial College London for the IWSLT 2019 evaluation campaign. The architecture consists of an automatic speech recognition (ASR) syst…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine TranslationMultimodal Machine Translation+3

E-THER: A Multimodal Dataset for Empathic AI -- Towards Emotional Mismatch Awareness

2025-09-02 · Sharjeel Tahir, Judith Johnson, Jumana Abu-Khalaf, Syed Afaq Ali Shah arxiv

A prevalent shortfall among current empathic AI systems is their inability to recognize when verbal expressions may not fully reflect underlying emotional states. This is because the existing datasets, used for the train…

Emotion Recognition

Cross-lingual Incongruences in the Annotation of Coreference

2019-06-01 · WS 2019 6 · Ekaterina Lapshinova-Koltunski, Sharid Lo{\'a}iciga, Christian Hardmeier, Pauline Krielke

In the present paper, we deal with incongruences in English-German multilingual coreference annotation and present automated methods to discover them. More specifically, we automatically detect full coreference chains in…

coreference-resolutionCoreference ResolutionTranslation

CF-Net: Conflict Fusion with Speaker Normalisation and Certainty Weighting for Ambivalence/Hesitancy Recognition

2026-07-15 · Tung Hung Bui, Hong Hai Nguyen, Van Thong Huynh arxiv

Detecting ambivalence and hesitancy (AH) in unconstrained video is challenging because the target signal is inherently ambiguous and expressed through subtle cross-modal incongruence rather than prototypical affect. We p…

FoCLIP: A Feature-Space Misalignment Framework for CLIP-Based Image Manipulation and Detection

2025-11-10 · Yulin Chen, Zeyuan Wang, Tianyuan Yu, Yingmei Wei 외 arxiv

The well-aligned attribute of CLIP-based models enables its effective application like CLIPscore as a widely adopted image quality assessment metric. However, such a CLIP-based metric is vulnerable for its delicate multi…

Image Quality AssessmentImage Manipulation