paper-with-me

홈 › Papers

Imagination improves Multimodal Translation

2017-05-11 · IJCNLP 2017 11 · Desmond Elliott, Ákos Kádár

We decompose multimodal translation into two sub-tasks: learning to translate and learning visually grounded representations. In a multitask learning framework, translations are learned in an attention-based encoder-decoder, and grounded representations are learned through image representation prediction. Our approach improves translation performance compared to the state of the art on the Multi30K dataset. Furthermore, it is equally effective if we train the image prediction task on the external MS COCO dataset, and we find improvements if we train the translation model on the external News Commentary parallel text.

📄 PDF Abstract BibTeX arXiv:1705.04350

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderPredictionTranslation

Similar Papers 제목 키워드 기반

Generative Imagination Elevates Machine Translation

2020-09-21 · NAACL 2021 4 · Quanyu Long, Mingxuan Wang, Lei LI

There are common semantics shared across text and images. Given a sentence in a source language, whether depicting the visual scene helps translation into a target language? Existing multimodal neural machine translation…

Machine TranslationMultimodal Machine TranslationSentenceTransfer Learning+1

Make Imagination Clearer! Stable Diffusion-based Visual Imagination for Multimodal Machine Translation

2024-12-17 · Andong Chen, Yuchen Song, Kehai Chen, Muyun Yang 외

Visual information has been introduced for enhancing machine translation (MT), and its effectiveness heavily relies on the availability of large amounts of bilingual parallel sentence pairs with manual image annotations.…

Language ModelingLanguage ModellingLarge Language ModelMachine Translation+3

SpatialImaginer: Towards Adaptive Visual Imagination for Spatial Reasoning

2026-04-19 · Yian Li, Yang Jiao, Bin Zhu, Tianwen Qian 외 arxiv

Spatial intelligence, which refers to the ability to reason about geometric and physical structure from visual observations, remains a core challenge for multimodal large language models. Despite promising performance, r…

multimodal generationSpatial Reasoning

Missing Modality Imagination Network for Emotion Recognition with Uncertain Missing Modalities

2021-08-01 · ACL 2021 5 · Jinming Zhao, Ruichen Li, Qin Jin

Multimodal fusion has been proved to improve emotion recognition performance in previous works. However, in real-world applications, we often encounter the problem of missing modality, and which modalities will be missin…

Emotion Recognition

Exploiting modality-invariant feature for robust multimodal emotion recognition with missing modalities

2022-10-27 · Haolin Zuo, Rui Liu, Jinming Zhao, Guanglai Gao 외

Multimodal emotion recognition leverages complementary information across modalities to gain performance. However, we cannot guarantee that the data of all modalities are always present in practice. In the studies to pre…

Emotion RecognitionMultimodal Emotion Recognition