paper-with-me

Papers

Noise-robust Cross-modal Interactive Learning with Text2Image Mask for Multi-modal Neural Machine Translation

2022-10-01 · COLING 2022 10 · Junjie Ye, Junjun Guo, Yan Xiang, Kaiwen Tan, Zhengtao Yu

Multi-modal neural machine translation (MNMT) aims to improve textual level machine translation performance in the presence of text-related images. Most of the previous works on MNMT focus on multi-modal fusion methods with full visual features. However, text and its corresponding image may not match exactly, visual noise is generally inevitable. The irrelevant image regions may mislead or distract the textual attention and cause model performance degradation. This paper proposes a noise-robust multi-modal interactive fusion approach with cross-modal relation-aware mask mechanism for MNMT. A text-image relation-aware attention module is constructed through the cross-modal interaction mask mechanism, and visual features are extracted based on the text-image interaction mask knowledge. Then a noise-robust multi-modal adaptive fusion approach is presented by fusion the relevant visual and textual features for machine translation. We validate our method on the Multi30K dataset. The experimental results show the superiority of our proposed model, and achieve the state-of-the-art scores in all En-De, En-Fr and En-Cs translation tasks.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

de-enMachine TranslationRelationTranslation

Similar Papers 제목 키워드 기반

ADaFuSE: Adaptive Diffusion-generated Image and Text Fusion for Interactive Text-to-Image Retrieval

2026-03-23 · Zhuocheng Zhang, Xingwu Zhang, Kangheng Liang, Guanxuan Li 외 arxiv

Recent advances in interactive text-to-image retrieval (I-TIR) use diffusion models to bridge the modality gap between the textual information need and the images to be searched, resulting in increased effectiveness. How…

Image Retrieval

Interactive Context-Aware Network for RGB-T Salient Object Detection

2022-11-11 · Yuxuan Wang, Feng Dong, Jinchao Zhu

Salient object detection (SOD) focuses on distinguishing the most conspicuous objects in the scene. However, most related works are based on RGB images, which lose massive useful information. Accordingly, with the maturi…

object-detectionObject DetectionRGB-T Salient Object DetectionSalient Object Detection

Dual Pseudo-Labels Interactive Self-Training for Semi-Supervised Visible-Infrared Person Re-Identification

2023-01-01 · ICCV 2023 1 · Jiangming Shi, Yachao Zhang, Xiangbo Yin, Yuan Xie 외

Visible-infrared person re-identification (VI-ReID) aims to match a specific person from a gallery of images captured from non-overlapping visible and infrared cameras. Most works focus on fully supervised VI-ReID, w…

Person Re-IdentificationPseudo Label

Knowledge Perceived Multi-modal Pretraining in E-commerce

2021-08-20 · Yushan Zhu, Huaixiao Tou, Wen Zhang, Ganqiang Ye 외

In this paper, we address multi-modal pretraining of product data in the field of E-commerce. Current multi-modal pretraining methods proposed for image and text modalities lack robustness in the face of modality-missing…

Language ModelingLanguage ModellingLink PredictionMasked Language Modeling

Text-DiFuse: An Interactive Multi-Modal Image Fusion Framework based on Text-modulated Diffusion Model

2024-10-31 · Hao Zhang, Lei Cao, Jiayi Ma

Existing multi-modal image fusion methods fail to address the compound degradations presented in source images, resulting in fusion images plagued by noise, color bias, improper exposure, \textit{etc}. Additionally, thes…

Semantic SegmentationSpecificity