paper-with-me

Papers

Cross-modal Language Generation using Pivot Stabilization for Web-scale Language Coverage

2020-05-01 · ACL 2020 6 · Ashish V. Thapliyal, Radu Soricut

Cross-modal language generation tasks such as image captioning are directly hurt in their ability to support non-English languages by the trend of data-hungry models combined with the lack of non-English annotations. We investigate potential solutions for combining existing language-generation annotations in English with translation capabilities in order to create solutions at web-scale in both domain and language coverage. We describe an approach called Pivot-Language Generation Stabilization (PLuGS), which leverages directly at training time both existing English annotations (gold data) as well as their machine-translated versions (silver data); at run-time, it generates first an English caption and then a corresponding target-language caption. We show that PLuGS models outperform other candidate solutions in evaluations performed over 5 different target languages, under a large-domain testset using images from the Open Images dataset. Furthermore, we find an interesting effect where the English captions generated by the PLuGS models are better than the captions generated by the original, monolingual English model.

📄 PDF Abstract BibTeX arXiv:2005.00246

Code (0)

등록된 구현이 없습니다.

Tasks

Image CaptioningText GenerationTranslation

Similar Papers 제목 키워드 기반

Large Multilingual Models Pivot Zero-Shot Multimodal Learning across Languages

2023-08-23 · Jinyi Hu, Yuan YAO, Chongyi Wang, Shan Wang 외

Recently there has been a significant surge in multimodal learning in terms of both image-to-text and text-to-image generation. However, the success is typically limited to English, leaving other languages largely behind…

Image GenerationImage to textLanguage ModelingLanguage Modelling+3

OpenOmni: Large Language Models Pivot Zero-shot Omnimodal Alignment across Language with Real-time Self-Aware Emotional Speech Synthesis

2025-01-08 · Run Luo, Ting-En Lin, Haonan Zhang, Yuchuan Wu 외

Recent advancements in omnimodal learning have been achieved in understanding and generation across images, text, and speech, though mainly within proprietary models. Limited omnimodal datasets and the inherent challenge…

DecoderEmotional Speech SynthesisLanguage ModelingLanguage Modelling+2

Aligning Source Visual and Target Language Domains for Unpaired Video Captioning

2022-11-22 · Fenglin Liu, Xian Wu, Chenyu You, Shen Ge 외

Training supervised video captioning model requires coupled video-caption pairs. However, for many targeted languages, sufficient paired data are not available. To this end, we introduce the unpaired video captioning tas…

TranslationVideo Captioning

MiniGPT-5: Interleaved Vision-and-Language Generation via Generative Vokens

2023-10-03 · Kaizhi Zheng, Xuehai He, Xin Eric Wang

The effectiveness of Multimodal Large Language Models (MLLMs) demonstrates a profound capability in multimodal understanding. However, the simultaneous generation of images with coherent texts is still underdeveloped. Ad…

Image Generationmultimodal generationReading ComprehensionText Generation

Cross2StrA: Unpaired Cross-lingual Image Captioning with Cross-lingual Cross-modal Structure-pivoted Alignment

2023-05-20 · Shengqiong Wu, Hao Fei, Wei Ji, Tat-Seng Chua

Unpaired cross-lingual image captioning has long suffered from irrelevancy and disfluency issues, due to the inconsistencies of the semantic scene and syntax attributes during transfer. In this work, we propose to addres…

Image CaptioningTranslation