paper-with-me

홈 › Papers

Unsupervised Multimodal Video-to-Video Translation via Self-Supervised Learning

2020-04-14 · Kangning Liu, Shuhang Gu, Andres Romero, Radu Timofte

Existing unsupervised video-to-video translation methods fail to produce translated videos which are frame-wise realistic, semantic information preserving and video-level consistent. In this work, we propose UVIT, a novel unsupervised video-to-video translation model. Our model decomposes the style and the content, uses the specialized encoder-decoder structure and propagates the inter-frame information through bidirectional recurrent neural network (RNN) units. The style-content decomposition mechanism enables us to achieve style consistent video translation results as well as provides us with a good interface for modality flexible translation. In addition, by changing the input frames and style codes incorporated in our translation, we propose a video interpolation loss, which captures temporal information within the sequence to train our building blocks in a self-supervised manner. Our model can produce photo-realistic, spatio-temporal consistent translated videos in a multimodal way. Subjective and objective experimental results validate the superiority of our model over existing methods. More details can be found on our project website: https://uvit.netlify.com

📄 PDF Abstract BibTeX arXiv:2004.06502

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderSelf-Supervised LearningTranslation

Similar Papers 제목 키워드 기반

BigVideo: A Large-scale Video Subtitle Translation Dataset for Multimodal Machine Translation

2023-05-23 · Liyan Kang, Luyang Huang, Ningxin Peng, Peihao Zhu 외

We present a large-scale video subtitle translation dataset, BigVideo, to facilitate the study of multi-modality machine translation. Compared with the widely used How2 and VaTeX datasets, BigVideo is more than 10 times …

Contrastive LearningMachine TranslationMultimodal Machine TranslationNMT+2

Unsupervised Video-to-Video Translation

2018-06-10 · ICLR 2019 5 · Dina Bashkirova, Ben Usman, Kate Saenko

Unsupervised image-to-image translation is a recently proposed task of translating an image to a different style or domain given only unpaired image examples at training time. In this paper, we formulate a new task of un…

Image-to-Image TranslationTranslationUnsupervised Image-To-Image Translation

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey

2026-04-13 · Bingzheng Qu, Kehai Chen, Xuefeng Bai, Min Zhang arxiv

Recent progress in multimodal large language models (MLLMs) is reshaping video translation from a cascaded pipeline of automatic speech recognition, machine translation, text-to-speech, and lip synchronization into a uni…

Multimodal ReasoningMachine TranslationSpeech Recognition

Unsupervised Sign Language Translation and Generation

2024-02-12 · Zhengsheng Guo, Zhiwei He, Wenxiang Jiao, Xing Wang 외

Motivated by the success of unsupervised neural machine translation (UNMT), we introduce an unsupervised sign language translation and generation network (USLNet), which learns from abundant single-modality (text and vid…

Machine TranslationSign Language TranslationTranslation

Keyframe Segmentation and Positional Encoding for Video-guided Machine Translation Challenge 2020

2020-06-23 · Tosho Hirasawa, Zhishen Yang, Mamoru Komachi, Naoaki Okazaki

Video-guided machine translation as one of multimodal neural machine translation tasks targeting on generating high-quality text translation by tangibly engaging both video and text. In this work, we presented our video-…

Machine TranslationTranslationVideo-Guided Machine Translation