paper-with-me

Papers

TRAVID: An End-to-End Video Translation Framework

2023-09-20 · Prottay Kumar Adhikary, Bandaru Sugandhi, Subhojit Ghimire, Santanu Pal, Partha Pakray

In today's globalized world, effective communication with people from diverse linguistic backgrounds has become increasingly crucial. While traditional methods of language translation, such as written text or voice-only translations, can accomplish the task, they often fail to capture the complete context and nuanced information conveyed through nonverbal cues like facial expressions and lip movements. In this paper, we present an end-to-end video translation system that not only translates spoken language but also synchronizes the translated speech with the lip movements of the speaker. Our system focuses on translating educational lectures in various Indian languages, and it is designed to be effective even in low-resource system settings. By incorporating lip movements that align with the target language and matching them with the speaker's voice using voice cloning techniques, our application offers an enhanced experience for students and users. This additional feature creates a more immersive and realistic learning environment, ultimately making the learning process more effective and engaging.

📄 PDF Abstract BibTeX arXiv:2309.11338

Code (0)

등록된 구현이 없습니다.

Tasks

TranslationVoice Cloning

Methods 이 논문이 사용한 방법론

fail 설명 없음
ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

UltraVideo: High-Quality UHD Video Dataset with Comprehensive Captions

2025-06-16 · Zhucun Xue, Jiangning Zhang, Teng Hu, Haoyang He 외

The quality of the video dataset (image quality, resolution, and fine-grained caption) greatly influences the performance of the video generation model. The growing demand for video applications sets higher requirements …

4k8kVideo Generation

RobustSora: De-Watermarked Benchmark for Robust AI-Generated Video Detection

2025-12-11 · Zhuo Wang, Xiliang Liu, Ligang Sun arxiv

The proliferation of AI-generated video models poses new challenges to information integrity and digital trust. A key confound, however, remains unaddressed: commercial generators embed visible overlay watermarks for pro…

Shortcut-V2V: Compression Framework for Video-to-Video Translation based on Temporal Redundancy Reduction

2023-08-15 · ICCV 2023 1 · Chaeyeon Chung, Yeojeong Park, Seunghwan Choi, Munkhsoyol Ganbat 외

Video-to-video translation aims to generate video frames of a target domain from an input video. Despite its usefulness, the existing networks require enormous computations, necessitating their model compression for wide…

Computational EfficiencyModel CompressionTranslation

Rerender A Video: Zero-Shot Text-Guided Video-to-Video Translation

2023-06-13 · Shuai Yang, Yifan Zhou, Ziwei Liu, Chen Change Loy

Large text-to-image diffusion models have exhibited impressive proficiency in generating high-quality images. However, when applying these models to video domain, ensuring temporal consistency across video frames remains…

Patch MatchingTranslation

Video-guided Machine Translation with Global Video Context

2026-04-08 · Jian Chen, JinZe Lv, Zi Long, XiangHua Fu arxiv

Video-guided Multimodal Translation (VMT) has advanced significantly in recent years. However, most existing methods rely on locally aligned video segments paired one-to-one with subtitles, limiting their ability to capt…

Machine Translation