paper-with-me

Papers

Text-to-Video: a Two-stage Framework for Zero-shot Identity-agnostic Talking-head Generation

2023-08-12 · Zhichao Wang, Mengyu Dai, Keld Lundgaard

The advent of ChatGPT has introduced innovative methods for information gathering and analysis. However, the information provided by ChatGPT is limited to text, and the visualization of this information remains constrained. Previous research has explored zero-shot text-to-video (TTV) approaches to transform text into videos. However, these methods lacked control over the identity of the generated audio, i.e., not identity-agnostic, hindering their effectiveness. To address this limitation, we propose a novel two-stage framework for person-agnostic video cloning, specifically focusing on TTV generation. In the first stage, we leverage pretrained zero-shot models to achieve text-to-speech (TTS) conversion. In the second stage, an audio-driven talking head generation method is employed to produce compelling videos privided the audio generated in the first stage. This paper presents a comparative analysis of different TTS and audio-driven talking head generation methods, identifying the most promising approach for future research and development. Some audio and videos samples can be found in the following link: https://github.com/ZhichaoWang970201/Text-to-Video/tree/main.

📄 PDF Abstract BibTeX arXiv:2308.06457

Code (1)

zhichaowang970201/text-to-video 공식 구현

Tasks

Talking Head Generationtext-to-speechText to Speech

Similar Papers 제목 키워드 기반

IM-Zero: Instance-level Motion Controllable Video Generation in a Zero-shot Manner

2025-01-01 · CVPR 2025 1 · YuYang Huang, Yabo Chen, Li Ding, Xiaopeng Zhang 외

Controllability of video generation has been recently concerned in addition to the quality of generated videos. The main challenge to controllable video generation is to synthesize videos based on user-specified inst…

Motion GenerationText-to-Video GenerationVideo Generation

X-Aligner: Composed Visual Retrieval without the Bells and Whistles

2026-01-23 · Yuqian Zheng, Mariana-Iuliana Georgescu arxiv

Composed Video Retrieval (CoVR) facilitates video retrieval by combining visual and textual queries. However, existing CoVR frameworks typically fuse multimodal inputs in a single stage, achieving only marginal gains ove…

Zero-shot GeneralizationImage RetrievalVideo Retrieval

VideoPoet: A Large Language Model for Zero-Shot Video Generation

2023-12-21 · Dan Kondratyuk, Lijun Yu, Xiuye Gu, José Lezama 외

We present VideoPoet, a language model capable of synthesizing high-quality video, with matching audio, from a large variety of conditioning signals. VideoPoet employs a decoder-only transformer architecture that process…

DecoderLanguage ModelingLanguage ModellingLarge Language Model+2

Zero-Shot Object Re-Identification in Egocentric Kitchen Videos via Multi-Stage SAM3 Feature Fusion

2026-05-25 · Dmytro Klepachevskyi, Alexander Wong, Sirisha Rambhatla, Yuhao Chen arxiv

Object re-identification (ReID) in egocentric kitchen videos is challenging due to rapid viewpoint changes, frequent occlusions, cluttered scenes, and large intra-class appearance variations. Objects may leave and re-ent…

Reason, Retrieve, Re-rank: A Zero-Shot Reasoning-Aware Framework for Composed Video Retrieval

2026-05-30 · Ali Alavi arxiv

Composed Video Retrieval (CoVR) seeks the target video that results from applying a free-form textual modification to a reference video. We address the \emph{Reason-Aware} CoVR (CoVR-R) challenge at the CVPR~2026 VidLLMs…

Video Retrieval