paper-with-me

Papers

Sync-DRAW: Automatic Video Generation using Deep Recurrent Attentive Architectures

2016-11-30 · Gaurav Mittal, Tanya Marwah, Vineeth N. Balasubramanian

This paper introduces a novel approach for generating videos called Synchronized Deep Recurrent Attentive Writer (Sync-DRAW). Sync-DRAW can also perform text-to-video generation which, to the best of our knowledge, makes it the first approach of its kind. It combines a Variational Autoencoder~(VAE) with a Recurrent Attention Mechanism in a novel manner to create a temporally dependent sequence of frames that are gradually formed over time. The recurrent attention mechanism in Sync-DRAW attends to each individual frame of the video in sychronization, while the VAE learns a latent distribution for the entire video at the global level. Our experiments with Bouncing MNIST, KTH and UCF-101 suggest that Sync-DRAW is efficient in learning the spatial and temporal information of the videos and generates frames with high structural integrity, and can generate videos from simple captions on these datasets. (Accepted as oral paper in ACM-Multimedia 2017)

📄 PDF Abstract BibTeX arXiv:1611.10314

Code (1)

Singularity42/Sync-DRAW 공식 구현 tf

Tasks

Text-to-Video GenerationVideo Generation

Methods 이 논문이 사용한 방법론

USD Coin Customer Service Number +1-833-534-1729 설명 없음

Similar Papers 제목 키워드 기반

Talking Face Generation by Conditional Recurrent Adversarial Network

2018-04-13 · Yang Song, Jingwen Zhu, Dawei Li, Xiaolong Wang 외

Given an arbitrary face image and an arbitrary speech clip, the proposed work attempts to generating the talking face video with accurate lip synchronization while maintaining smooth transition of both lip and facial mov…

Constrained Lip-synchronizationFace GenerationTalking Face GenerationVideo Generation

Video-based Music Generation

2026-02-05 · Serkan Sulun arxiv

As the volume of video content on the internet grows rapidly, finding a suitable soundtrack remains a significant challenge. This thesis presents EMSYNC (EMotion and SYNChronization), a fast, free, and automatic solution…

Emotion ClassificationMusic Generation

Speech-Synchronized Whiteboard Generation via VLM-Driven Structured Drawing Representations

2026-03-26 · Suraj Prasad, Pinak Mahapatra arxiv

Creating whiteboard-style educational videos demands precise coordination between freehand illustrations and spoken narration, yet no existing method addresses this multimodal synchronization problem with structured, rep…

Learning Asynchronous and Sparse Human-Object Interaction in Videos

2021-03-03 · CVPR 2021 1 · Romero Morais, Vuong Le, Svetha Venkatesh, Truyen Tran

Human activities can be learned from video. With effective modeling it is possible to discover not only the action labels but also the temporal structures of the activities such as the progression of the sub-activities. …

Human-Object Interaction DetectionObject

AnyoneNet: Synchronized Speech and Talking Head Generation for Arbitrary Person

2021-08-09 · Xinsheng Wang, Qicong Xie, Jihua Zhu, Lei Xie 외

Automatically generating videos in which synthesized speech is synchronized with lip movements in a talking head has great potential in many human-computer interaction scenarios. In this paper, we present an automatic me…

Talking Head Generationtext-to-speechText to Speech