paper-with-me

홈 › Papers

Vlogger: Make Your Dream A Vlog

2024-01-17 · CVPR 2024 1 · Shaobin Zhuang, Kunchang Li, Xinyuan Chen, Yaohui Wang, Ziwei Liu, Yu Qiao, Yali Wang

In this work, we present Vlogger, a generic AI system for generating a minute-level video blog (i.e., vlog) of user descriptions. Different from short videos with a few seconds, vlog often contains a complex storyline with diversified scenes, which is challenging for most existing video generation approaches. To break through this bottleneck, our Vlogger smartly leverages Large Language Model (LLM) as Director and decomposes a long video generation task of vlog into four key stages, where we invoke various foundation models to play the critical roles of vlog professionals, including (1) Script, (2) Actor, (3) ShowMaker, and (4) Voicer. With such a design of mimicking human beings, our Vlogger can generate vlogs through explainable cooperation of top-down planning and bottom-up shooting. Moreover, we introduce a novel video diffusion model, ShowMaker, which serves as a videographer in our Vlogger for generating the video snippet of each shooting scene. By incorporating Script and Actor attentively as textual and visual prompts, it can effectively enhance spatial-temporal coherence in the snippet. Besides, we design a concise mixed training paradigm for ShowMaker, boosting its capacity for both T2V generation and prediction. Finally, the extensive experiments show that our method achieves state-of-the-art performance on zero-shot T2V generation and prediction tasks. More importantly, Vlogger can generate over 5-minute vlogs from open-world descriptions, without loss of video coherence on script and actor. The code and model is all available at https://github.com/zhuangshaobin/Vlogger.

📄 PDF Abstract BibTeX arXiv:2401.09414

Code (2)

zhuangshaobin/vlogger 공식 구현 pytorch
vchitect/vlogger pytorch

Tasks

Language ModellingLarge Language ModelVideo Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

A Vlogger-augmented Graph Neural Network Model for Micro-video Recommendation

2024-05-28 · Weijiang Lai, Beihong Jin, Beibei Li, Yiyuan Zheng 외

Existing micro-video recommendation models exploit the interactions between users and micro-videos and/or multi-modal information of micro-videos to predict the next micro-video a user will watch, ignoring the informatio…

Contrastive LearningGraph Neural Network

VLOGGER: Multimodal Diffusion for Embodied Avatar Synthesis

2024-03-13 · CVPR 2025 1 · Enric Corona, Andrei Zanfir, Eduard Gabriel Bazavan, Nikos Kolotouros 외

We propose VLOGGER, a method for audio-driven human video generation from a single input image of a person, which builds on the success of recent generative diffusion models. Our method consists of 1) a stochastic human-…

Face DetectionVideo EditingVideo Generation

Identifying the sentiment styles of YouTube's vloggers

2018-08-29 · EMNLP 2018 10 · Bennett Kleinberg, Maximilian Mozes, Isabelle van der Vegt

Vlogs provide a rich public source of data in a novel setting. This paper examined the continuous sentiment styles employed in 27,333 vlogs using a dynamic intra-textual approach to sentiment analysis. Using unsupervised…

ClusteringSentiment Analysis

VDCook:DIY video data cook your MLLMs

2026-03-04 · Chengwei Wu arxiv

We introduce VDCook: a self-evolving video data operating system, a configurable video data construction platform for researchers and vertical domain teams. Users initiate data requests via natural language queries and a…

Natural Language QueriesScene SegmentationVideo Retrieval

Sentiment Analysis on YouTube Smart Phone Unboxing Video Reviews in Sri Lanka

2023-02-04 · Sherina Sally

Product-related reviews are based on users' experiences that are mostly shared on videos in YouTube. It is the second most popular website globally in 2021. People prefer to watch videos on recently released products pri…

Sentiment Analysis