paper-with-me

Papers

Probabilistic Video Generation using Holistic Attribute Control

2018-03-21 · ECCV 2018 9 · Jiawei He, Andreas Lehrmann, Joseph Marino, Greg Mori, Leonid Sigal

Videos express highly structured spatio-temporal patterns of visual data. A video can be thought of as being governed by two factors: (i) temporally invariant (e.g., person identity), or slowly varying (e.g., activity), attribute-induced appearance, encoding the persistent content of each frame, and (ii) an inter-frame motion or scene dynamics (e.g., encoding evolution of the person ex-ecuting the action). Based on this intuition, we propose a generative framework for video generation and future prediction. The proposed framework generates a video (short clip) by decoding samples sequentially drawn from a latent space distribution into full video frames. Variational Autoencoders (VAEs) are used as a means of encoding/decoding frames into/from the latent space and RNN as a wayto model the dynamics in the latent space. We improve the video generation consistency through temporally-conditional sampling and quality by structuring the latent space with attribute controls; ensuring that attributes can be both inferred and conditioned on during learning/generation. As a result, given attributes and/orthe first frame, our model is able to generate diverse but highly consistent sets ofvideo sequences, accounting for the inherent uncertainty in the prediction task. Experimental results on Chair CAD, Weizmann Human Action, and MIT-Flickr datasets, along with detailed comparison to the state-of-the-art, verify effectiveness of the framework.

📄 PDF Abstract BibTeX arXiv:1803.08085

Code (1)

charlescheng0117/pytorch-VideoVAE pytorch

Tasks

AttributeFuture predictionVideo Generation

Similar Papers 제목 키워드 기반

Toward Controlled Generation of Text

2017-03-02 · ICML 2017 8 · Zhiting Hu, Zichao Yang, Xiaodan Liang, Ruslan Salakhutdinov 외

Generic generation and manipulation of text is challenging and has limited success compared to recent deep generative modeling in visual domain. This paper aims at generating plausible natural language sentences, whose a…

AttributeSentence

TRACE Back from the Future: A Probabilistic Reasoning Approach to Controllable Language Generation

2025-04-25 · Gwen Yidou Weng, Benjie Wang, Guy Van Den Broeck

As large language models (LMs) advance, there is an increasing need to control their outputs to align with human values (e.g., detoxification) or desired attributes (e.g., personalization, topic). However, autoregressive…

AttributeText Generation

Towards Accurate Emotion-Attributed Video Captioning via Fine-grained Emotion-Cause Pair Extraction

2026-06-07 · Weidong Chen, Cheng Ye, Zhendong Mao, Liping Wang 외 arxiv

Emotional Video Captioning (EVC) is a challenging task that aims to generate factually accurate and emotionally rich descriptions for videos. Existing EVC methods leverage holistic visual features to mine global emotiona…

Emotion-Cause Pair ExtractionVideo Captioning

Archon: A Unified Multimodal Model for Holistic Digital Human Generation

2026-05-28 · Chong Bao, Shichen Liu, Lijun Yu, David Futschik 외 arxiv

Digital humans are fundamental to immersive interaction, yet creating a unified model for holistic modalities, including text, audio, motion, and visual content, remains an open challenge. In this paper, we present Archo…

Determining the best attributes for surveillance video keywords generation

2016-02-21 · Liangchen Liu, Arnold Wiliem, Shaokang Chen, Kun Zhao 외

Automatic video keyword generation is one of the key ingredients in reducing the burden of security officers in analyzing surveillance videos. Keywords or attributes are generally chosen manually based on expert knowledg…

Attribute