paper-with-me

Papers

AKiRa: Augmentation Kit on Rays for optical video generation

2024-12-18 · CVPR 2025 1 · Xi Wang, Robin Courant, Marc Christie, Vicky Kalogeiton

Recent advances in text-conditioned video diffusion have greatly improved video quality. However, these methods offer limited or sometimes no control to users on camera aspects, including dynamic camera motion, zoom, distorted lens and focus shifts. These motion and optical aspects are crucial for adding controllability and cinematic elements to generation frameworks, ultimately resulting in visual content that draws focus, enhances mood, and guides emotions according to filmmakers' controls. In this paper, we aim to close the gap between controllable video generation and camera optics. To achieve this, we propose AKiRa (Augmentation Kit on Rays), a novel augmentation framework that builds and trains a camera adapter with a complex camera model over an existing video generation backbone. It enables fine-tuned control over camera motion as well as complex optical parameters (focal length, distortion, aperture) to achieve cinematic effects such as zoom, fisheye effect, and bokeh. Extensive experiments demonstrate AKiRa's effectiveness in combining and composing camera optics while outperforming all state-of-the-art methods. This work sets a new landmark in controlled and optically enhanced video generation, paving the way for future optical video generation methods.

📄 PDF Abstract BibTeX arXiv:2412.14158

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
Adapter 설명 없음
Focus 설명 없음

Similar Papers 제목 키워드 기반

Learning Spatio-Temporal Features with Two-Stream Deep 3D CNNs for Lipreading

2019-05-04 · Xinshuo Weng, Kris Kitani

We focus on the word-level visual lipreading, which requires recognizing the word being spoken, given only the video but not the audio. State-of-the-art methods explore the use of end-to-end neural networks, including a …

General ClassificationLipreadingOptical Flow Estimation

VanGogh: A Unified Multimodal Diffusion-based Framework for Video Colorization

2025-01-16 · Zixun Fang, Zhiheng Liu, Kai Zhu, Yu Liu 외

Video colorization aims to transform grayscale videos into vivid color representations while maintaining temporal consistency and structural integrity. Existing video colorization methods often suffer from color bleeding…

ColorizationOptical Flow Estimation

Efficient Proxy Raytracer for Optical Systems using Implicit Neural Representations

2025-07-28 · Shiva Sinaei, Chuanjun Zheng, Kaan Akşit, Daisuke Iwai arxiv

Ray tracing is a widely used technique for modeling optical systems, involving sequential surface-by-surface computations, which can be computationally intensive. We propose Ray2Ray, a novel method that leverages implici…

CLTS-GAN: Color-Lighting-Texture-Specular Reflection Augmentation for Colonoscopy

2022-06-29 · Shawn Mathew, Saad Nadeem, Arie Kaufman

Automated analysis of optical colonoscopy (OC) video frames (to assist endoscopists during OC) is challenging due to variations in color, lighting, texture, and specular reflections. Previous methods either remove some o…

FloVD: Optical Flow Meets Video Diffusion Model for Enhanced Camera-Controlled Video Synthesis

2025-02-12 · CVPR 2025 1 · Wonjoon Jin, Qi Dai, Chong Luo, Seung-Hwan Baek 외

This paper presents FloVD, a novel optical-flow-based video diffusion model for camera-controllable video generation. FloVD leverages optical flow maps to represent motions of the camera and moving objects. This approach…

Motion SynthesisOptical Flow EstimationVideo Generation