paper-with-me

Papers

TokenDial: Continuous Attribute Control for Text-to-Video Generation in Visual Dial Space

2026-03-29 · Zhixuan Liu, Peter Schaldenbrand, Yijun Li, Long Mai, Aniruddha Mahapatra, Cusuh Ham, Jean Oh, Jui-Hsien Wang arxiv

In video diffusion transformers, visual patch tokens maintain explicit correspondence to space and time. We hypothesize that their channel dimension can serve as a semantic control space, which we call Visual Dial Space V+. In this space, additive directions can be broadcast to the token stream to control appearance or motion attributes, enabling slider-style edits such as making a generated person look older or run faster. To verify the hypothesis, we present TokenDial, a framework for learning attribute directions in the proposed V+ space. TokenDial keeps the pretrained video generator frozen and optimizes only additive directions. Rather than requiring paired edited videos, TokenDial supervises each direction through its induced effect on generated videos: the edited video should move along the desired attribute while the remaining content stays stable. The learned directions become reusable visual dials that support continuous appearance and motion control, explicit spatiotemporal localization, composition, and reuse across prompts, resolutions, and video lengths. Experiments and human studies show that TokenDial achieves stronger slider controllability and better content preservation than prior video editing and slider-based methods.

📄 PDF Abstract BibTeX arXiv:2603.27520

Code (0)

등록된 구현이 없습니다.

Tasks

Text-to-Video Generation

Similar Papers 제목 키워드 기반

Text Slider: Efficient and Plug-and-Play Continuous Concept Control for Image/Video Synthesis via LoRA Adapters

2025-09-23 · Pin-Yen Chiu, I-Sheng Fang, Jun-Cheng Chen arxiv

Recent advances in diffusion models have significantly improved image and video synthesis. In addition, several concept control methods have been proposed to enable fine-grained, continuous, and flexible control over fre…

Continuous Control

Learning Continuous 3D Words for Text-to-Image Generation

2024-02-13 · CVPR 2024 1 · Ta-Ying Cheng, Matheus Gadelha, Thibault Groueix, Matthew Fisher 외

Current controls over diffusion models (e.g., through text or ControlNet) for image generation fall short in recognizing abstract, continuous attributes like illumination direction or non-rigid shape change. In this pape…

Image GenerationText to Image GenerationText-to-Image Generation

Continuously Controllable Facial Expression Editing in Talking Face Videos

2022-09-17 · Zhiyao Sun, Yu-Hui Wen, Tian Lv, Yanan sun 외

Recently audio-driven talking face video generation has attracted considerable attention. However, very few researches address the issue of emotional editing of these talking face videos with continuously controllable ex…

Image-to-Image TranslationVideo Generation

PERSE: Personalized 3D Generative Avatars from A Single Portrait

2024-12-30 · CVPR 2025 1 · Hyunsoo Cha, Inhee Lee, Hanbyul Joo

We present PERSE, a method for building an animatable personalized generative avatar from a reference portrait. Our avatar model enables facial attribute editing in a continuous and disentangled latent space to control e…

Attribute

Continuous, Subject-Specific Attribute Control in T2I Models by Identifying Semantic Directions

2024-03-25 · CVPR 2025 1 · Stefan Andreas Baumann, Felix Krause, Michael Neumayr, Nick Stracke 외

In recent years, advances in text-to-image (T2I) diffusion models have substantially elevated the quality of their generated images. However, achieving fine-grained control over attributes remains a challenge due to the …

Attribute