TokenDial: Continuous Attribute Control for Text-to-Video Generation in Visual Dial Space
In video diffusion transformers, visual patch tokens maintain explicit correspondence to space and time. We hypothesize that their channel dimension can serve as a semantic control space, which we call Visual Dial Space V+. In this space, additive directions can be broadcast to the token stream to control appearance or motion attributes, enabling slider-style edits such as making a generated person look older or run faster. To verify the hypothesis, we present TokenDial, a framework for learning attribute directions in the proposed V+ space. TokenDial keeps the pretrained video generator frozen and optimizes only additive directions. Rather than requiring paired edited videos, TokenDial supervises each direction through its induced effect on generated videos: the edited video should move along the desired attribute while the remaining content stays stable. The learned directions become reusable visual dials that support continuous appearance and motion control, explicit spatiotemporal localization, composition, and reuse across prompts, resolutions, and video lengths. Experiments and human studies show that TokenDial achieves stronger slider controllability and better content preservation than prior video editing and slider-based methods.
Code (0)
등록된 구현이 없습니다.
Tasks
Text-to-Video GenerationSimilar Papers 제목 키워드 기반
Text Slider: Efficient and Plug-and-Play Continuous Concept Control for Image/Video Synthesis via LoRA Adapters
Recent advances in diffusion models have significantly improved image and video synthesis. In addition, several concept control methods have been proposed to enable fine-grained, continuous, and flexible control over fre…
Continuous ControlLearning Continuous 3D Words for Text-to-Image Generation
Current controls over diffusion models (e.g., through text or ControlNet) for image generation fall short in recognizing abstract, continuous attributes like illumination direction or non-rigid shape change. In this pape…
Image GenerationText to Image GenerationText-to-Image GenerationContinuously Controllable Facial Expression Editing in Talking Face Videos
Recently audio-driven talking face video generation has attracted considerable attention. However, very few researches address the issue of emotional editing of these talking face videos with continuously controllable ex…
Image-to-Image TranslationVideo GenerationPERSE: Personalized 3D Generative Avatars from A Single Portrait
We present PERSE, a method for building an animatable personalized generative avatar from a reference portrait. Our avatar model enables facial attribute editing in a continuous and disentangled latent space to control e…
AttributeContinuous, Subject-Specific Attribute Control in T2I Models by Identifying Semantic Directions
In recent years, advances in text-to-image (T2I) diffusion models have substantially elevated the quality of their generated images. However, achieving fine-grained control over attributes remains a challenge due to the …
Attribute