paper-with-me

Papers

Activation Steering of Video Generation Models via Reduced-Order Linear Optimal Control

2026-06-03 · Jihoon Hong, Alice Chan, Qiyue Dai, Julian Skifstad, Glen Chou arxiv

Text-to-video (T2V) models trained on large-scale web data can generate undesired content, motivating interventions that reduce harmful outputs without sacrificing visual quality. Activation steering offers an attractive mechanistic alternative to finetuning and prompt filtering, but existing T2V steering methods remain limited, typically applying coarse, non-anticipative interventions that can lead to oversteering and content degradation. To close this gap, we propose Latent Activation Linear-Quadratic Regulator (LA-LQR), a reduced-order optimal control framework for minimally invasive T2V steering. LA-LQR formulates T2V inference as a dynamical system and computes closed-loop feedback interventions that steer activations toward desired feature setpoints while penalizing unnecessary perturbations. To make optimal control feasible for high-dimensional video activations, we project activations onto a low-dimensional, task-relevant subspace derived from contrastive prompt pairs, estimate local linear dynamics in this latent space, and solve a latent LQR problem to obtain timestep- and layer-specific steering signals. We provide theoretical bounds relating latent setpoint tracking to raw activation-space feature control, and empirically validate the fidelity of the reduced latent dynamics. On concept steering and video safety benchmarks, LA-LQR reduces unsafe generations relative to baselines, while preserving prompt fidelity and visual quality.

📄 PDF Abstract BibTeX arXiv:2606.04775

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

Steering Robustness into World Action Models via Mechanistic Interpretability and Optimal Control

2026-07-16 · Jihoon Hong, Julian Skifstad, Qiyue Dai, Alice Chan 외 arxiv

World Action Models (WAMs) enable semantically- and physically-informed control but are brittle under distribution shift. In this work, we use mechanistic interpretability to study how robustness-relevant perturbations a…

Pulling The REINS: Training-Free Safety Alignment of Video Diffusion Models via Representation Steering

2026-06-15 · Rohit Kundu, Arindam Dutta, Sarosij Bose, Athula Balachandran 외 arxiv

Open-weight video diffusion models can generate photorealistic unsafe content, from violence to misinformation, yet existing defenses either require expensive safety fine-tuning that degrades general capability, or apply…

Video Generation

Mechanistic Knobs in LLMs: Retrieving and Steering High-Order Semantic Features via Sparse Autoencoders

2026-01-06 · Ruikang Zhang, Shuo Wang, Qi Su arxiv

Recent work in Mechanistic Interpretability (MI) has enabled the identification and intervention of internal features in Large Language Models (LLMs). However, a persistent challenge lies in linking such internal feature…

Multi-property Steering of Large Language Models with Dynamic Activation Composition

2024-06-25 · Daniel Scalena, Gabriele Sarti, Malvina Nissim

Activation steering methods were shown to be effective in conditioning language model generation by additively intervening over models' intermediate representations. However, the evaluation of these techniques has so far…

Language ModelingLanguage Modelling

Steering Video Diffusion Transformers with Massive Activations

2026-03-18 · Xianhang Cheng, Yujian Zheng, Zhenyu Xie, Tingting Liao 외 arxiv

Despite rapid progress in video diffusion transformers, how their internal model signals can be leveraged with minimal overhead to enhance video generation quality remains underexplored. In this work, we study the role o…

Video Generation