paper-with-me

홈 › Papers

Words in Motion: Extracting Interpretable Control Vectors for Motion Transformers

2024-06-17 · Omer Sahin Tas, Royden Wagner

Transformer-based models generate hidden states that are difficult to interpret. In this work, we analyze hidden states and modify them at inference, with a focus on motion forecasting. We use linear probing to analyze whether interpretable features are embedded in hidden states. Our experiments reveal high probing accuracy, indicating latent space regularities with functionally important directions. Building on this, we use the directions between hidden states with opposing features to fit control vectors. At inference, we add our control vectors to hidden states and evaluate their impact on predictions. Remarkably, such modifications preserve the feasibility of predictions. We further refine our control vectors using sparse autoencoders (SAEs). This leads to more linear changes in predictions when scaling control vectors. Our approach enables mechanistic interpretation as well as zero-shot generalization to unseen dataset characteristics with negligible computational overhead.

📄 PDF Abstract BibTeX arXiv:2406.11624

Code (1)

kit-mrt/future-motion 공식 구현 pytorch

Tasks

Motion ForecastingZero-shot Generalization

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

LingoMotion: An Interpretable and Unambiguous Symbolic Representation for Human Motion

2026-03-13 · Yao Zhang, Zhuchenyang Liu, Yu Xiao arxiv

Existing representations for human motion, such as MotionGPT, often operate as black-box latent vectors with limited interpretability and build on joint positions which can cause ambiguity. Inspired by the hierarchical s…

Guilt by Association: Emotion Intensities in Lexical Representations

2021-04-18 · EMNLP 2021 11 · Shahab Raji, Gerard de Melo

What do word vector representations reveal about the emotions associated with words? In this study, we consider the task of estimating word-level emotion intensity scores for specific emotions, exploring unsupervised, su…

Valence-Arousal Subspace in LLMs: Circular Emotion Geometry and Multi-Behavioral Control

2026-04-03 · Lihao Sun, Lewen Yan, Xiaoya Lu, Andrew Lee 외 arxiv

We show that emotion vectors in LLMs are organized by a two-dimensional valence-arousal (VA) subspace exhibiting circular geometry. Through principal component decomposition and ridge regression, we recover meaningful VA…

Disentangling Latent Emotions of Word Embeddings on Complex Emotional Narratives

2019-08-15 · Zhengxuan Wu, Yueyi Jiang

Word embedding models such as GloVe are widely used in natural language processing (NLP) research to convert words into vectors. Here, we provide a preliminary guide to probe latent emotions in text through GloVe word ve…

Word Embeddings

Adjusting Interpretable Dimensions in Embedding Space with Human Judgments

2024-04-03 · Katrin Erk, Marianna Apidianaki

Embedding spaces contain interpretable dimensions indicating gender, formality in style, or even object properties. This has been observed multiple times. Such interpretable dimensions are becoming valuable tools in diff…

Object