paper-with-me

홈 › Papers

Representation Engineering: A Top-Down Approach to AI Transparency

2023-10-02 · Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, Shashwat Goel, Nathaniel Li, Michael J. Byun, Zifan Wang, Alex Mallen, Steven Basart, Sanmi Koyejo, Dawn Song, Matt Fredrikson, J. Zico Kolter, Dan Hendrycks

In this paper, we identify and characterize the emerging area of representation engineering (RepE), an approach to enhancing the transparency of AI systems that draws on insights from cognitive neuroscience. RepE places population-level representations, rather than neurons or circuits, at the center of analysis, equipping us with novel methods for monitoring and manipulating high-level cognitive phenomena in deep neural networks (DNNs). We provide baselines and an initial analysis of RepE techniques, showing that they offer simple yet effective solutions for improving our understanding and control of large language models. We showcase how these methods can provide traction on a wide range of safety-relevant problems, including honesty, harmlessness, power-seeking, and more, demonstrating the promise of top-down transparency research. We hope that this work catalyzes further exploration of RepE and fosters advancements in the transparency and safety of AI systems.

📄 PDF Abstract BibTeX arXiv:2310.01405

Code (5)

andyzoujm/representation-engineering 공식 구현 pytorch
cma1114/activation_steering pytorch
kaiyuhe998/rulearn_idea
steering-vectors/steering-vectors pytorch
sunblaze-ucb/political_leaning_RepE pytorch

Tasks

Question Answering

Similar Papers 제목 키워드 기반

Why Representation Engineering Works: A Theoretical and Empirical Study in Vision-Language Models

2025-03-25 · Bowei Tian, Xuntao Lyu, Meng Liu, Hongyi Wang 외

Representation Engineering (RepE) has emerged as a powerful paradigm for enhancing AI transparency by focusing on high-level representations rather than individual neurons or circuits. It has proven effective in improvin…

DescriptiveFairness

A Timeline and Analysis for Representation Plasticity in Large Language Models

2024-10-08 · Akshat Kannan

The ability to steer AI behavior is crucial to preventing its long term dangerous and catastrophic potential. Representation Engineering (RepE) has emerged as a novel, powerful method to steer internal model behaviors, s…

Information Ecosystem Reengineering via Public Sector Knowledge Representation

2025-08-21 · Mayukh Bagchi arxiv

Information Ecosystem Reengineering (IER) -- the technological reconditioning of information sources, services, and systems within a complex information ecosystem -- is a foundational challenge in the digital transformat…

Decision Making

Towards Accountability for Machine Learning Datasets: Practices from Software Engineering and Infrastructure

2020-10-23 · Ben Hutchinson, Andrew Smart, Alex Hanna, Emily Denton 외

Rising concern for the societal implications of artificial intelligence systems has inspired demands for greater transparency and accountability. However the datasets which empower machine learning are often used, shared…

BIG-bench Machine LearningDecision Making

DELTA: Variational Disentangled Learning for Privacy-Preserving Data Reprogramming

2025-08-31 · Arun Vignesh Malarkkan, Haoyue Bai, Anjali Kaushik, Yanjie Fu arxiv

In real-world applications, domain data often contains identifiable or sensitive attributes, is subject to strict regulations (e.g., HIPAA, GDPR), and requires explicit data feature engineering for interpretability and t…

Reinforcement LearningFeature Engineering