paper-with-me

홈 › Papers

The Geometry of Distributed Representations for Better Alignment, Attenuated Bias, and Improved Interpretability

2020-11-25 · Sunipa Dev

High-dimensional representations for words, text, images, knowledge graphs and other structured data are commonly used in different paradigms of machine learning and data mining. These representations have different degrees of interpretability, with efficient distributed representations coming at the cost of the loss of feature to dimension mapping. This implies that there is obfuscation in the way concepts are captured in these embedding spaces. Its effects are seen in many representations and tasks, one particularly problematic one being in language representations where the societal biases, learned from underlying data, are captured and occluded in unknown dimensions and subspaces. As a result, invalid associations (such as different races and their association with a polar notion of good versus bad) are made and propagated by the representations, leading to unfair outcomes in different tasks where they are used. This work addresses some of these problems pertaining to the transparency and interpretability of such representations. A primary focus is the detection, quantification, and mitigation of socially biased associations in language representation.

📄 PDF Abstract BibTeX arXiv:2011.12465

Code (1)

sunipa/Abs-Orientation 공식 구현

Tasks

Knowledge Graphs

Methods 이 논문이 사용한 방법론

Interpretability 설명 없음

Similar Papers 제목 키워드 기반

Geometry of Human Perceptual Domains Emerges Transiently in LLM Representations

2026-05-27 · Simardeep Singh, Paras Chopra arxiv

While large language models (LLMs) are trained purely on textual data, prior work has shown that their internal representations can exhibit rich geometric structure in embedding space. Building on this line of work, we i…

Optimizing Against Safety Representations: Activation-Guided Adversarial Suffixes and the Geometry of Refusal

2026-07-09 · Ege Çakar, Hannah Guan, Kayden Kehe arxiv

Behavioral alignment in large language models often masks fragile internal safety representations. Recent work suggests that refusal behavior is mediated by low-dimensional directions in activation space. This raises que…

The geometry of online conversations and the causal antecedents of conflictual discourse

2026-02-17 · Carlo Santagiustina, Caterina Cruciani arxiv

This article investigates the causal antecedents of conflictual language and the geometry of interaction in online threaded conversations related to climate change. We employ three annotation dimensions, inferred through…

Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling

2025-07-10 · Haoyu Wu, Diankun Wu, Tianyu He, Junliang Guo 외 arxiv

Videos inherently represent 2D projections of a dynamic 3D world. However, our analysis suggests that video diffusion models trained solely on raw video data often fail to capture meaningful geometric-aware structure in …

Video Generation

Emotions Where Art Thou: Understanding and Characterizing the Emotional Latent Space of Large Language Models

2025-10-24 · Benjamin Reichman, Adar Avsian, Larry Heck arxiv

This work investigates how large language models (LLMs) internally represent emotion by analyzing the geometry of their hidden-state space. The paper identifies a low-dimensional emotional manifold and shows that emotion…