paper-with-me

홈 › Papers

Interpretability of Language Models via Task Spaces

2024-06-10 · Lucas Weber, Jaap Jumelet, Elia Bruni, Dieuwke Hupkes

The usual way to interpret language models (LMs) is to test their performance on different benchmarks and subsequently infer their internal processes. In this paper, we present an alternative approach, concentrating on the quality of LM processing, with a focus on their language abilities. To this end, we construct 'linguistic task spaces' -- representations of an LM's language conceptualisation -- that shed light on the connections LMs draw between language phenomena. Task spaces are based on the interactions of the learning signals from different linguistic phenomena, which we assess via a method we call 'similarity probing'. To disentangle the learning signals of linguistic phenomena, we further introduce a method called 'fine-tuning via gradient differentials' (FTGD). We apply our methods to language models of three different scales and find that larger models generalise better to overarching general concepts for linguistic tasks, making better use of their shared structure. Further, the distributedness of linguistic processing increases with pre-training through increased parameter sharing between related linguistic tasks. The overall generalisation patterns are mostly stable throughout training and not marked by incisive stages, potentially explaining the lack of successful curriculum strategies for LMs.

📄 PDF Abstract BibTeX arXiv:2406.06441

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Learning Disentangled Semantic Spaces of Explanations via Invertible Neural Networks

2023-05-02 · Yingji Zhang, Danilo S. Carvalho, André Freitas

Disentangled latent spaces usually have better semantic separability and geometrical properties, which leads to better interpretability and more controllable data generation. While this has been well investigated in Comp…

DisentanglementSentenceStyle Transfer

Are Embedding Spaces Interpretable? Results of an Intrusion Detection Evaluation on a Large French Corpus

2022-06-01 · LREC 2022 6 · Thibault Prouteau, Nicolas Dugué, Nathalie Camelin, Sylvain Meignier

Word embedding methods allow to represent words as vectors in a space that is structured using word co-occurrences so that words with close meanings are close in this space. These vectors are then provided as input to au…

Intrusion Detection

Indic-TunedLens: Interpreting Multilingual Models in Indian Languages

2026-01-29 · Mihir Panchal, Deeksha Varshney, Mamta, Asif Ekbal arxiv

Multilingual large language models (LLMs) are increasingly deployed in linguistically diverse regions like India, yet most interpretability tools remain tailored to English. Prior work reveals that LLMs often operate in …

The Geometry of Distributed Representations for Better Alignment, Attenuated Bias, and Improved Interpretability

2020-11-25 · Sunipa Dev

High-dimensional representations for words, text, images, knowledge graphs and other structured data are commonly used in different paradigms of machine learning and data mining. These representations have different degr…

Knowledge Graphs

When Meanings Meet: Investigating the Emergence and Quality of Shared Concept Spaces during Multilingual Language Model Training

2026-01-30 · Felicia Körner, Max Müller-Eberstein, Anna Korhonen, Barbara Plank arxiv

Training Large Language Models (LLMs) with high multilingual coverage is becoming increasingly important -- especially when monolingual resources are scarce. Recent studies have found that LLMs process multilingual input…

Cross-Lingual Transfer