paper-with-me

Papers

Enhancing Interpretability using Human Similarity Judgements to Prune Word Embeddings

2023-10-16 · Natalia Flechas Manrique, Wanqian Bao, Aurelie Herbelot, Uri Hasson

Interpretability methods in NLP aim to provide insights into the semantics underlying specific system architectures. Focusing on word embeddings, we present a supervised-learning method that, for a given domain (e.g., sports, professions), identifies a subset of model features that strongly improve prediction of human similarity judgments. We show this method keeps only 20-40% of the original embeddings, for 8 independent semantic domains, and that it retains different feature sets across domains. We then present two approaches for interpreting the semantics of the retained features. The first obtains the scores of the domain words (co-hyponyms) on the first principal component of the retained embeddings, and extracts terms whose co-occurrence with the co-hyponyms tracks these scores' profile. This analysis reveals that humans differentiate e.g. sports based on how gender-inclusive and international they are. The second approach uses the retained sets as variables in a probing task that predicts values along 65 semantically annotated dimensions for a dataset of 535 words. The features retained for professions are best at predicting cognitive, emotional and social dimensions, whereas features retained for fruits or vegetables best predict the gustation (taste) dimension. We discuss implications for alignment between AI systems and human knowledge.

📄 PDF Abstract BibTeX arXiv:2310.10262

Code (0)

등록된 구현이 없습니다.

Tasks

Word Embeddings

Similar Papers 제목 키워드 기반

Word Sense Distance in Human Similarity Judgements and Contextualised Word Embeddings

2020-06-01 · PaM 2020 6 · Janosch Haber, Massimo Poesio

Homonymy is often used to showcase one of the advantages of context-sensitive word embedding techniques such as ELMo and BERT. In this paper we want to shift the focus to the related but less exhaustively explored phenom…

Word Embeddings

An Interpretable Music Similarity Measure Based on Path Interestingness

2021-08-03 · Giovanni Gabbolini, Derek Bridge

We introduce a novel and interpretable path-based music similarity measure. Our similarity measure assumes that items, such as songs and artists, and information about those items are represented in a knowledge graph. We…

Investigating Context Effects in Similarity Judgements in Large Language Models

2024-08-20 · Sagar Uprety, Amit Kumar Jaiswal, Haiming Liu, Dawei Song

Large Language Models (LLMs) have revolutionised the capability of AI models in comprehending and generating natural language text. They are increasingly being used to empower and deploy agents in real-world scenarios, w…

Modifications of Machine Translation Evaluation Metrics by Using Word Embeddings

2016-12-01 · WS 2016 12 · Haozhou Wang, Paola Merlo

Traditional machine translation evaluation metrics such as BLEU and WER have been widely used, but these metrics have poor correlations with human judgements because they badly represent word similarity and impose strict…

Machine TranslationSemantic Textual SimilarityTranslationWord Embeddings+1

A Linked Aggregate Code for Processing Faces (Revised Version)

2020-09-17 · Michael Lyons, Kazunori Morikawa

A model of face representation, inspired by the biology of the visual system, is compared to experimental data on the perception of facial similarity. The face representation model uses aggregate primary visual cortex (V…