paper-with-me

Papers

Evaluating Steering Techniques using Human Similarity Judgments

2025-05-25 · Zach Studdiford, Timothy T. Rogers, Siddharth Suresh, Kushin Mukherjee

Current evaluations of Large Language Model (LLM) steering techniques focus on task-specific performance, overlooking how well steered representations align with human cognition. Using a well-established triadic similarity judgment task, we assessed steered LLMs on their ability to flexibly judge similarity between concepts based on size or kind. We found that prompt-based steering methods outperformed other methods both in terms of steering accuracy and model-to-human alignment. We also found LLMs were biased towards 'kind' similarity and struggled with 'size' alignment. This evaluation approach, grounded in human cognition, adds further support to the efficacy of prompt-based steering and reveals privileged representational axes in LLMs prior to steering.

📄 PDF Abstract BibTeX arXiv:2505.19333

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingLarge Language Model

Methods 이 논문이 사용한 방법론

Focus 설명 없음
ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Towards a Gold Standard for Evaluating Danish Word Embeddings

2020-05-01 · LREC 2020 5 · Nina Schneidermann, Rasmus Hvingelby, Bolette Pedersen

This paper presents the process of compiling a model-agnostic similarity goal standard for evaluating Danish word embeddings based on human judgments made by 42 native speakers of Danish. Word embeddings resemble semanti…

Semantic SimilaritySemantic Textual SimilarityWord Embeddings

Evaluating Ways of Adapting Word Similarity

2019-08-01 · WS 2019 8 · Libby Barak, Adele Goldberg

People judge pairwise similarity by deciding which aspects of the words{'} meanings are relevant for the comparison of the given pair. However, computational representations of meaning rely on dimensions of the vector re…

Word Similarity

Between Rules and Reality: On the Context Sensitivity of LLM Moral Judgment

2026-03-24 · Adrian Sauter, Mona Schirmer arxiv

A human's moral decision depends heavily on the context. Yet research on LLM morality has largely studied fixed scenarios. We address this gap by introducing Contextual MoralChoice, a dataset of moral dilemmas with syste…

Words are all you need? Language as an approximation for human similarity judgments

2022-06-08 · Raja Marjieh, Pol van Rijn, Ilia Sucholutsky, Theodore R. Sumers 외

Human similarity judgments are a powerful supervision signal for machine learning applications based on techniques such as contrastive learning, information retrieval, and model alignment, but classical methods for colle…

AllContrastive LearningInformation RetrievalRetrieval+1

Sentence Mover's Similarity: Automatic Evaluation for Multi-Sentence Texts

2019-07-01 · ACL 2019 7 · Elizabeth Clark, Asli Celikyilmaz, Noah A. Smith

For evaluating machine-generated texts, automatic methods hold the promise of avoiding collection of human judgments, which can be expensive and time-consuming. The most common automatic metrics, like BLEU and ROUGE, dep…

Reinforcement LearningSemantic SimilaritySemantic Textual SimilaritySentence+1