paper-with-me

홈 › Papers

Metaphors are a Source of Cross-Domain Misalignment of Large Reasoning Models

2026-01-06 · Zhibo Hu, Chen Wang, Yanfeng Shu, Hye-young Paik, Liming Zhu arxiv

Earlier research has shown that metaphors influence human decision-making, raising the question of whether metaphors also influence large language models (LLMs)' reasoning pathways, given that their training data contain a large number of metaphors. In this work, we investigate the problem in the scope of the emergent misalignment problem, where LLMs can generalize patterns learned from misaligned content in one domain to another domain. We find strong evidence that metaphors in training data contribute to cross-domain misalignment in LLMs' reasoning outputs. With metaphor-based interventions during continued pre-training and fine-tuning for inducing misalignment, models exhibit significantly different degrees of emergent cross-domain misalignment. We also observe similar effects in re-alignment settings. As we further investigate this phenomenon, we find that metaphors are linked to the activation of latent features in large reasoning models. By monitoring these latent features, we design a detector that predicts misaligned content with high accuracy.

📄 PDF Abstract BibTeX arXiv:2601.03388

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Multi-Cultural Repository of Automatically Discovered Linguistic and Conceptual Metaphors

2014-05-01 · LREC 2014 5 · Samira Shaikh, Tomek Strzalkowski, Ting Liu, George Aaron Broadwell 외

In this article, we present details about our ongoing work towards building a repository of Linguistic and Conceptual Metaphors. This resource is being developed as part of our research effort into the large-scale detect…

Resources for the Detection of Conventionalized Metaphors in Four Languages

2014-05-01 · LREC 2014 5 · Lori Levin, Teruko Mitamura, Brian MacWhinney, Davida Fromm 외

This paper describes a suite of tools for extracting conventionalized metaphors in English, Spanish, Farsi, and Russian. The method depends on three significant resources for each language: a corpus of conventionalized m…

Not all ANIMALs are equal: metaphorical framing through source domains and semantic frames

2026-04-22 · Yulia Otmakhova, Matteo Guida, Lea Frermann arxiv

Metaphors are powerful framing devices, yet their source domains alone do not fully explain the specific associations they evoke. We argue that the interplay between source domains and semantic frames determines how meta…

Two Approaches to Metaphor Detection

2014-05-01 · LREC 2014 5 · Brian MacWhinney, Davida Fromm

Methods for automatic detection and interpretation of metaphors have focused on analysis and utilization of the ways in which metaphors violate selectional preferences (Martin, 2006). Detection and interpretation process…

Vocal Bursts Valence Prediction

Can GPT replace human raters? Validity and reliability of machine-generated norms for metaphors

2025-12-13 · Veronica Mangiaterra, Hamad Al-Azary, Chiara Barattieri di San Pietro, Paolo Canal 외 arxiv

As Large Language Models (LLMs) are increasingly being used in scientific research, the issue of their trustworthiness becomes crucial. In psycholinguistics, LLMs have been recently employed in automatically augmenting h…