paper-with-me

홈 › Papers

Position: Anthropomorphic Misalignment Research Needs Stronger Evidence

2026-05-29 · Vansh Gupta, Peter Nutter, Samuel Stante, Andreas Krause, Florian Tramèr, Lukas Fluri, Xin Chen, Anna Hedström arxiv

We argue that many Anthropomorphic Misalignment Research (AMR) studies need stronger evidence to ensure that they can provide a robust foundation for critical safety decisions, such as model deployment and regulation. By evaluating failure modes across different misalignment concepts, such as deception, emergent misalignment, and sycophancy, we show how conceptual ambiguity, non-robust datasets, experimental design, and insufficient causal interventions can lead to overinterpretation of model behaviors. This position paper aims to offer guidance on evidentiary considerations that can help improve methodological rigor in AMR. To achieve this, we provide a clear call to action through a proposed framework of evidence levels and a diagnostic checklist. These shared standards will enable more productive scientific discourse and ensure that claims about AI risks rest on solid empirical foundations.

📄 PDF Abstract BibTeX arXiv:2606.07612

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Thinking beyond the anthropomorphic paradigm benefits LLM research

2025-02-13 · Lujain Ibrahim, Myra Cheng

Anthropomorphism, or the attribution of human traits to technology, is an automatic and unconscious response that occurs even in those with advanced technical expertise. In this position paper, we analyze hundreds of tho…

Articles

Character as a Latent Variable in Large Language Models: A Mechanistic Account of Emergent Misalignment and Conditional Safety Failures

2026-01-30 · Yanghao Su, Wenbo Zhou, Tianwei Zhang, Qiu Han 외 arxiv

Emergent Misalignment refers to a failure mode in which fine-tuning large language models (LLMs) on narrowly scoped data induces broadly misaligned behavior. Prior explanations mainly attribute this phenomenon to the gen…

Do Robots Really Need Anthropomorphic Hands? A Comparison of Human and Robotic Hands

2025-08-07 · Alexander Fabisch, Wadhah Zai El Amri, Chandandeep Singh, Nicolás Navarro-Guerrero arxiv

Human manipulation skills represent a pinnacle of voluntary motor functions, requiring the coordination of many degrees of freedom and the processing of high-dimensional sensor input to achieve remarkable dexterity. Thus…

Anthropomorphic Features for On-Line Signatures

2025-01-15 · Moises Diaz, Miguel A. Ferrer, Jose J. Quintana

Many features have been proposed in on-line signature verification. Generally, these features rely on the position of the on-line signature samples and their dynamic properties, as recorded by a tablet. This paper propos…

Position

Time Waits for No One! Analysis and Challenges of Temporal Misalignment

2021-11-14 · NAACL 2022 7 · Kelvin Luu, Daniel Khashabi, Suchin Gururangan, Karishma Mandyam 외

When an NLP model is trained on text data from one time period and tested or deployed on data from another, the resulting temporal misalignment can degrade end-task performance. In this work, we establish a suite of eigh…