Functional Emotions or Situational Contexts? A Discriminating Test from the Mythos Preview System Card
The Claude Mythos Preview system card deploys emotion vectors, sparse autoencoder (SAE) features, and activation verbalisers to study model internals during misaligned behaviour. The two primary toolkits are not jointly reported on the most alignment-relevant episodes. This note identifies two hypotheses that are qualitatively consistent with the published results: that the emotion vectors track functional emotions that causally drive behaviour, or that they are a projection of a richer situational-context structure onto human emotional axes. The hypotheses can be distinguished by cross-referencing the two toolkits on episodes where only one is currently reported: most directly, applying emotion probes to the strategic concealment episodes analysed only with SAE features. If emotion probes show flat activation while SAE features are strongly active, the alignment-relevant structure lies outside the emotion subspace. Which hypothesis is correct determines whether emotion-based monitoring will robustly detect dangerous model behaviour or systematically miss it.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Emotional Neural Language Generation Grounded in Situational Contexts
Emotional language generation is one of the keys to human-like artificial intelligence. Humans use different type of emotions depending on the situation of the conversation. Emotions also play an important role in mediat…
Language ModelingLanguage ModellingText GenerationDepressionEmo: A novel dataset for multilabel classification of depression emotions
Emotions are integral to human social interactions, with diverse responses elicited by various situational contexts. Particularly, the prevalence of negative emotional states has been correlated with negative outcomes fo…
text-classificationText ClassificationEmotion Concepts and their Function in a Large Language Model
Large language models (LLMs) sometimes appear to exhibit emotional reactions. We investigate why this is the case in Claude Sonnet 4.5 and explore implications for alignment-relevant behavior. We find internal representa…
An Appraisal-Based Chain-Of-Emotion Architecture for Affective Language Model Game Agents
The development of believable, natural, and interactive digital artificial agents is a field of growing interest. Theoretical uncertainties and technical barriers present considerable challenges to the field, particularl…
Emotional IntelligenceLanguage ModelingLanguage ModellingEmotion pattern detection on facial videos using functional statistics
There is an increasing scientific interest in automatically analysing and understanding human behavior, with particular reference to the evolution of facial expressions and the recognition of the corresponding emotions. …
Emotion Recognition