paper-with-me

Papers

In Machina N400: Pinpointing Where a Causal Language Model Detects Semantic Violations

2025-11-24 · Christos-Nikolaos Zacharopoulos, Revekka Kyriakoglou arxiv

How and where does a transformer notice that a sentence has gone semantically off the rails? To explore this question, we evaluated the causal language model (phi-2) using a carefully curated corpus, with sentences that concluded plausibly or implausibly. Our analysis focused on the hidden states sampled at each model layer. To investigate how violations are encoded, we utilized two complementary probes. First, we conducted a per-layer detection using a linear probe. Our findings revealed that a simple linear decoder struggled to distinguish between plausible and implausible endings in the lowest third of the model's layers. However, its accuracy sharply increased in the middle blocks, reaching a peak just before the top layers. Second, we examined the effective dimensionality of the encoded violation. Initially, the violation widens the representational subspace, followed by a collapse after a mid-stack bottleneck. This might indicate an exploratory phase that transitions into rapid consolidation. Taken together, these results contemplate the idea of alignment with classical psycholinguistic findings in human reading, where semantic anomalies are detected only after syntactic resolution, occurring later in the online processing sequence.

📄 PDF Abstract BibTeX arXiv:2511.19232

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

TextMachina: Seamless Generation of Machine-Generated Text Datasets

2024-01-08 · Areg Mikael Sarvazyan, José Ángel González, Marc Franco-Salvador

Recent advancements in Large Language Models (LLMs) have led to high-quality Machine-Generated Text (MGT), giving rise to countless new use cases and applications. However, easy access to LLMs is posing new challenges du…

Boundary Detection

Local Utility and Multivariate Risk Aversion

2021-02-08 · Arthur Charpentier, Alfred Galichon, Marc Henry

We revisit Machina's local utility as a tool to analyze attitudes to multivariate risks. We show that for non-expected utility maximizers choosing between multivariate prospects, aversion to multivariate mean preserving …

DexMachina: Functional Retargeting for Bimanual Dexterous Manipulation

2025-05-30 · Zhao Mandi, Yifan Hou, Dieter Fox, Yashraj Narang 외

We study the problem of functional retargeting: learning dexterous manipulation policies to track object states from human hand-object demonstrations. We focus on long-horizon, bimanual tasks with articulated objects, wh…

Object

ReCCoVER: Detecting Causal Confusion for Explainable Reinforcement Learning

2022-03-21 · Jasmina Gajcin, Ivana Dusparic

Despite notable results in various fields over the recent years, deep reinforcement learning (DRL) algorithms lack transparency, affecting user trust and hindering their deployment to high-risk tasks. Causal confusion re…

Deep Reinforcement Learningfeature selectionreinforcement-learningReinforcement Learning+1

Comparing Causal Frameworks: Potential Outcomes, Structural Models, Graphs, and Abstractions

2023-06-25 · NeurIPS 2023 11

The aim of this paper is to make clear and precise the relationship between the Rubin causal model (RCM) and structural causal model (SCM) frameworks for causal inference. Adopting a neutral logical perspective, and draw…

Causal Inference