False perspectives on human language: why statistics needs linguistics
A sharp tension exists about the nature of human language between two opposite parties: those who believe that statistical surface distributions, in particular using measures like surprisal, provide a better understanding of language processing, vs. those who believe that discrete hierarchical structures implementing linguistic information such as syntactic ones are a better tool. In this paper, we show that this dichotomy is a false one. Relying on the fact that statistical measures can be defined on the basis of either structural or non-structural models, we provide empirical evidence that only models of surprisal that reflect syntactic structure are able to account for language regularities.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Language Statistics and False Belief Reasoning: Evidence from 41 Open-Weight LMs
Research on mental state reasoning in language models (LMs) has the potential to inform theories of human social cognition--such as the theory that mental state reasoning emerges in part from language exposure--and our u…
Towards Low-Resource Alignment to Diverse Perspectives with Sparse Feedback
As language models have a greater impact on society, it is important to ensure they are aligned to a diverse range of perspectives and are able to reflect nuance in human values. However, the most popular training paradi…
Hate Speech DetectionReinforcement Learning from Human Feedback: Whose Culture, Whose Values, Whose Perspectives?
We argue for the epistemic and ethical advantages of pluralism in Reinforcement Learning from Human Feedback (RLHF) in the context of Large Language Models (LLM). Drawing on social epistemology and pluralist philosophy o…
Philosophyreinforcement-learningReinforcement LearningData Science Students Perspectives on Learning Analytics: An Application of Human-Led and LLM Content Analysis
Objective This study is part of a series of initiatives at a UK university designed to cultivate a deep understanding of students' perspectives on analytics that resonate with their unique learning needs. It explores col…
Large Language ModelRAGRetrieval-augmented GenerationAttention to Non-Adopters
Although language model-based chat systems are increasingly used in daily life, most Americans remain non-adopters of chat-based LLMs -- as of June 2025, 66% had never used ChatGPT. At the same time, LLM development and …