Evaluating The Impact of Stimulus Quality in Investigations of LLM Language Performance
Recent studies employing Large Language Models (LLMs) to test the Argument from the Poverty of the Stimulus (APS) have yielded contrasting results across syntactic phenomena. This paper investigates the hypothesis that characteristics of the stimuli used in recent studies, including lexical ambiguities and structural complexities, may confound model performance. A methodology is proposed for re-evaluating LLM competence on syntactic prediction, focusing on GPT-2. This involves: 1) establishing a baseline on previously used (both filtered and unfiltered) stimuli, and 2) generating a new, refined dataset using a state-of-the-art (SOTA) generative LLM (Gemini 2.5 Pro Preview) guided by linguistically-informed templates designed to mitigate identified confounds. Our preliminary findings indicate that GPT-2 demonstrates notably improved performance on these refined PG stimuli compared to baselines, suggesting that stimulus quality significantly influences outcomes in surprisal-based evaluations of LLM syntactic competency.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Evaluating neuronal codes for inference using Fisher information
Many studies have explored the impact of response variability on the quality of sensory codes. The source of this variability is almost always assumed to be intrinsic to the brain. However, when inferring a particular st…
Neurophysiological Investigation of Context Modulation based on Musical Stimulus
There are numerous studies which suggest that perhaps music is truly the language of emotions. Music seems to have an almost willful, evasive quality, defying simple explanation, and indeed requires deeper neurophysiolog…
Industrialized Deception: The Collateral Effects of LLM-Generated Misinformation on Digital Ecosystems
Generative AI and misinformation research has evolved since our 2024 survey. This paper presents an updated perspective, transitioning from literature review to practical countermeasures. We report on changes in the thre…
High-Resource Methodological Bias in Low-Resource Investigations
The central bottleneck for low-resource NLP is typically regarded to be the quantity of accessible data, overlooking the contribution of data quality. This is particularly seen in the development and evaluation of low-re…
Machine TranslationPOSPOS TaggingTranslation+1Beyond Output Correctness: Benchmarking and Evaluating Large Language Model Reasoning in Coding Tasks
Large language models (LLMs) increasingly rely on explicit reasoning to solve coding tasks, yet evaluating the quality of this reasoning remains challenging. Existing reasoning evaluators are not designed for coding, and…
Code Generation