paper-with-me

Papers

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean

2025-06-02 · Sungho Kim, Nayeon Kim, Taehee Jeon, SangKeun Lee

We introduce the $\underline{Ko}rean \underline{G}rammar \underline{E}valuation Bench\underline{M}ark (KoGEM)$, designed to assess the linguistic competence of LLMs and humans in Korean. KoGEM consists of 1.5k multiple-choice QA pairs covering five main categories and 16 subcategories. The zero-shot evaluation of 27 LLMs of various sizes and types reveals that while LLMs perform remarkably well on straightforward tasks requiring primarily definitional knowledge, they struggle with tasks that demand the integration of real-world experiential knowledge, such as phonological rules and pronunciation. Furthermore, our in-depth analysis suggests that incorporating such experiential knowledge could enhance the linguistic competence of LLMs. With KoGEM, we not only highlight the limitations of current LLMs in linguistic competence but also uncover hidden facets of LLMs in linguistic competence, paving the way for enhancing comprehensive language understanding. Our code and dataset are available at: https://github.com/SungHo3268/KoGEM.

📄 PDF Abstract BibTeX arXiv:2506.01237

Code (1)

sungho3268/kogem 공식 구현 pytorch

Tasks

Multiple-choice

Similar Papers 제목 키워드 기반

Natural Language Processing RELIES on Linguistics

2024-05-09 · Juri Opitz, Shira Wein, Nathan Schneider

Large Language Models (LLMs) have become capable of generating highly fluent text in certain languages, without modules specially designed to capture grammar or semantic coherence. What does this mean for the future of l…

An Iterative Polishing Framework based on Quality Aware Masked Language Model for Chinese Poetry Generation

2019-11-29 · Liming Deng, Jie Wang, Hangming Liang, Hui Chen 외

Owing to its unique literal and aesthetical characteristics, automatic generation of Chinese poetry is still challenging in Artificial Intelligence, which can hardly be straightforwardly realized by end-to-end methods. I…

DecoderLanguage ModelingLanguage ModellingMulti-Task Learning

Dissociating language and thought in large language models

2023-01-16 · Kyle Mahowald, Anna A. Ivanova, Idan A. Blank, Nancy Kanwisher 외

Large Language Models (LLMs) have come closest among all models to date to mastering human language, yet opinions about their linguistic and cognitive capabilities remain split. Here, we evaluate LLMs using a distinction…

On the Nature of BERT: Correlating Fine-Tuning and Linguistic Competence

2022-10-01 · COLING 2022 10 · Federica Merendi, Felice Dell’Orletta, Giulia Venturi

Several studies in the literature on the interpretation of Neural Language Models (NLM) focus on the linguistic generalization abilities of pre-trained models. However, little attention is paid to how the linguistic know…

Holmes: A Benchmark to Assess the Linguistic Competence of Language Models

2024-04-29 · Andreas Waldis, Yotam Perlitz, Leshem Choshen, Yufang Hou 외

We introduce Holmes, a new benchmark designed to assess language models (LMs) linguistic competence - their unconscious understanding of linguistic phenomena. Specifically, we use classifier-based probing to examine LMs'…

Part-Of-Speech Tagging