paper-with-me

홈 › Papers

Thunder-KoNUBench: A Corpus-Aligned Benchmark for Korean Negation Understanding

2026-01-08 · Sungmok Jung, Yeonkyoung So, Joonhak Lee, Sangho Kim, Yelim Ahn, Jaejin Lee arxiv

Although negation is known to challenge large language models (LLMs), benchmarks for evaluating negation understanding-especially in Korean-are scarce. We conduct a corpus-based analysis of Korean negation and show that LLM performance degrades under negation. We then introduce Thunder-KoNUBench, a sentence-level negation understanding benchmark that reflects the empirical distribution of Korean negation phenomena. Evaluating 47 LLMs on Thunder-KoNUBench, we analyze the effects of model size and instruction tuning, and perform error analysis to better understand model behavior. We further show that fine-tuning on Thunder-KoNUBench improves negation understanding and broader contextual comprehension in Korean.

📄 PDF Abstract BibTeX arXiv:2601.04693

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Less Is More: Reducing Token Counts Without Compromising Performance

2025-06-18 · Gyeongje Cho, Yeonkyoung So, Sangmin Lee, Jaejin Lee arxiv

Tokenization directly affects the inference efficiency of large language models, since fragmented tokenization increases sequence length and generation cost. Although longer, multi-word tokens can reduce fertility, naive…

Enriching the Korean Learner Corpus with Multi-reference Annotations and Rubric-Based Scoring

2025-05-01 · Jayoung Song, Kyungtae Lim, Jungyeul Park

Despite growing global interest in Korean language education, there remains a significant lack of learner corpora tailored to Korean L2 writing. To address this gap, we enhance the KoLLA Korean learner corpus by adding m…

DiversityGrammatical Error Correction

GECKO: Generative Language Model for English, Code and Korean

2024-05-24 · Sungwoo Oh, Donggyu Kim

We introduce GECKO, a bilingual large language model (LLM) optimized for Korean and English, along with programming languages. GECKO is pretrained on the balanced, high-quality corpus of Korean and English employing LLaM…

kmmluLanguage ModelingLanguage ModellingLarge Language Model+1

KoSpeech: Open-Source Toolkit for End-to-End Korean Speech Recognition

2020-09-07 · Soohwan Kim, Seyoung Bae, Cheolhwang Won

We present KoSpeech, an open-source software, which is modular and extensible end-to-end Korean automatic speech recognition (ASR) toolkit based on the deep learning library PyTorch. Several automatic speech recognition …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

UD-KSL Treebank v1.3: A semi-automated framework for aligning XPOS-extracted units with UPOS tags

2025-06-10 · Hakyung Sung, Gyu-Ho Shin, Chanyoung Lee, You Kyung Sung 외

The present study extends recent work on Universal Dependencies annotations for second-language (L2) Korean by introducing a semi-automated framework that identifies morphosyntactic constructions from XPOS sequences and …

Dependency Parsing