paper-with-me

홈 › Papers

EVOKE: Emotion Vocabulary Of Korean and English

2026-02-11 · Yoonwon Jung, Hagyeong Shin, Benjamin K. Bergen arxiv

This paper introduces EVOKE (Emotion Vocabulary of Korean and English), a Korean-English parallel dataset of emotion words. The dataset offers comprehensive coverage of emotion words in each language, in addition to many-to-many translations between words in the two languages and identification of language-specific emotion words. The dataset contains 1,426 Korean words and 1,397 English words, and we systematically annotate 819 Korean and 924 English adjectives and verbs. We also annotate multiple meanings of each word and their relationships, identifying polysemous emotion words and emotion-related metaphors. The dataset is, to our knowledge, the most systematic and theory-agnostic dataset of emotion words in both Korean and English to date. It can serve as a practical tool for emotion science, psycholinguistics, computational linguistics, and natural language processing, allowing researchers to adopt different views on the resource reflecting their needs and theoretical perspectives. The dataset is publicly available at https://github.com/yoonwonj/EVOKE.

📄 PDF Abstract BibTeX arXiv:2602.10414

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Efficient and Effective Vocabulary Expansion Towards Multilingual Large Language Models

2024-02-22 · Seungduk Kim, Seungtaek Choi, Myeongho Jeong

This report introduces \texttt{EEVE-Korean-v1.0}, a Korean adaptation of large language models that exhibit remarkable capabilities across English and Korean text understanding. Building on recent highly capable but Engl…

GECKO: Generative Language Model for English, Code and Korean

2024-05-24 · Sungwoo Oh, Donggyu Kim

We introduce GECKO, a bilingual large language model (LLM) optimized for Korean and English, along with programming languages. GECKO is pretrained on the balanced, high-quality corpus of Korean and English employing LLaM…

kmmluLanguage ModelingLanguage ModellingLarge Language Model+1

Korean-Specific Emotion Annotation Procedure Using N-Gram-Based Distant Supervision and Korean-Specific-Feature-Based Distant Supervision

2020-05-01 · LREC 2020 5 · Young-Jun Lee, Chae-Gyun Lim, Ho-Jin Choi

Detecting emotions from texts is considerably important in an NLP task, but it has the limitation of the scarcity of manually labeled data. To overcome this limitation, many researchers have annotated unlabeled data with…

Emotion Classification

Optimizing Korean-Centric LLMs via Token Pruning

2026-04-17 · Hoyeol Kim, Hyeonwoo Kim arxiv

This paper presents a systematic benchmark of state-of-the-art multilingual large language models (LLMs) adapted via token pruning - a compression technique that eliminates tokens and embedding parameters corresponding t…

Instruction FollowingMachine Translation

HAE-RAE Bench: Evaluation of Korean Knowledge in Language Models

2023-09-06 · Guijin Son, Hanwool Lee, Suwan Kim, Huiseo Kim 외

Large language models (LLMs) trained on massive corpora demonstrate impressive capabilities in a wide range of tasks. While there are ongoing efforts to adapt these models to languages beyond English, the attention given…

General KnowledgeLogical ReasoningReading ComprehensionRetrieval