paper-with-me

홈 › Papers

NLP Datasets for Idiom and Figurative Language Tasks

2025-11-20 · Blake Matheny, Phuong Minh Nguyen, Minh Le Nguyen, Stephanie Reynolds arxiv

Idiomatic and figurative language form a large portion of colloquial speech and writing. With social media, this informal language has become more easily observable to people and trainers of large language models (LLMs) alike. While the advantage of large corpora seems like the solution to all machine learning and Natural Language Processing (NLP) problems, idioms and figurative language continue to elude LLMs. Finetuning approaches are proving to be optimal, but better and larger datasets can help narrow this gap even further. The datasets presented in this paper provide one answer, while offering a diverse set of categories on which to build new models and develop new approaches. A selection of recent idiom and figurative language datasets were used to acquire a combined idiom list, which was used to retrieve context sequences from a large corpus. One large-scale dataset of potential idiomatic and figurative language expressions and two additional human-annotated datasets of definite idiomatic and figurative language expressions were created to evaluate the baseline ability of pre-trained language models in handling figurative meaning through idiom recognition (detection) tasks. The resulting datasets were post-processed for model agnostic training compatibility, utilized in training, and evaluated on slot labeling and sequence tagging.

📄 PDF Abstract BibTeX arXiv:2511.16345

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Bilingual Lexical Access and Cognate Idiom Comprehension

2020-12-01 · COLING (CogALex) 2020 12 · Eve Fleisig

Language transfer can facilitate learning L2 words whose form and meaning are similar to L1 words, or hinder speakers when the languages differ. L2 idioms introduce another layer of challenge, as language transfer could …

Beyond Understanding: Evaluating the Pragmatic Gap in LLMs' Cultural Processing of Figurative Language

2025-10-27 · Mena Attia, Aashiq Muhamed, Mai Alkhamissi, Thamar Solorio 외 arxiv

We present a comprehensive evaluation of the ability of large language models (LLMs) to process culturally grounded language, specifically to understand and pragmatically use figurative expressions that encode local know…

English Proverbs

Multilingual Idioms in Sentences and Conversations Across High-, Medium-, and Low-Resource Languages

2026-06-01 · Saeed Almheiri, Bilal Elbouardi, Salsabila Zahirah Pranida, Irina Nikishina 외 arxiv

Idiomatic expressions pose a major challenge for multilingual NLP because their meanings shift between figurative and literal usage, often requiring context for accurate interpretation. Prior work has focused on high-res…

When Words Don't Mean What They Say: Figurative Understanding in Bengali Idioms

2026-02-13 · Adib Sakhawat, Shamim Ara Parveen, Md Ruhul Amin, Shamim Al Mahmud 외 arxiv

Figurative language understanding remains a significant challenge for Large Language Models (LLMs), especially for low-resource languages. To address this, we introduce a new idiom dataset, a large-scale, culturally-grou…

FFE-Hallu:Hallucinations in Fixed Figurative Expressions:Benchmark of Idioms and Proverbs in the Persian Language

2026-01-27 · Faezeh Hosseini, Mohammadali Yousefzadeh, Yadollah Yaghoobzadeh arxiv

Figurative language, particularly fixed figurative expressions (FFEs) such as idioms and proverbs, poses persistent challenges for large language models (LLMs). Unlike literal phrases, FFEs are culturally grounded, large…