ePiC: Employing Proverbs in Context as a Benchmark for Abstract Language Understanding
While large language models have shown exciting progress on several NLP benchmarks, evaluating their ability for complex analogical reasoning remains under-explored. Here, we introduce a high-quality crowdsourced dataset of narratives for employing proverbs in context as a benchmark for abstract language understanding. The dataset provides fine-grained annotation of aligned spans between proverbs and narratives, and contains minimal lexical overlaps between narratives and proverbs, ensuring that models need to go beyond surface-level reasoning to succeed. We explore three tasks: (1) proverb recommendation and alignment prediction, (2) narrative generation for a given proverb and topic, and (3) identifying narratives with similar motifs. Our experiments show that neural language models struggle on these tasks compared to humans, and these tasks pose multiple learning challenges.
Code (1)
Similar Papers 제목 키워드 기반
Are Multilingual LLMs Culturally-Diverse Reasoners? An Investigation into Multicultural Proverbs and Sayings
Large language models (LLMs) are highly adept at question answering and reasoning tasks, but when reasoning in a situational context, human expectations vary depending on the relevant cultural common ground. As languages…
Question AnsweringSame Lesson, Different Story: Cross-Lingual Reconstruction of Cultural Narratives in Large Language Models
The evaluation of cultural grounding context becomes complex when multiple cultures convey the same moral lesson. This challenge is particularly relevant to large language models (LLMs), which produce narratives across a…
Semantic SimilarityMasalBench: A Benchmark for Contextual and Cross-Cultural Understanding of Persian Proverbs in LLMs
In recent years, multilingual Large Language Models (LLMs) have become an inseparable part of daily life, making it crucial for them to master the rules of conversational language in order to communicate effectively with…
English ProverbsProverbs Run in Pairs: Evaluating Proverb Translation Capability of Large Language Model
Despite achieving remarkable performance, machine translation (MT) research remains underexplored in terms of translating cultural elements in languages, such as idioms, proverbs, and colloquial expressions. This paper i…
Language ModelingLanguage ModellingLarge Language ModelMachine Translation+2Pupil size behavior during on line processing of sentences
In the present work we analyzed the pupil size behavior of forty subjects while they read well defined sentences with different contextual predictability (i.e., regular sentences and proverbs). In general, pupil size inc…
Sentence