paper-with-me

Papers

Multilingual Idioms in Sentences and Conversations Across High-, Medium-, and Low-Resource Languages

2026-06-01 · Saeed Almheiri, Bilal Elbouardi, Salsabila Zahirah Pranida, Irina Nikishina, Ashwath Rao B, Parameswari Krishnamurthy, Muhammad Cendekia Airlangga, Rifo Ahmad Genadi, Nguyen Phan Gia Bao, Amir Hossein Yari, Hawau Olamide Toyin, Nurdaulet Mukhituly, Mena Attia, Besher Hassan, Ahmad Fathan Hidayatullah, Tatsuki Kuribayashi, Haonan Li, Suma Bhat, Fajri Koto arxiv

Idiomatic expressions pose a major challenge for multilingual NLP because their meanings shift between figurative and literal usage, often requiring context for accurate interpretation. Prior work has focused on high-resource languages typically evaluates isolated idiom-meaning questions, overlooking realistic discourse. We introduce MIDI, a multilingual idiom dataset spanning 3 high-, 3 medium-, and 12 low-resource languages, curated by native speakers. Unlike previous datasets, MIDI provides idioms embedded in both sentence-level and conversational contexts, capturing both literal and figurative readings. Benchmarking state-of-the-art models shows that idiom comprehension degrades in low-resource languages and that, in all resource tiers, literal interpretations are substantially harder than figurative ones. Conversational context improves performance but does not eliminate these disparities. Through controlled tests and interventions on hidden representations, we further separate memorization from reasoning, exposing core limitations of current models.

📄 PDF Abstract BibTeX arXiv:2606.02147

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

LIDIOMS: A Multilingual Linked Idioms Data Set

2018-02-22 · LREC 2018 5 · Diego Moussallem, Mohamed Ahmed Sherif, Diego Esteves, Marcos Zampieri 외

In this paper, we describe the LIDIOMS data set, a multilingual RDF representation of idioms currently containing five languages: English, German, Italian, Portuguese, and Russian. The data set is intended to support nat…

Crossing the Threshold: Idiomatic Machine Translation through Retrieval Augmentation and Loss Weighting

2023-10-10 · Emmy Liu, Aditi Chaudhary, Graham Neubig

Idioms are common in everyday language, but often pose a challenge to translators because their meanings do not follow from the meanings of their parts. Despite significant advances, machine translation systems still str…

4kMachine TranslationRetrievalTranslation

Handling Idioms in Symbolic Multilingual Natural Language Generation

2022-06-01 · LREC (MWE) 2022 6 · Michaelle Dubé, François Lareau

While idioms are usually very rigid in their expression, they sometimes allow a certain level of freedom in their usage, with modifiers or complements splitting them or being syntactically attached to internal nodes rath…

Text Generation

ID10M: Idiom Identification in 10 Languages

2022-07-01 · Findings (NAACL) 2022 7 · Simone Tedeschi, Federico Martelli, Roberto Navigli

Idioms are phrases which present a figurative meaning that cannot be (completely) derived by looking at the meaning of their individual components.Identifying and understanding idioms in context is a crucial goal and a k…

Natural Language Understanding

"Be My Cheese?": Cultural Nuance Benchmarking for Machine Translation in Multilingual LLMs

2026-02-04 · Madison Van Doren, Casey Ford, Jennifer Barajas, Riley VanMeter 외 arxiv

We present a large-scale human evaluation benchmark for assessing cultural localisation in machine translation produced by state-of-the-art multilingual large language models (LLMs). Existing MT benchmarks emphasise toke…

Machine Translation