paper-with-me

Papers

LIDIOMS: A Multilingual Linked Idioms Data Set

2018-02-22 · LREC 2018 5 · Diego Moussallem, Mohamed Ahmed Sherif, Diego Esteves, Marcos Zampieri, Axel-Cyrille Ngonga Ngomo

In this paper, we describe the LIDIOMS data set, a multilingual RDF representation of idioms currently containing five languages: English, German, Italian, Portuguese, and Russian. The data set is intended to support natural language processing applications by providing links between idioms across languages. The underlying data was crawled and integrated from various sources. To ensure the quality of the crawled data, all idioms were evaluated by at least two native speakers. Herein, we present the model devised for structuring the data. We also provide the details of linking LIDIOMS to well-known multilingual data sets such as BabelNet. The resulting data set complies with best practices according to Linguistic Linked Open Data Community.

📄 PDF Abstract BibTeX arXiv:1802.08148

Code (1)

dice-group/LIdioms 공식 구현

Similar Papers 제목 키워드 기반

Handling Idioms in Symbolic Multilingual Natural Language Generation

2022-06-01 · LREC (MWE) 2022 6 · Michaelle Dubé, François Lareau

While idioms are usually very rigid in their expression, they sometimes allow a certain level of freedom in their usage, with modifiers or complements splitting them or being syntactically attached to internal nodes rath…

Text Generation

ID10M: Idiom Identification in 10 Languages

2022-07-01 · Findings (NAACL) 2022 7 · Simone Tedeschi, Federico Martelli, Roberto Navigli

Idioms are phrases which present a figurative meaning that cannot be (completely) derived by looking at the meaning of their individual components.Identifying and understanding idioms in context is a crucial goal and a k…

Natural Language Understanding

The Mediomatix Corpus: Parallel Data for Romansh Language Varieties via Comparable Schoolbooks

2025-08-22 · Zachary Hopton, Jannis Vamvas, Andrin Büchler, Anna Rutkiewicz 외 arxiv

The five idioms (i.e., varieties) of the Romansh language are largely standardized and are taught in the schools of the respective communities in Switzerland. In this paper, we present the first parallel corpus of Romans…

Machine Translation

Memorization or Reasoning? Exploring the Idiom Understanding of LLMs

2025-05-22 · Jisu Kim, Youngwoo Shin, Uiji Hwang, Jihun Choi 외

Idioms have long posed a challenge due to their unique linguistic properties, which set them apart from other common expressions. While recent studies have leveraged large language models (LLMs) to handle idioms across v…

Machine TranslationMemorizationSentence

Multilingual Idioms in Sentences and Conversations Across High-, Medium-, and Low-Resource Languages

2026-06-01 · Saeed Almheiri, Bilal Elbouardi, Salsabila Zahirah Pranida, Irina Nikishina 외 arxiv

Idiomatic expressions pose a major challenge for multilingual NLP because their meanings shift between figurative and literal usage, often requiring context for accurate interpretation. Prior work has focused on high-res…