paper-with-me

Papers

ID10M: Idiom Identification in 10 Languages

2022-07-01 · Findings (NAACL) 2022 7 · Simone Tedeschi, Federico Martelli, Roberto Navigli

Idioms are phrases which present a figurative meaning that cannot be (completely) derived by looking at the meaning of their individual components.Identifying and understanding idioms in context is a crucial goal and a key challenge in a wide range of Natural Language Understanding tasks. Although efforts have been undertaken in this direction, the automatic identification and understanding of idioms is still a largely under-investigated area, especially when operating in a multilingual scenario. In this paper, we address such limitations and put forward several new contributions: we propose a novel multilingual Transformer-based system for the identification of idioms; we produce a high-quality automatically-created training dataset in 10 languages, along with a novel manually-curated evaluation benchmark; finally, we carry out a thorough performance analysis and release our evaluation suite at https://github.com/Babelscape/ID10M.

📄 PDF Abstract BibTeX

Code (1)

babelscape/id10m 공식 구현 pytorch

Tasks

Natural Language Understanding

Similar Papers 제목 키워드 기반

Gamified Crowdsourcing for Idiom Corpora Construction

2021-02-01 · Gülşen Eryiğit, Ali Şentaş, Johanna Monti

Learning idiomatic expressions is seen as one of the most challenging stages in second language learning because of their unpredictable meaning. A similar situation holds for their identification within natural language …

Machine Translation

VarIDE at PARSEME Shared Task 2018: Are Variants Really as Alike as Two Peas in a Pod?

2018-08-01 · COLING 2018 8 · Caroline Pasquer, Carlos Ramisch, Agata Savary, Jean-Yves Antoine

We describe the VarIDE system (standing for Variant IDEntification) which participated in the edition 1.1 of the PARSEME shared task on automatic identification of verbal multiword expressions (VMWEs). Our system focuses…

LIDIOMS: A Multilingual Linked Idioms Data Set

2018-02-22 · LREC 2018 5 · Diego Moussallem, Mohamed Ahmed Sherif, Diego Esteves, Marcos Zampieri 외

In this paper, we describe the LIDIOMS data set, a multilingual RDF representation of idioms currently containing five languages: English, German, Italian, Portuguese, and Russian. The data set is intended to support nat…

Multilingual Idioms in Sentences and Conversations Across High-, Medium-, and Low-Resource Languages

2026-06-01 · Saeed Almheiri, Bilal Elbouardi, Salsabila Zahirah Pranida, Irina Nikishina 외 arxiv

Idiomatic expressions pose a major challenge for multilingual NLP because their meanings shift between figurative and literal usage, often requiring context for accurate interpretation. Prior work has focused on high-res…

Idiom Type Identification with Smoothed Lexical Features and a Maximum Margin Classifier

2017-09-01 · RANLP 2017 9 · Giancarlo Salton, Robert Ross, John Kelleher

In our work we address limitations in the state-of-the-art in idiom type identification. We investigate different approaches for a lexical fixedness metric, a component of the state-of the-art model. We also show that ou…

BIG-bench Machine LearningVocal Bursts Type Prediction