paper-with-me

Papers

Handling Idioms in Symbolic Multilingual Natural Language Generation

2022-06-01 · LREC (MWE) 2022 6 · Michaelle Dubé, François Lareau

While idioms are usually very rigid in their expression, they sometimes allow a certain level of freedom in their usage, with modifiers or complements splitting them or being syntactically attached to internal nodes rather than to the root (e.g., “take something with a big grain of salt”). This means that they cannot always be handled as ready-made strings in rule-based natural language generation systems. Having access to the internal syntactic structure of an idiom allows for more subtle processing. We propose a way to enumerate all possible language-independent n-node trees and to map particular idioms of a language onto these generic syntactic patterns. Using this method, we integrate the idioms from the LN-fr into GenDR, a multilingual realizer. Our implementation covers nearly 98% of LN-fr’s idioms with high precision, and can easily be extended or ported to other languages.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Text Generation

Similar Papers 제목 키워드 기반

LIDIOMS: A Multilingual Linked Idioms Data Set

2018-02-22 · LREC 2018 5 · Diego Moussallem, Mohamed Ahmed Sherif, Diego Esteves, Marcos Zampieri 외

In this paper, we describe the LIDIOMS data set, a multilingual RDF representation of idioms currently containing five languages: English, German, Italian, Portuguese, and Russian. The data set is intended to support nat…

ID10M: Idiom Identification in 10 Languages

2022-07-01 · Findings (NAACL) 2022 7 · Simone Tedeschi, Federico Martelli, Roberto Navigli

Idioms are phrases which present a figurative meaning that cannot be (completely) derived by looking at the meaning of their individual components.Identifying and understanding idioms in context is a crucial goal and a k…

Natural Language Understanding

Translation Of Telugu-Marathi and Vice-Versa using Rule Based Machine Translation

2014-06-16 · Siddhartha Ghosh, Sujata Thamke, Kalyani U. R. S

In todays digital world automated Machine Translation of one language to another has covered a long way to achieve different kinds of success stories. Whereas Babel Fish supports a good number of foreign languages and on…

Machine TranslationTranslation

Xmodel-1.5: An 1B-scale Multilingual LLM

2024-11-15 · Wang Qun, Liu Yang, Lin Qingquan, Jiang Ling

We introduce Xmodel-1.5, a 1-billion-parameter multilingual large language model pretrained on 2 trillion tokens, designed for balanced performance and scalability. Unlike most large models that use the BPE tokenizer, Xm…

Language ModelingLanguage ModellingLarge Language Model

The Mediomatix Corpus: Parallel Data for Romansh Language Varieties via Comparable Schoolbooks

2025-08-22 · Zachary Hopton, Jannis Vamvas, Andrin Büchler, Anna Rutkiewicz 외 arxiv

The five idioms (i.e., varieties) of the Romansh language are largely standardized and are taught in the schools of the respective communities in Switzerland. In this paper, we present the first parallel corpus of Romans…

Machine Translation