paper-with-me

홈 › Papers

Corpus and dictionary development for classifiers/quantifiers towards a French-Japanese machine translation

2016-12-01 · WS 2016 12 · Mutsuko Tomokiyo, Christian Boitet

Although quantifiers/classifiers expressions occur frequently in everyday communications or written documents, there is no description for them in classical bilingual paper dictionaries, nor in machine-readable dictionaries. The paper describes a corpus and dictionary development for quantifiers/classifiers, and their usage in the framework of French-Japanese machine translation (MT). They often cause problems of lexical ambiguity and of set phrase recognition during analysis, in particular for a long-distance language pair like French and Japanese. For the development of a dictionary aiming at ambiguity resolution for expressions including quantifiers and classifiers which may be ambiguous with common nouns, we have annotated our corpus with UWs (interlingual lexemes) of UNL (Universal Networking Language) found on the UNL-jp dictionary. The extraction of potential classifiers/quantifiers from corpus is made by UNLexplorer web service. Keywords : classifiers, quantifiers, phraseology study, corpus annotation, UNL (Universal Networking Language), UWs dictionary, Tori Bank, French-Japanese machine translation (MT).

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationTranslation

Similar Papers 제목 키워드 기반

Development of a classifiers/quantifiers dictionary towards French-Japanese MT

2019-02-21 · MTSummit 2017 9 · Mutsuko Tomokiyo, Mathieu Mangeot, Christian Boitet

Although classifiers/quantifiers (CQs) expressions appear frequently in everyday communications or written documents, they are described neither in classical bilingual paper dictionaries , nor in machine-readable diction…

Machine TranslationTranslation

A new semantically annotated corpus with syntactic-semantic and cross-lingual senses

2026-05-27 · Myriam Rakho, Eric Laporte, Matthieu Constant arxiv

We describe a new sense-tagged corpus for word sense disambiguation. The corpus is constituted of instances of 20 French polysemous verbs. Each verb instance is annotated with three sense labels: (1) the actual translati…

Word Sense Disambiguation

A new semantically annotated corpus with syntactic-semantic and cross-lingual senses

2012-05-01 · LREC 2012 5 · Myriam Rakho, {\'E}ric Laporte, Matthieu Constant

In this article, we describe a new sense-tagged corpus for Word Sense Disambiguation. The corpus is constituted of instances of 20 French polysemous verbs. Each verb instance is annotated with three sense labels: (1) the…

Machine TranslationTranslationWord Sense Disambiguation

MotàMot project: conversion of a French-Khmer published dictionary for building a multilingual lexical system

2014-05-22 · LREC 2014 5 · Mathieu Mangeot

Economic issues related to the information processing techniques are very important. The development of such technologies is a major asset for developing countries like Cambodia and Laos, and emerging ones like Vietnam, …

Mot\`aMot project: conversion of a French-Khmer published dictionary for building a multilingual lexical system

2014-05-01 · LREC 2014 5 · Mathieu Mangeot

Economic issues related to the information processing techniques are very important. The development of such technologies is a major asset for developing countries like Cambodia and Laos, and emerging ones like Vietnam, …