paper-with-me

Papers

Annotated Corpora for Word Alignment between Japanese and English and its Evaluation with MAP-based Word Aligner

2012-05-01 · LREC 2012 5 · Tsuyoshi Okita

This paper presents two annotated corpora for word alignment between Japanese and English. We annotated on top of the IWSLT-2006 and the NTCIR-8 corpora. The IWSLT-2006 corpus is in the domain of travel conversation while the NTCIR-8 corpus is in the domain of patent. We annotated the first 500 sentence pairs from the IWSLT-2006 corpus and the first 100 sentence pairs from the NTCIR-8 corpus. After mentioned the annotation guideline, we present two evaluation algorithms how to use such hand-annotated corpora: although one is a well-known algorithm for word alignment researchers, one is novel which intends to evaluate a MAP-based word aligner of Okita et al. (2010b).

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationSentenceWord Alignment

Similar Papers 제목 키워드 기반

Building Large-Scale Japanese Pronunciation-Annotated Corpora for Reading Heteronymous Logograms

2022-06-01 · LREC 2022 6 · Fumikazu Sato, Naoki Yoshinaga, Masaru Kitsuregawa

Although screen readers enable visually impaired people to read written text via speech, the ambiguities in pronunciations of heteronyms cause wrong reading, which has a serious impact on the text understanding. Especial…

Sentence

Original-Transcribed Text Alignment for Manyosyu Written by Old Japanese Language

2016-12-01 · WS 2016 12 · Teruaki Oka, Tomoaki Kono

We are constructing an annotated diachronic corpora of the Japanese language. In part of thiswork, we construct a corpus of Manyosyu, which is an old Japanese poetry anthology. In thispaper, we describe how to align the …

Machine TranslationTranslation

Wikification for Scriptio Continua

2016-05-01 · LREC 2016 5 · Yugo Murawaki, Shinsuke Mori

The fact that Japanese employs scriptio continua, or a writing system without spaces, complicates the first step of an NLP pipeline. Word segmentation is widely used in Japanese language processing, and lexical knowledge…

Segmentation

A Japanese Word Dependency Corpus

2014-05-01 · LREC 2014 5 · Shinsuke Mori, Hideki Ogura, Tetsuro Sasada

In this paper, we present a corpus annotated with dependency relationships in Japanese. It contains about 30 thousand sentences in various domains. Six domains in Balanced Corpus of Contemporary Written Japanese have par…

ArticlesDependency ParsingMachine TranslationSentence+1

Line-a-line: A Tool for Annotating Word-Alignments

2020-05-01 · LREC 2020 5 · Maria Skeppstedt, Magnus Ahltorp, Gunnar Eriksson, Rickard Domeij

We here describe line-a-line, a web-based tool for manual annotation of word-alignments in sentence-aligned parallel corpora. The graphical user interface, which builds on a design template from the Jigsaw system for inv…

Multilingual Word EmbeddingsSentenceWord AlignmentWord Embeddings