paper-with-me

홈 › Papers

A Japanese Word Dependency Corpus

2014-05-01 · LREC 2014 5 · Shinsuke Mori, Hideki Ogura, Tetsuro Sasada

In this paper, we present a corpus annotated with dependency relationships in Japanese. It contains about 30 thousand sentences in various domains. Six domains in Balanced Corpus of Contemporary Written Japanese have part-of-speech and pronunciation annotation as well. Dictionary example sentences have pronunciation annotation and cover basic vocabulary in Japanese with English sentence equivalent. Economic newspaper articles also have pronunciation annotation and the topics are similar to those of Penn Treebank. Invention disclosures do not have other annotation, but it has a clear application, machine translation. The unit of our corpus is word like other languages contrary to existing Japanese corpora whose unit is phrase called bunsetsu. Each sentence is manually segmented into words. We first present the specification of our corpus. Then we give a detailed explanation about our standard of word dependency. We also report some preliminary results of an MST-based dependency parser on our corpus.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

ArticlesDependency ParsingMachine TranslationSentenceTranslation

Similar Papers 제목 키워드 기반

UD-Japanese BCCWJ: Universal Dependencies Annotation for the Balanced Corpus of Contemporary Written Japanese

2018-11-01 · WS 2018 11 · Mai Omura, Masayuki Asahara

In this paper, we describe a corpus UD Japanese-BCCWJ that was created by converting the Balanced Corpus of Contemporary Written Japanese (BCCWJ), a Japanese language corpus, to adhere to the UD annotation schema. The BC…

On the Definition of Japanese Word

2019-06-24 · Yugo Murawaki

The annotation guidelines for Universal Dependencies (UD) stipulate that the basic units of dependency annotation are syntactic words, but it is not clear what are syntactic words in Japanese. Departing from the long tra…

Dependency Parsing

BCCWJ-DepPara: A Syntactic Annotation Treebank on the `Balanced Corpus of Contemporary Written Japanese'

2016-12-01 · WS 2016 12 · Masayuki Asahara, Yuji Matsumoto

Paratactic syntactic structures are difficult to represent in syntactic dependency tree structures. As such, we propose an annotation schema for syntactic dependency annotation of Japanese, in which coordinate structures…

Dependency Parsing

`BonTen' -- Corpus Concordance System for `NINJAL Web Japanese Corpus'

2016-12-01 · COLING 2016 12 · Masayuki Asahara, Kazuya Kawahara, Yuya Takei, Hideto Masuoka 외

The National Institute for Japanese Language and Linguistics, Japan (NINJAL) has undertaken a corpus compilation project to construct a web corpus for linguistic research comprising ten billion words. The project is divi…

Morphological Analysis

Dependency-Based Relative Positional Encoding for Transformer NMT

2019-09-01 · RANLP 2019 9 · Yutaro Omote, Akihiro Tamura, Takashi Ninomiya

This paper proposes a new Transformer neural machine translation model that incorporates syntactic distances between two source words into the relative position representations of the self-attention mechanism. In particu…

Machine TranslationNMTPositionTranslation