paper-with-me

홈 › Papers

`BonTen' -- Corpus Concordance System for `NINJAL Web Japanese Corpus'

2016-12-01 · COLING 2016 12 · Masayuki Asahara, Kazuya Kawahara, Yuya Takei, Hideto Masuoka, Yasuko Ohba, Yuki Torii, Toru Morii, Yuki Tanaka, Kikuo Maekawa, Sachi Kato, Hikari Konishi

The National Institute for Japanese Language and Linguistics, Japan (NINJAL) has undertaken a corpus compilation project to construct a web corpus for linguistic research comprising ten billion words. The project is divided into four parts: page collection, linguistic analysis, development of the corpus concordance system, and preservation. This article presents the corpus concordance system named {`}BonTen{'} which enables the ten-billion-scaled corpus to be queried by string, a sequence of morphological information or a subtree of the syntactic dependency structure.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Morphological Analysis

Similar Papers 제목 키워드 기반

KOTONOHA: A Corpus Concordance System for Skewer-Searching NINJAL Corpora

2020-05-01 · LREC 2020 5 · Teruaki Oka, Yuichi Ishimoto, Yutaka Yagi, Takenori Nakamura 외

The National Institute for Japanese Language and Linguistics, Japan (NINJAL, Japan), has developed several types of corpora. For each corpus NINJAL provided an online search environment, {`}Chunagon{'}, which is a morpho…

Parsed Corpus as a Source for Testing Generalizations in Japanese Syntax

2019-07-01 · LILT 2019 7 · Hideki Kishimoto, Prashant Pardeshi

In this paper, we discuss constituent ordering generalizations in Japanese. Japanese has SOV as its basic order, but a significant range of argument order variations brought about by ‘scrambling’ is permitted. Although s…

Probing the nature of an island constraint with a parsed corpus

2019-07-01 · LILT 2019 7 · Yusuke Kubota, Ai Kubota

This paper presents a case study of the use of the NINJAL Parsed Corpus of Modern Japanese (NPCMJ) for syntactic research. NPCMJ is the first phrase structure-based treebank for Japanese that is specifically designed for…

Constructing Web-Accessible Semantic Role Labels and Frames for Japanese as Additions to the NPCMJ Parsed Corpus

2020-05-01 · LREC 2020 5 · Koichi Takeuchi, Alastair Butler, Iku Nagasaki, Takuya Okamura 외

As part of constructing the NINJAL Parsed Corpus of Modern Japanese (NPCMJ), a web-accessible language resource, we are adding frame information for predicates, together with two types of semantic role labels that mark t…

AMR ParsingSemantic Parsing

The development of a web corpus of Hindi language and corpus-based comparative studies to Japanese

2016-12-01 · WS 2016 12 · Miki Nishioka, Shiro Akasegawa

In this paper, we discuss our creation of a web corpus of spoken Hindi (COSH), one of the Indo-Aryan languages spoken mainly in the Indian subcontinent. We also point out notable problems we{'}ve encountered in the web c…