paper-with-me

홈 › Papers

Probing the nature of an island constraint with a parsed corpus

2019-07-01 · LILT 2019 7 · Yusuke Kubota, Ai Kubota

This paper presents a case study of the use of the NINJAL Parsed Corpus of Modern Japanese (NPCMJ) for syntactic research. NPCMJ is the first phrase structure-based treebank for Japanese that is specifically designed for application in linguistic (in addition to NLP) research. After discussing some basic methodological issues pertaining to the use of treebanks for theoretical linguistics research, we introduce our case study on the status of the Coordinate Structure Constraint (CSC) in Japanese, showing that NPCMJ enables us to easily retrieve examples that support one of the key claims of Kubota and Lee (2015): that the CSC should be viewed as a pragmatic, rather than a syntactic constraint. The corpus-based study we conducted moreover revealed a previously unnoticed tendency that was highly relevant for further clarifying the principles governing the empirical data in question. We conclude the paper by briefly discussing some further methodological issues brought up by our case study pertaining to the relationship between linguistic research and corpus development.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Active Islanding Detection Using Pulse Compression Probing

2024-06-08 · Nicholas Piaquadio, N. Eva Wu, Morteza Sarailoo

An islanding detection scheme is developed using pulse compression probing (PCP). A state space system realization is taken from the probing output. The nu-gap metric is applied to compare the measured system to fully in…

Lemmatising Verbs in Middle English Corpora: The Benefit of Enriching the Penn-Helsinki Parsed Corpus of Middle English 2 (PPCME2), the Parsed Corpus of Middle English Poetry (PCMEP), and A Parsed Linguistic Atlas of Early Middle English (PLAEME)

2020-05-01 · LREC 2020 5 · Carola Trips, Michael Percillier

This paper describes the lemmatisation of three annotated corpora of Middle English{---}the Penn-Helsinki Parsed Corpus of Middle English 2 (PPCME2), the Parsed Corpus of Middle English Poetry (PCMEP), and A Parsed Lingu…

LEMMA

Parsing Icelandic Al\thingi Transcripts: Parliamentary Speeches as a Genre

2020-05-01 · LREC 2020 5 · Kristj{\'a}n R{\'u}narsson, Einar Freyr Sigur{\dh}sson

We introduce a corpus of transcripts from Al{\th}ingi, the Icelandic parliament. The corpus is syntactically parsed for phrase structure according to the annotation scheme of the Icelandic Parsed Historical Corpus (IcePa…

The Icelandic Parsed Historical Corpus (IcePaHC)

2012-05-01 · LREC 2012 5 · Eir{\'\i}kur R{\"o}gnvaldsson, Anton Karl Ingason, Einar Freyr Sigur{\dh}sson, Joel Wallenberg

We describe the background for and building of IcePaHC, a one million word parsed historical corpus of Icelandic which has just been finished. This corpus which is completely free and open contains fragments of 60 texts …

Surface Statistics of an Unknown Language Indicate How to Parse It

2018-01-01 · TACL 2018 1 · Dingquan Wang, Jason Eisner

We introduce a novel framework for delexicalized dependency parsing in a new language. We show that useful features of the target language can be extracted automatically from an unparsed corpus, which consists only of go…

Dependency ParsingPOS