paper-with-me

홈 › Papers

Reading-Time Annotations for ``Balanced Corpus of Contemporary Written Japanese''

2016-12-01 · COLING 2016 12 · Masayuki Asahara, Hajime Ono, Edson T. Miyamoto

The Dundee Eyetracking Corpus contains eyetracking data collected while native speakers of English and French read newspaper editorial articles. Similar resources for other languages are still rare, especially for languages in which words are not overtly delimited with spaces. This is a report on a project to build an eyetracking corpus for Japanese. Measurements were collected while 24 native speakers of Japanese read excerpts from the Balanced Corpus of Contemporary Written Japanese Texts were presented with or without segmentation (i.e. with or without space at the boundaries between bunsetsu segmentations) and with two types of methodologies (eyetracking and self-paced reading presentation). Readers{'} background information including vocabulary-size estimation and Japanese reading-span score were also collected. As an example of the possible uses for the corpus, we also report analyses investigating the phenomena of anti-locality.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Articles

Similar Papers 제목 키워드 기반

Between Reading Time and Syntactic/Semantic Categories

2017-11-01 · IJCNLP 2017 11 · Masayuki Asahara, Sachi Kato

This article presents a contrastive analysis between reading time and syntactic/semantic categories in Japanese. We overlaid the reading time annotation of BCCWJ-EyeTrack and a syntactic/semantic category information ann…

UD-Japanese BCCWJ: Universal Dependencies Annotation for the Balanced Corpus of Contemporary Written Japanese

2018-11-01 · WS 2018 11 · Mai Omura, Masayuki Asahara

In this paper, we describe a corpus UD Japanese-BCCWJ that was created by converting the Balanced Corpus of Contemporary Written Japanese (BCCWJ), a Japanese language corpus, to adhere to the UD annotation schema. The BC…

Beyond SoNaR: towards the facilitation of large corpus building efforts

2012-05-01 · LREC 2012 5 · Martin Reynaert, Ineke Schuurman, V{\'e}ronique Hoste, Nelleke Oostdijk 외

In this paper we report on the experiences gained in the recent construction of the SoNaR corpus, a 500 MW reference corpus of contemporary, written Dutch. It shows what can realistically be done within the confines of a…

Annotation of ‘Word List by Semantic Principles’ Labels for the Balanced Corpus of Contemporary Written Japanese

2018-12-01 · PACLIC 2018 12 · Sachi Kato, Masayuki Asahara, Makoto Yamazaki

Cloze Distillation: Improving Neural Language Models with Human Next-Word Prediction

2020-11-01 · CONLL 2020 · Tiwalayo Eisape, Noga Zaslavsky, Roger Levy

Contemporary autoregressive language models (LMs) trained purely on corpus data have been shown to capture numerous features of human incremental processing. However, past work has also suggested dissociations between co…