paper-with-me

홈 › Papers

Word Familiarity and Frequency

2018-06-09 · Kumiko Tanaka-Ishii, Hiroshi Terada

Word frequency is assumed to correlate with word familiarity, but the strength of this correlation has not been thoroughly investigated. In this paper, we report on our analysis of the correlation between a word familiarity rating list obtained through a psycholinguistic experiment and the log-frequency obtained from various corpora of different kinds and sizes (up to the terabyte scale) for English and Japanese. Major findings are threefold: First, for a given corpus, familiarity is necessary for a word to achieve high frequency, but familiar words are not necessarily frequent. Second, correlation increases with the corpus data size. Third, a corpus of spoken language correlates better than one of written language. These findings suggest that cognitive familiarity ratings are correlated to frequency, but more highly to that of spoken rather than written language.

📄 PDF Abstract BibTeX arXiv:1806.03431

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Dispersion Measures as Predictors of Lexical Decision Time, Word Familiarity, and Lexical Complexity

2025-01-11 · Adam Nohejl, Taro Watanabe

Various measures of dispersion have been proposed to paint a fuller picture of a word's distribution in a corpus, but only little has been done to validate them externally. We evaluate a wide range of dispersion measures…

What makes a word hard to learn? Modeling L1 influence on English vocabulary difficulty

2026-05-12 · Jonas Mayer Martins, Zhuojing Huang, Aaricia Herygers, Lisa Beinborn arxiv

What makes a word difficult to learn, and how does the difficulty depend on the learner's native language? We computationally model vocabulary difficulty for English learners whose first language is Spanish, German, or C…

Beyond Film Subtitles: Is YouTube the Best Approximation of Spoken Vocabulary?

2024-10-04 · Adam Nohejl, Frederikus Hudi, Eunike Andriani Kardinata, Shintaro Ozaki 외

Word frequency is a key variable in psycholinguistics, useful for modeling human familiarity with words even in the era of large language models (LLMs). Frequency in film subtitles has proved to be a particularly good ap…

Lexical Complexity PredictionWord Embeddings

Word Familiarity Rate Estimation Using a Bayesian Linear Mixed Model

2019-11-01 · WS 2019 11 · Masayuki Asahara

This paper presents research on word familiarity rate estimation using the {`}Word List by Semantic Principles{'}. We collected rating information on 96,557 words in the {`}Word List by Semantic Principles{'} via Yahoo! …

Effects of Lexical Properties on Viewing Time per Word in Autistic and Neurotypical Readers

2017-09-01 · WS 2017 9 · Sanja {\v{S}}tajner, Victoria Yaneva, Ruslan Mitkov, Simone Paolo Ponzetto

Eye tracking studies from the past few decades have shaped the way we think of word complexity and cognitive load: words that are long, rare and ambiguous are more difficult to read. However, online processing techniques…

Lexical Simplification