paper-with-me

홈 › Papers

Dispersion Measures as Predictors of Lexical Decision Time, Word Familiarity, and Lexical Complexity

2025-01-11 · Adam Nohejl, Taro Watanabe

Various measures of dispersion have been proposed to paint a fuller picture of a word's distribution in a corpus, but only little has been done to validate them externally. We evaluate a wide range of dispersion measures as predictors of lexical decision time, word familiarity, and lexical complexity in five diverse languages. We find that the logarithm of range is not only a better predictor than log-frequency across all tasks and languages, but that it is also the most powerful additional variable to log-frequency, consistently outperforming the more complex dispersion measures. We discuss the effects of corpus part granularity and logarithmic transformation, shedding light on contradictory results of previous studies.

📄 PDF Abstract BibTeX arXiv:2501.06536

Code (1)

naist-nlp/tubelex

Similar Papers 제목 키워드 기반

An experimental and computational study of an Estonian single-person word naming

2025-09-03 · Kaidi Lõo, Arvi Tavast, Maria Heitmeier, Harald Baayen arxiv

This study investigates lexical processing in Estonian. A large-scale single-subject experiment is reported that combines the word naming task with eye-tracking. Five response variables (first fixation duration, total fi…

How trial-to-trial learning shapes mappings in the mental lexicon: Modelling Lexical Decision with Linear Discriminative Learning

2022-07-01 · Maria Heitmeier, Yu-Ying Chuang, R. Harald Baayen

Trial-to-trial effects have been found in a number of studies, indicating that processing a stimulus influences responses in subsequent trials. A special case are priming effects which have been modelled successfully wit…

Additive modelsIncremental Learning

FABRA: French Aggregator-Based Readability Assessment toolkit

2022-06-01 · LREC 2022 6 · Rodrigo Wilkens, David Alfter, Xiaoou Wang, Alice Pintard 외

In this paper, we present the FABRA: readability toolkit based on the aggregation of a large number of readability predictor variables. The toolkit is implemented as a service-oriented architecture, which obviates the ne…

Diversity

Same Error, Different Function: The Optimizer as an Implicit Prior in Financial Time Series

2026-03-03 · Federico Vittorio Cortesi, Giuseppe Iannone, Giulia Crippa, Tomaso Poggio 외 arxiv

Neural networks applied to financial time series operate in a regime of underspecification, where model predictors achieve indistinguishable out-of-sample error. Using large-scale volatility forecasting for S$\&$P 500 st…

A good space: Lexical predictors in word space evaluation

2012-05-01 · LREC 2012 5 · Christian Smith, Henrik Danielsson, Arne J{\"o}nsson

Vector space models benefit from using an outside corpus to train the model. It is, however, unclear what constitutes a good training corpus. We have investigated the effect on summary quality when using various language…

Text CategorizationWord Sense Disambiguation