paper-with-me

Papers

Using large language models to estimate features of multi-word expressions: Concreteness, valence, arousal

2024-08-16 · Gonzalo Martínez, Juan Diego Molero, Sandra González, Javier Conde, Marc Brysbaert, Pedro Reviriego

This study investigates the potential of large language models (LLMs) to provide accurate estimates of concreteness, valence and arousal for multi-word expressions. Unlike previous artificial intelligence (AI) methods, LLMs can capture the nuanced meanings of multi-word expressions. We systematically evaluated ChatGPT-4o's ability to predict concreteness, valence and arousal. In Study 1, ChatGPT-4o showed strong correlations with human concreteness ratings (r = .8) for multi-word expressions. In Study 2, these findings were repeated for valence and arousal ratings of individual words, matching or outperforming previous AI models. Study 3 extended the prevalence and arousal analysis to multi-word expressions and showed promising results despite the lack of large-scale human benchmarks. These findings highlight the potential of LLMs for generating valuable psycholinguistic data related to multiword expressions. To help researchers with stimulus selection, we provide datasets with AI norms of concreteness, valence and arousal for 126,397 English single words and 63,680 multi-word expressions

📄 PDF Abstract BibTeX arXiv:2408.16012

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

To Words and Beyond: Probing Large Language Models for Sentence-Level Psycholinguistic Norms of Memorability and Reading Times

2026-03-12 · Thomas Hikaru Clark, Carlos Arriaga, Javier Conde, Gonzalo Martínez 외 arxiv

Large Language Models (LLMs) have recently been shown to produce estimates of psycholinguistic norms, such as valence, arousal, or concreteness, for words and multiword expressions, that correlate with human judgments. T…

Word Features for Latent Dirichlet Allocation

2010-12-01 · NeurIPS 2010 12 · James Petterson, Wray Buntine, Shravan M. Narayanamurthy, Tibério S. Caetano 외

We extend Latent Dirichlet Allocation (LDA) by explicitly allowing for the encoding of side information in the distribution over words. This results in a variety of new capabilities, such as improved estimates for infreq…

Quantifying the redundancy between prosody and text

2023-11-28 · Lukas Wolf, Tiago Pimentel, Evelina Fedorenko, Ryan Cotterell 외

Prosody -- the suprasegmental component of speech, including pitch, loudness, and tempo -- carries critical aspects of meaning. However, the relationship between the information conveyed by prosody vs. by the words thems…

Word Embeddings

It Means More if It Sounds Good: Yet Another Hypothesis Concerning the Evolution of Polysemous Words

2020-03-12 · Ivan P. Yamshchikov, Cyrille Merleau Nono Saha, Igor Samenko, Jürgen Jost

This position paper looks into the formation of language and shows ties between structural properties of the words in the English language and their polysemy. Using Ollivier-Ricci curvature over a large graph of synonyms…

Position

Temperature-scaling surprisal estimates improve fit to human reading times -- but does it do so for the "right reasons"?

2023-11-15 · Tong Liu, Iza Škrjanec, Vera Demberg

A wide body of evidence shows that human language processing difficulty is predicted by the information-theoretic measure surprisal, a word's negative log probability in context. However, it is still unclear how to best …

Language ModellingLarge Language Model