Context Matters: Recovering Human Semantic Structure from Machine Learning Analysis of Large-Scale Text Corpora
Applying machine learning algorithms to large-scale, text-based corpora (embeddings) presents a unique opportunity to investigate at scale how human semantic knowledge is organized and how people use it to judge fundamental relationships, such as similarity between concepts. However, efforts to date have shown a substantial discrepancy between algorithm predictions and empirical judgments. Here, we introduce a novel approach of generating embeddings motivated by the psychological theory that semantic context plays a critical role in human judgments. Specifically, we train state-of-the-art machine learning algorithms using contextually-constrained text corpora and show that this greatly improves predictions of similarity judgments and feature ratings. By improving the correspondence between representations derived using embeddings generated by machine learning methods and empirical measurements of human judgments, the approach we describe helps advance the use of large-scale text corpora to understand the structure of human semantic representations.
Code (0)
등록된 구현이 없습니다.
Tasks
BIG-bench Machine LearningEmpirical JudgmentsSimilar Papers 제목 키워드 기반
Social Structure Matters in 3D Human-Human Interaction Generation
Although text-to-motion generation has achieved strong progress in synthesizing realistic single-person motions from language, extending it to text-driven 3D human-human interaction (HHI) remains non-trivial, as HHI requ…
Take a Prior from Other Tasks for Severe Blur Removal
Recovering clear structures from severely blurry inputs is a challenging problem due to the large movements between the camera and the scene. Although some works apply segmentation maps on human face images for deblurrin…
DeblurringImage DeblurringKnowledge DistillationWhen Polysemy Matters: Modeling Semantic Categorization with Word Embeddings
Recent work using word embeddings to model semantic categorization have indicated that static models outperform the more recent contextual class of models (Majewska et al, 2021). In this paper, we consider polysemy as a …
Word EmbeddingsSemantic projection: recovering human knowledge of multiple, distinct object features from word embeddings
The words of a language reflect the structure of the human mind, allowing us to transmit thoughts between individuals. However, language can represent only a subset of our rich and detailed cognitive architecture. Here, …
Word EmbeddingsCharacterizing Model-Native Skills
Skills are a natural unit for describing what a language model can do and how its behavior can be changed. However, existing characterizations rely on human-written taxonomies, textual descriptions, or manual profiling p…