paper-with-me

홈 › Papers

Quantifying the visual concreteness of words and topics in multimodal datasets

2018-04-18 · NAACL 2018 6 · Jack Hessel, David Mimno, Lillian Lee

Multimodal machine learning algorithms aim to learn visual-textual correspondences. Previous work suggests that concepts with concrete visual manifestations may be easier to learn than concepts with abstract ones. We give an algorithm for automatically computing the visual concreteness of words and topics within multimodal datasets. We apply the approach in four settings, ranging from image captions to images/text scraped from historical books. In addition to enabling explorations of concepts in multimodal datasets, our concreteness scores predict the capacity of machine learning algorithms to learn textual/visual relationships. We find that 1) concrete concepts are indeed easier to learn; 2) the large number of algorithms we consider have similar failure cases; 3) the precise positive relationship between concreteness and performance varies between datasets. We conclude with recommendations for using concreteness scores to facilitate future multimodal research.

📄 PDF Abstract BibTeX arXiv:1804.06786

Code (1)

victorssilva/concreteness 공식 구현 pytorch

Tasks

BIG-bench Machine LearningImage Captioning

Similar Papers 제목 키워드 기반

ICC: Quantifying Image Caption Concreteness for Multimodal Dataset Curation

2024-03-02 · Moran Yanuka, Morris Alper, Hadar Averbuch-Elor, Raja Giryes

Web-scale training on paired text-image data is becoming increasingly central to multimodal learning, but is challenged by the highly noisy nature of datasets in the wild. Standard data filtering approaches succeed in re…

Sentence

Integrating Vision and Language Datasets to Measure Word Concreteness

2017-11-01 · IJCNLP 2017 11 · Gitit Kehat, James Pustejovsky

We present and take advantage of the inherent visualizability properties of words in visual corpora (the textual components of vision-language datasets) to compute concreteness scores for words. Our simple method does no…

Image CaptioningImage RetrievalQuestion Answering

Grounded Concreteness: Human-Like Concreteness Sensitivity in Vision-Language Models

2026-01-26 · Aryan Roy, Zekun Wang, Christopher J. MacLellan arxiv

Do vision--language models (VLMs) develop more human-like sensitivity to linguistic concreteness than text-only large language models (LLMs) when both are evaluated with text-only prompts? We study this question with a c…

Real Images, Worse Judgments: Evaluating Vision-Language Models on Concreteness and Imagery

2026-05-26 · Yifan Jiang, Ruoxi Ning, Sheng Yao, Freda Shi arxiv

Visual inputs are often assumed to improve language understanding in multimodal models. We examine this assumption by asking whether vision-language models (VLMs) can distinguish useful visual evidence from incidental im…

Color Me Intrigued: Quantifying Usage of Colors in Fiction

2023-01-09 · Siyan Li

We present preliminary results in quantitative analyses of color usage in selected authors' works from LitBank. Using Glasgow Norms, human ratings on 5000+ words, we measure attributes of nouns dependent on color terms. …