paper-with-me

Papers

ICC: Quantifying Image Caption Concreteness for Multimodal Dataset Curation

2024-03-02 · Moran Yanuka, Morris Alper, Hadar Averbuch-Elor, Raja Giryes

Web-scale training on paired text-image data is becoming increasingly central to multimodal learning, but is challenged by the highly noisy nature of datasets in the wild. Standard data filtering approaches succeed in removing mismatched text-image pairs, but permit semantically related but highly abstract or subjective text. These approaches lack the fine-grained ability to isolate the most concrete samples that provide the strongest signal for learning in a noisy dataset. In this work, we propose a new metric, image caption concreteness, that evaluates caption text without an image reference to measure its concreteness and relevancy for use in multimodal learning. Our approach leverages strong foundation models for measuring visual-semantic information loss in multimodal representations. We demonstrate that this strongly correlates with human evaluation of concreteness in both single-word and sentence-level texts. Moreover, we show that curation using ICC complements existing approaches: It succeeds in selecting the highest quality samples from multimodal web-scale datasets to allow for efficient training in resource-constrained settings.

📄 PDF Abstract BibTeX arXiv:2403.01306

Code (1)

moranyanuka/icc_code 공식 구현 pytorch

Tasks

Sentence

Similar Papers 제목 키워드 기반

Quantifying the visual concreteness of words and topics in multimodal datasets

2018-04-18 · NAACL 2018 6 · Jack Hessel, David Mimno, Lillian Lee

Multimodal machine learning algorithms aim to learn visual-textual correspondences. Previous work suggests that concepts with concrete visual manifestations may be easier to learn than concepts with abstract ones. We giv…

BIG-bench Machine LearningImage Captioning

Grounded Concreteness: Human-Like Concreteness Sensitivity in Vision-Language Models

2026-01-26 · Aryan Roy, Zekun Wang, Christopher J. MacLellan arxiv

Do vision--language models (VLMs) develop more human-like sensitivity to linguistic concreteness than text-only large language models (LLMs) when both are evaluated with text-only prompts? We study this question with a c…

Uncovering Visual-Semantic Psycholinguistic Properties from the Distributional Structure of Text Embedding Spac

2025-05-29 · Si Wu, Sebastian Bruch

Imageability (potential of text to evoke a mental image) and concreteness (perceptibility of text) are two psycholinguistic properties that link visual and semantic spaces. It is little surprise that computational method…

Quantifying the Gaps Between Translation and Native Perception in Training for Multimodal, Multilingual Retrieval

2024-10-02 · Kyle Buettner, Adriana Kovashka

There is a scarcity of multilingual vision-language models that properly account for the perceptual differences that are reflected in image captions across languages and cultures. In this work, through a multimodal, mult…

Image CaptioningRetrieval

Asymmetric Idiosyncrasies in Multimodal Models

2026-02-26 · Muzi Tao, Chufan Shi, Huijuan Wang, Shengbang Tong 외 arxiv

In this work, we study idiosyncrasies in the caption models and their downstream impact on text-to-image models. We design a systematic analysis: given either a generated caption or the corresponding image, we train neur…

Text Classification