paper-with-me

홈 › Papers

A Picture is Worth More Than 77 Text Tokens: Evaluating CLIP-Style Models on Dense Captions

2023-12-14 · CVPR 2024 1 · Jack Urbanek, Florian Bordes, Pietro Astolfi, Mary Williamson, Vasu Sharma, Adriana Romero-Soriano

Curation methods for massive vision-language datasets trade off between dataset size and quality. However, even the highest quality of available curated captions are far too short to capture the rich visual detail in an image. To show the value of dense and highly-aligned image-text pairs, we collect the Densely Captioned Images (DCI) dataset, containing 7805 natural images human-annotated with mask-aligned descriptions averaging above 1000 words each. With precise and reliable captions associated with specific parts of an image, we can evaluate vision-language models' (VLMs) understanding of image content with a novel task that matches each caption with its corresponding subcrop. As current models are often limited to 77 text tokens, we also introduce a summarized version (sDCI) in which each caption length is limited. We show that modern techniques that make progress on standard benchmarks do not correspond with significant improvement on our sDCI based benchmark. Lastly, we finetune CLIP using sDCI and show significant improvements over the baseline despite a small training set. By releasing the first human annotated dense image captioning dataset, we hope to enable the development of new benchmarks or fine-tuning recipes for the next generation of VLMs to come.

📄 PDF Abstract BibTeX arXiv:2312.08578

Code (1)

facebookresearch/dci 공식 구현

Tasks

Image Captioning

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

Are Some Words Worth More than Others?

2020-10-12 · EMNLP (Eval4NLP) 2020 11 · Shiran Dudy, Steven Bedrick

Current evaluation metrics for language modeling and generation rely heavily on the accuracy of predicted (or generated) words as compared to a reference ground truth. While important, token-level accuracy only captures …

Language ModelingLanguage ModellingPrediction

A Picture is Worth a Thousand Words? An Empirical Study of Aggregation Strategies for Visual Financial Document Retrieval

2026-05-14 · Ho Hung Lim, Yi Yang arxiv

Visual RAG has offered an alternative to traditional RAG. It treats documents as images and uses vision encoders to obtain vision patch tokens. However, hundreds of patch tokens per document create retrieval and storage …

Make A Long Image Short: Adaptive Token Length for Vision Transformers

2021-12-03 · Yichen Zhu, Yuqin Zhu, Jie Du, Yi Wang 외

The vision transformer splits each image into a sequence of tokens with fixed length and processes the tokens in the same way as words in natural language processing. More tokens normally lead to better performance but c…

Action Recognitionimage-classificationImage Classification

Is a Picture Worth Ten Thousand Words in a Review Dataset?

2016-06-23 · Roberto Camacho Barranco, Laura M. Rodriguez, Rebecca Urbina, M. Shahriar Hossain

While textual reviews have become prominent in many recommendation-based systems, automated frameworks to provide relevant visual cues against text reviews where pictures are not available is a new form of task confronte…

TAG

Conical Visual Concentration for Efficient Large Vision-Language Models

2025-01-01 · CVPR 2025 1 · Long Xing, Qidong Huang, Xiaoyi Dong, Jiajie Lu 외

In large vision-language models (LVLMs), images serve as inputs that carry a wealth of information. As the idiom "A picture is worth a thousand words" implies, representing a single image in current LVLMs can require…