paper-with-me

Papers

UOUO: Uncontextualized Uncommon Objects for Measuring Knowledge Horizons of Vision Language Models

2024-07-25 · Xinyu Pi, Mingyuan Wu, Jize Jiang, Haozhen Zheng, Beitong Tian, ChengXiang Zhai, Klara Nahrstedt, Zhiting Hu

Smaller-scale Vision-Langauge Models (VLMs) often claim to perform on par with larger models in general-domain visual grounding and question-answering benchmarks while offering advantages in computational efficiency and storage. However, their ability to handle rare objects, which fall into the long tail of data distributions, is less understood. To rigorously evaluate this aspect, we introduce the "Uncontextualized Uncommon Objects" (UOUO) benchmark. This benchmark focuses on systematically testing VLMs with both large and small parameter counts on rare and specialized objects. Our comprehensive analysis reveals that while smaller VLMs maintain competitive performance on common datasets, they significantly underperform on tasks involving uncommon objects. We also propose an advanced, scalable pipeline for data collection and cleaning, ensuring the UOUO benchmark provides high-quality, challenging instances. These findings highlight the need to consider long-tail distributions when assessing the true capabilities of VLMs.

📄 PDF Abstract BibTeX arXiv:2407.18391

Code (0)

등록된 구현이 없습니다.

Tasks

Computational EfficiencyQuestion AnsweringVisual Grounding

Similar Papers 제목 키워드 기반

Generating Uncontextualized and Contextualized Questions for Document-Level Event Argument Extraction

2024-04-07 · Md Nayem Uddin, Enfa Rose George, Eduardo Blanco, Steven Corman

This paper presents multiple question generation strategies for document-level event argument extraction. These strategies do not require human involvement and result in uncontextualized questions as well as contextualiz…

Event Argument ExtractionQuestion GenerationQuestion-Generation

FOCUS: Familiar Objects in Common and Uncommon Settings

2021-10-07 · Priyatham Kattakinda, Soheil Feizi

Standard training datasets for deep learning often contain objects in common settings (e.g., "a horse on grass" or "a ship in water") since they are usually collected by randomly scraping the web. Uncommon and rare setti…

An Attribute-Enriched Dataset and Auto-Annotated Pipeline for Open Detection

2024-09-10 · Pengfei Qi, Yifei Zhang, Wenqiang Li, Youwen Hu 외

Detecting objects of interest through language often presents challenges, particularly with objects that are uncommon or complex to describe, due to perceptual discrepancies between automated models and human annotators.…

AttributeObjectobject-detectionObject Detection

Conformal Disentanglement: A Neural Framework for Perspective Synthesis and Differentiation

2024-08-27 · George A. Kevrekidis, Eleni D. Koronaki, Yannis G. Kevrekidis

For multiple scientific endeavors it is common to measure a phenomenon of interest in more than one ways. We make observations of objects from several different perspectives in space, at different points in time; we may …

Disentanglement

UnCommon Objects in 3D

2025-01-13 · CVPR 2025 1 · Xingchen Liu, Piyush Tayal, Jianyuan Wang, Jesus Zarzar 외

We introduce Uncommon Objects in 3D (uCO3D), a new object-centric dataset for 3D deep learning and 3D generative AI. uCO3D is the largest publicly-available collection of high-resolution videos of objects with 3D annotat…

Object