paper-with-me

Papers

GenCeption: Evaluate Multimodal LLMs with Unlabeled Unimodal Data

2024-02-22 · Lele Cao, Valentin Buchner, Zineb Senane, Fangkai Yang

Multimodal Large Language Models (MLLMs) are typically assessed using expensive annotated multimodal benchmarks, which often lag behind the rapidly evolving demands of MLLM evaluation. This paper outlines and validates GenCeption, a novel, annotation-free evaluation method that requires only unimodal data to measure inter-modality semantic coherence and inversely assesses MLLMs' tendency to hallucinate. This approach eliminates the need for costly data annotation, minimizes the risk of training data contamination, results in slower benchmark saturation, and avoids the illusion of emerging abilities. Inspired by the DrawCeption game, GenCeption begins with a non-textual sample and proceeds through iterative description and generation steps. The semantic drift across iterations is quantified using the GC@T metric. Based on the GenCeption method, we establish the MMECeption benchmark for evaluating Vision LLMs (VLLMs), and compare performance of several popular VLLMs and human annotators. Our empirical results validate GenCeption's effectiveness, demonstrating strong correlations with established VLLM benchmarks. VLLMs still significantly lack behind human performance and struggle especially with text-intensive tasks.

📄 PDF Abstract BibTeX arXiv:2402.14973

Code (1)

llcresearch/genception 공식 구현

Tasks

Benchmarking

Similar Papers 제목 키워드 기반

MESEN: Exploit Multimodal Data to Design Unimodal Human Activity Recognition with Few Labels

2024-04-02 · Lilin Xu, Chaojie Gu, Rui Tan, Shibo He 외

Human activity recognition (HAR) will be an essential function of various emerging applications. However, HAR typically encounters challenges related to modality limitations and label scarcity, leading to an application …

Activity RecognitionContrastive LearningHuman Activity Recognition

Multimodal Knowledge Expansion

2021-03-26 · ICCV 2021 10 · Zihui Xue, Sucheng Ren, Zhengqi Gao, Hang Zhao

The popularity of multimodal sensors and the accessibility of the Internet have brought us a massive amount of unlabeled multimodal data. Since existing datasets and well-trained models are primarily unimodal, the modali…

DenoisingKnowledge DistillationSemantic Segmentation

Video Generation Models are General-Purpose Vision Learners

2026-07-10 · Letian Wang, Chuhan Zhang, Rishabh Kabra, Jasper Uijlings 외 arxiv

Driven by next-token prediction, NLP shifted from task-specific models into powerful generalist foundation models. What, then, is the equivalent catalyst needed to achieve a general-purpose model in computer vision? In t…

Text-to-Video GenerationCamera Pose Estimation

Instruction-Tuned Video-Audio Models Elucidate Functional Specialization in the Brain

2025-06-09 · Subba Reddy Oota, Khushbu Pahwa, Prachi Jindal, Satya Sai Srinath Namburi 외

Recent voxel-wise multimodal brain encoding studies have shown that multimodal large language models (MLLMs) exhibit a higher degree of brain alignment compared to unimodal models in both unimodal and multimodal stimulus…

Disentanglement

Improving Unimodal Inference with Multimodal Transformers

2023-11-16 · Kateryna Chumachenko, Alexandros Iosifidis, Moncef Gabbouj

This paper proposes an approach for improving performance of unimodal models with multimodal training. Our approach involves a multi-branch architecture that incorporates unimodal models with a multimodal transformer-bas…

Emotion RecognitionGesture RecognitionHand Gesture RecognitionHand-Gesture Recognition+1