paper-with-me

Papers

JourneyBench: A Challenging One-Stop Vision-Language Understanding Benchmark of Generated Images

2024-09-19 · Zhecan Wang, Junzhang Liu, Chia-Wei Tang, Hani AlOmari, Anushka Sivakumar, Rui Sun, Wenhao Li, Md. Atabuzzaman, Hammad Ayyubi, Haoxuan You, Alvi Ishmam, Kai-Wei Chang, Shih-Fu Chang, Chris Thomas

Existing vision-language understanding benchmarks largely consist of images of objects in their usual contexts. As a consequence, recent multimodal large language models can perform well with only a shallow visual understanding by relying on background language biases. Thus, strong performance on these benchmarks does not necessarily correlate with strong visual understanding. In this paper, we release JourneyBench, a comprehensive human-annotated benchmark of generated images designed to assess the model's fine-grained multimodal reasoning abilities across five tasks: complementary multimodal chain of thought, multi-image VQA, imaginary image captioning, VQA with hallucination triggers, and fine-grained retrieval with sample-specific distractors. Unlike existing benchmarks, JourneyBench explicitly requires fine-grained multimodal reasoning in unusual imaginary scenarios where language bias and holistic image gist are insufficient. We benchmark state-of-the-art models on JourneyBench and analyze performance along a number of fine-grained dimensions. Results across all five tasks show that JourneyBench is exceptionally challenging for even the best models, indicating that models' visual reasoning abilities are not as strong as they first appear. We discuss the implications of our findings and propose avenues for further research.

📄 PDF Abstract BibTeX arXiv:2409.12953

Code (1)

journeybench/journeybench 공식 구현

Tasks

HallucinationImage CaptioningMultimodal ReasoningVisual Question Answering (VQA)Visual Reasoning

Similar Papers 제목 키워드 기반

Beyond IVR: Benchmarking Customer Support LLM Agents for Business-Adherence

2026-01-02 · Sumanth Balaji, Piyush Mishra, Aashraya Sachdeva, Suraj Agrawal arxiv

Traditional customer support systems, such as Interactive Voice Response (IVR), rely on rigid scripts and lack the flexibility required for handling complex, policy-driven tasks. While large language model (LLM) agents o…

Learning to Stop: A Simple yet Effective Approach to Urban Vision-Language Navigation

2020-09-28 · Findings of the Association for Computational Linguistics 2020 · Jiannan Xiang, Xin Eric Wang, William Yang Wang

Vision-and-Language Navigation (VLN) is a natural language grounding task where an agent learns to follow language instructions and navigate to specified destinations in real-world environments. A key challenge is to rec…

NavigateVision and Language NavigationVision-Language Navigation

STOP: Integrated Spatial-Temporal Dynamic Prompting for Video Understanding

2025-03-20 · CVPR 2025 1 · Zichen Liu, Kunlun Xu, Bing Su, Xu Zou 외

Pre-trained on tremendous image-text pairs, vision-language models like CLIP have demonstrated promising zero-shot generalization across numerous image-based tasks. However, extending these capabilities to video tasks re…

Video UnderstandingZero-shot Generalization

Toward Building General Foundation Models for Language, Vision, and Vision-Language Understanding Tasks

2023-01-12 · Xinsong Zhang, Yan Zeng, Jipeng Zhang, Hang Li

Foundation models or pre-trained models have substantially improved the performance of various language, vision, and vision-language understanding tasks. However, existing foundation models can only perform the best in o…

Cross-Modal RetrievalOpen-Ended Question AnsweringVisual GroundingVisual Question Answering (VQA)+1

ListOps: A Diagnostic Dataset for Latent Tree Learning

2018-04-17 · NAACL 2018 6 · Nikita Nangia, Samuel R. Bowman

Latent tree learning models learn to parse a sentence without syntactic supervision, and use that parse to build the sentence representation. Existing work on such models has shown that, while they perform well on tasks …

DiagnosticListOpsSentenceSentence Classification+1