paper-with-me

Papers

CrossWordBench: Evaluating the Reasoning Capabilities of LLMs and LVLMs with Controllable Puzzle Generation

2025-03-30 · Jixuan Leng, Chengsong Huang, Langlin Huang, Bill Yuchen Lin, William W. Cohen, Haohan Wang, Jiaxin Huang

Existing reasoning evaluation frameworks for Large Language Models (LLMs) and Large Vision-Language Models (LVLMs) predominantly either assess text-based reasoning or vision-language understanding capabilities, with limited dynamic interplay between textual and visual constraints. To address this limitation, we introduce CrossWordBench, a benchmark designed to evaluate the reasoning capabilities of both LLMs and LVLMs through the medium of crossword puzzles-a task requiring multimodal adherence to semantic constraints from text-based clues and intersectional constraints from visual grid structures. CrossWordBench leverages a controllable puzzle generation framework that produces puzzles in multiple formats (text and image) and offers different evaluation strategies ranging from direct puzzle solving to interactive modes. Our extensive evaluation of over 20 models reveals that reasoning LLMs outperform non-reasoning models substantially by effectively leveraging crossing-letter constraints. We further demonstrate that LVLMs struggle with the task, showing a strong correlation between their puzzle-solving performance and grid-parsing accuracy. Our findings offer insights into the limitations of the reasoning capabilities of current LLMs and LVLMs, and provide an effective approach for creating multimodal constrained tasks for future evaluations.

📄 PDF Abstract BibTeX arXiv:2504.00043

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

TouchStone: Evaluating Vision-Language Models by Language Models

2023-08-31 · Shuai Bai, Shusheng Yang, Jinze Bai, Peng Wang 외

Large vision-language models (LVLMs) have recently witnessed rapid advancements, exhibiting a remarkable capacity for perceiving, understanding, and processing visual information by connecting visual receptor with large …

Visual Storytelling

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models

2025-05-13 · Pritam Sarkar, Ali Etemad

Despite recent advances in video understanding, the capabilities of Large Video Language Models (LVLMs) to perform video-based causal reasoning remains underexplored, largely due to the absence of relevant and dedicated …

FormMultiple-choiceVideo RecognitionVideo Understanding

ReForm-Eval: Evaluating Large Vision Language Models via Unified Re-Formulation of Task-Oriented Benchmarks

2023-10-04 · Zejun Li, Ye Wang, Mengfei Du, Qingwen Liu 외

Recent years have witnessed remarkable progress in the development of large vision-language models (LVLMs). Benefiting from the strong language backbones and efficient cross-modal alignment strategies, LVLMs exhibit surp…

cross-modal alignment

Beyond the Hype: A dispassionate look at vision-language models in medical scenario

2024-08-16 · Yang Nan, Huichi Zhou, Xiaodan Xing, Guang Yang

Recent advancements in Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities across diverse tasks, garnering significant attention in AI communities. However, their performance and reliability in…

Question AnsweringSpatial ReasoningVisual Question AnsweringVisual Question Answering (VQA)

Coherent Multimodal Reasoning with Iterative Self-Evaluation for Vision-Language Models

2025-08-04 · Wenjie Luo, Ruocheng Li, Shanshan Zhu, Julian Perry arxiv

Despite significant advancements, current large language models (LLMs) and vision-language models (LVLMs) continue to struggle with complex, multi-step, cross-modal common sense reasoning tasks, often exhibiting a lack o…

Common Sense ReasoningMultimodal Reasoning