paper-with-me

Papers

TDBench: Benchmarking Vision-Language Models in Understanding Top-Down Images

2025-04-01 · Kaiyuan Hou, Minghui Zhao, Lilin Xu, Yuang Fan, Xiaofan Jiang

The rapid emergence of Vision-Language Models (VLMs) has significantly advanced multimodal understanding, enabling applications in scene comprehension and visual reasoning. While these models have been primarily evaluated and developed for front-view image understanding, their capabilities in interpreting top-down images have received limited attention, partly due to the scarcity of diverse top-down datasets and the challenges in collecting such data. In contrast, top-down vision provides explicit spatial overviews and improved contextual understanding of scenes, making it particularly valuable for tasks like autonomous navigation, aerial imaging, and spatial planning. In this work, we address this gap by introducing TDBench, a comprehensive benchmark for VLMs in top-down image understanding. TDBench is constructed from public top-down view datasets and high-quality simulated images, including diverse real-world and synthetic scenarios. TDBench consists of visual question-answer pairs across ten evaluation dimensions of image understanding. Moreover, we conduct four case studies that commonly happen in real-world scenarios but are less explored. By revealing the strengths and limitations of existing VLM through evaluation results, we hope TDBench to provide insights for motivating future research. Project homepage: https://github.com/Columbia-ICSL/TDBench

📄 PDF Abstract BibTeX arXiv:2504.03748

Code (1)

columbia-icsl/tdbench 공식 구현

Tasks

Autonomous NavigationBenchmarkingVisual Reasoning

Similar Papers 제목 키워드 기반

ProtDBench: A Unified Benchmark of Protein Binder Design and Evaluation

2026-05-05 · Cong Liu, Milong Ren, Jiaqi Guan, Chengyue Gong 외 arxiv

Recent advances in de novo protein binder design have enabled increasing experimental validation, yet reported in silico metrics remain difficult to interpret or compare across studies due to non-standardized evaluation …

Computational Efficiency

On Learning Representations for Tabular Data Distillation

2025-01-23 · Inwon Kang, Parikshit Ram, Yi Zhou, Horst Samulowitz 외

Dataset distillation generates a small set of information-rich instances from a large dataset, resulting in reduced storage requirements, privacy or copyright risks, and computational costs for downstream modeling, thoug…

Dataset DistillationRepresentation Learning

Harnessing Temporal Databases for Systematic Evaluation of Factual Time-Sensitive Question-Answering in Large Language Models

2025-08-04 · Soyeon Kim, Jindong Wang, Xing Xie, Steven Euijong Whang arxiv

Facts change over time, making it essential for Large Language Models (LLMs) to handle time-sensitive factual knowledge accurately and reliably. Although factual Time-Sensitive Question-Answering (TSQA) tasks have been w…

Benchmarking performance, explainability, and evaluation strategies of vision-language models for surgery: Challenges and opportunities

2025-05-16 · Jiajun Cheng, Xianwu Zhao, Shan Lin

Minimally invasive surgery (MIS) presents significant visual and technical challenges, including surgical instrument classification and understanding surgical action involving instruments, verbs, and anatomical targets. …

Benchmarking

Spatial Reasoning in Foundation Models: Benchmarking Object-Centric Spatial Understanding

2025-09-26 · Vahid Mirjalili, Ramin Giahi, Sriram Kollipara, Akshay Kekuda 외 arxiv

Spatial understanding is a critical capability for vision foundation models. While recent advances in large vision models or vision-language models (VLMs) have expanded recognition capabilities, most benchmarks emphasize…

Relational ReasoningScene UnderstandingSpatial Reasoning