paper-with-me

홈 › Papers

Evaluating Small Vision-Language Models on Distance-Dependent Traffic Perception

2025-10-09 · Nikos Theodoridis, Tim Brophy, Reenu Mohandas, Ganesh Sistu, Fiachra Collins, Anthony Scanlan, Ciaran Eising arxiv

Vision-Language Models (VLMs) are becoming increasingly powerful, demonstrating strong performance on a variety of tasks that require both visual and textual understanding. Their strong generalisation abilities make them a promising component for automated driving systems, which must handle unexpected corner cases. However, to be trusted in such safety-critical applications, a model must first possess a reliable perception system. Moreover, since critical objects and agents in traffic scenes are often at a distance, we require systems that are not "shortsighted", i.e., systems with strong perception capabilities at both close (up to 20 meters) and long (30+ meters) range. With this in mind, we introduce Distance-Annotated Traffic Perception Question Answering (DTPQA), the first Visual Question Answering (VQA) benchmark focused solely on perception-based questions in traffic scenes, enriched with distance annotations. By excluding questions that require reasoning, we ensure that model performance reflects perception capabilities alone. Since automated driving hardware has limited processing power and cannot support large VLMs, our study centers on smaller VLMs. More specifically, we evaluate several state-of-the-art (SOTA) small VLMs on DTPQA and show that, despite the simplicity of the questions, these models significantly underperform compared to humans (~60% average accuracy for the best-performing small VLM versus ~85% human performance). However, it is important to note that the human sample size was relatively small, which imposes statistical limitations. We also identify specific perception tasks, such as distinguishing left from right, that remain particularly challenging for these models.

📄 PDF Abstract BibTeX arXiv:2510.08352

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question Answering

Similar Papers 제목 키워드 기반

From Model-centered to Human-Centered: Revision Distance as a Metric for Text Evaluation in LLMs-based Applications

2024-04-10 · Yongqiang Ma, Lizhi Qing, Jiawei Liu, Yangyang Kang 외

Evaluating large language models (LLMs) is fundamental, particularly in the context of practical applications. Conventional evaluation methods, typically designed primarily for LLM development, yield numerical scores tha…

Fréchet Distance for Offline Evaluation of Information Retrieval Systems with Sparse Labels

2024-01-31 · Negar Arabzadeh, Charles L. A. Clarke

The rapid advancement of natural language processing, information retrieval (IR), computer vision, and other technologies has presented significant challenges in evaluating the performance of these systems. One of the ma…

Image GenerationInformation RetrievalRetrievalText to Image Generation+1

Evaluating Multimodal LLMs as Generalist Vision-Language-Action Agents for Drone Control: Commanding, Approaching, Tracking and Searching

2026-09-01 · Jaewoo Park, Minyoung Lee, Sukmin Seo, Moonbin Yim 외 hf

Multimodal Large Language Models (MLLMs) are strong perceivers of images and video. We ask how far that reach extends into acting: dropping an MLLM directly into a drone's control loop, with its entire action space decla…

lex4all: A language-independent tool for building and evaluating pronunciation lexicons for small-vocabulary speech recognition

2014-06-01 · ACL 2014 6 · Anjana Vakil, Max Paulus, Alexis Palmer, Michaela Regneri
speech-recognitionSpeech Recognition

Contrast Sets for Evaluating Language-Guided Robot Policies

2024-06-19 · Abrar Anwar, Rohan Gupta, Jesse Thomason

Robot evaluations in language-guided, real world settings are time-consuming and often sample only a small space of potential instructions across complex scenes. In this work, we introduce contrast sets for robotics as a…

Vision and Language Navigation