Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering
In this paper, we propose a Grid-based Local and Global Area Transcription (Grid-LoGAT) system for Video Question Answering (VideoQA). The system operates in two phases. First, extracting text transcripts from video frames using a Vision-Language Model (VLM). Next, processing questions using these transcripts to generate answers through a Large Language Model (LLM). This design ensures image privacy by deploying the VLM on edge devices and the LLM in the cloud. To improve transcript quality, we propose grid-based visual prompting, which extracts intricate local details from each grid cell and integrates them with global information. Evaluation results show that Grid-LoGAT, using the open-source VLM (LLaVA-1.6-7B) and LLM (Llama-3.1-8B), outperforms state-of-the-art methods with similar baseline models on NExT-QA and STAR-QA datasets with an accuracy of 65.9% and 50.11% respectively. Additionally, our method surpasses the non-grid version by 24 points on localization-based questions we created using NExT-QA. (This paper is accepted by IEEE ICIP 2025.)
Code (0)
등록된 구현이 없습니다.
Tasks
Language ModelingLanguage ModellingLarge Language ModelQuestion AnsweringVideo Question AnsweringVisual PromptingSimilar Papers 제목 키워드 기반
Localization in Aerial Imagery with Grid Maps using LocGAN
In this work, we present LocGAN, our localization approach based on a geo-referenced aerial imagery and LiDAR grid maps. Currently, most self-localization approaches relate the current sensor observations to a map genera…
Deep Representation Learning and Clustering of Traffic Scenarios
Determining the traffic scenario space is a major challenge for the homologation and coverage assessment of automated driving functions. In contrast to current approaches that are mainly scenario-based and rely on expert…
ClusteringRepresentation LearningRetrievalExperience Dependent Formation of Global Coherent Representation of Environment by Grid Cells and Head Direction Cells
The grid firing patterns are thought to provide an efficient intrinsic metric capable of supporting universal spatial metric for mammalian spatial navigation in all environments. However, whether spatial representations …
HippocampusSimultaneous Localization and MappingosmAG-Nav: A Hierarchical Semantic Topometric Navigation Stack for Robust Lifelong Indoor Autonomy
The deployment of mobile robots in large-scale, multi-floor environments demands navigation systems that achieve spatial scalability without compromising local kinematic precision. Traditional navigation stacks, reliant …
Coherency and Online Signal Selection Based Wide Area Control of Wind Integrated Power Grid
This paper introduces a novel method of designing wide area control (WAC) based on a discrete linear quadratic regulator and Kalman filtering based state-estimation that can be applied for real-time damping of interarea …
State Estimation