paper-with-me

Papers

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering

2025-05-30 · Md Intisar Chowdhury, Kittinun Aukkapinyo, Hiroshi Fujimura, Joo Ann Woo, Wasu Wasusatein, Fadoua Ghourabi

In this paper, we propose a Grid-based Local and Global Area Transcription (Grid-LoGAT) system for Video Question Answering (VideoQA). The system operates in two phases. First, extracting text transcripts from video frames using a Vision-Language Model (VLM). Next, processing questions using these transcripts to generate answers through a Large Language Model (LLM). This design ensures image privacy by deploying the VLM on edge devices and the LLM in the cloud. To improve transcript quality, we propose grid-based visual prompting, which extracts intricate local details from each grid cell and integrates them with global information. Evaluation results show that Grid-LoGAT, using the open-source VLM (LLaVA-1.6-7B) and LLM (Llama-3.1-8B), outperforms state-of-the-art methods with similar baseline models on NExT-QA and STAR-QA datasets with an accuracy of 65.9% and 50.11% respectively. Additionally, our method surpasses the non-grid version by 24 points on localization-based questions we created using NExT-QA. (This paper is accepted by IEEE ICIP 2025.)

📄 PDF Abstract BibTeX arXiv:2505.24371

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingLarge Language ModelQuestion AnsweringVideo Question AnsweringVisual Prompting

Similar Papers 제목 키워드 기반

Localization in Aerial Imagery with Grid Maps using LocGAN

2019-06-04 · Haohao Hu, Junyi Zhu, Sascha Wirges, Martin Lauer

In this work, we present LocGAN, our localization approach based on a geo-referenced aerial imagery and LiDAR grid maps. Currently, most self-localization approaches relate the current sensor observations to a map genera…

Deep Representation Learning and Clustering of Traffic Scenarios

2020-07-15 · Nick Harmening, Marin Biloš, Stephan Günnemann

Determining the traffic scenario space is a major challenge for the homologation and coverage assessment of automated driving functions. In contrast to current approaches that are mainly scenario-based and rely on expert…

ClusteringRepresentation LearningRetrieval

Experience Dependent Formation of Global Coherent Representation of Environment by Grid Cells and Head Direction Cells

2019-10-11

The grid firing patterns are thought to provide an efficient intrinsic metric capable of supporting universal spatial metric for mammalian spatial navigation in all environments. However, whether spatial representations …

HippocampusSimultaneous Localization and Mapping

osmAG-Nav: A Hierarchical Semantic Topometric Navigation Stack for Robust Lifelong Indoor Autonomy

2026-03-30 · Yongqi Zhang, Jiajie Zhang, Chengqian Li, Fujing Xie 외 arxiv

The deployment of mobile robots in large-scale, multi-floor environments demands navigation systems that achieve spatial scalability without compromising local kinematic precision. Traditional navigation stacks, reliant …

Coherency and Online Signal Selection Based Wide Area Control of Wind Integrated Power Grid

2019-07-16

This paper introduces a novel method of designing wide area control (WAC) based on a discrete linear quadratic regulator and Kalman filtering based state-estimation that can be applied for real-time damping of interarea …

State Estimation