paper-with-me

홈 › Papers

Object-centric Video Question Answering with Visual Grounding and Referring

2025-07-25 · Haochen Wang, Qirui Chen, Cilin Yan, Jiayin Cai, Xiaolong Jiang, Yao Hu, Weidi Xie, Stratis Gavves arxiv

Video Large Language Models (VideoLLMs) have recently demonstrated remarkable progress in general video understanding. However, existing models primarily focus on high-level comprehension and are limited to text-only responses, restricting the flexibility for object-centric, multiround interactions. In this paper, we make three contributions: (i) we address these limitations by introducing a VideoLLM model, capable of performing both object referring for input and grounding for output in video reasoning tasks, i.e., allowing users to interact with videos using both textual and visual prompts; (ii) we propose STOM (Spatial-Temporal Overlay Module), a novel approach that propagates arbitrary visual prompts input at any single timestamp to the remaining frames within a video; (iii) we present VideoInfer, a manually curated object-centric video instruction dataset featuring questionanswering pairs that require reasoning. We conduct comprehensive experiments on VideoInfer and other existing benchmarks across video question answering and referring object segmentation. The results on 12 benchmarks of 6 tasks show that our proposed model consistently outperforms baselines in both video question answering and segmentation, underscoring its robustness in multimodal, object-centric video and image understanding. Project page: https://qirui-chen.github.io/RGA3-release/.

📄 PDF Abstract BibTeX arXiv:2507.19599

Code (0)

등록된 구현이 없습니다.

Tasks

Video Question AnsweringObject SegmentationVisual Grounding

Similar Papers 제목 키워드 기반

DMC$^3$: Dual-Modal Counterfactual Contrastive Construction for Egocentric Video Question Answering

2025-10-23 · Jiayi Zou, Chaofan Chen, Bing-Kun Bao, Changsheng Xu arxiv

Egocentric Video Question Answering (Egocentric VideoQA) plays an important role in egocentric video understanding, which refers to answering questions based on first-person videos. Although existing methods have made pr…

Video Question Answering

Video Question Answering for People with Visual Impairments Using an Egocentric 360-Degree Camera

2024-05-30 · Inpyo Song, Minjun Joo, Joonhyung Kwon, JangWon Lee

This paper addresses the daily challenges encountered by visually impaired individuals, such as limited access to information, navigation difficulties, and barriers to social interaction. To alleviate these challenges, w…

Question AnsweringVideo Question AnsweringVisual Question Answering

Object-Centric Representation Learning for Video Question Answering

2021-04-12 · Long Hoang Dang, Thao Minh Le, Vuong Le, Truyen Tran

Video question answering (Video QA) presents a powerful testbed for human-like intelligent behaviors. The task demands new capabilities to integrate video processing, language understanding, binding abstract linguistic c…

ObjectQuestion AnsweringRelational ReasoningRepresentation Learning+1

SuperMemory-VQA: An Egocentric Visual Question-Answering Benchmark for Long-Horizon Memory

2026-05-30 · Samiul Alam, Shakhrul Iman Siam, Michael J. Proulx, James Fort 외 arxiv

AI glasses present a compelling platform for AI agents to serve as personalized memory assistants. To be genuinely useful, such systems must move beyond short-term video comprehension and address memory gaps that humans …

Visual Question AnsweringAction Recognition

HCQA @ Ego4D EgoSchema Challenge 2024

2024-06-22 · Haoyu Zhang, Yuquan Xie, Yisen Feng, Zaijing Li 외

In this report, we present our champion solution for Ego4D EgoSchema Challenge in CVPR 2024. To deeply integrate the powerful egocentric captioning model and question reasoning model, we propose a novel Hierarchical Comp…

Caption GenerationEgoSchemaMultiple-choice+2