paper-with-me

홈 › Papers

Reconstruction as a Bridge for Event-Based Visual Question Answering

2025-12-12 · Hanyue Lou, Jiayi Zhou, Yang Zhang, Boyu Li, Yi Wang, Guangnan Ye, Boxin Shi arxiv

Integrating event cameras with Multimodal Large Language Models (MLLMs) promises general scene understanding in challenging visual conditions, yet requires navigating a trade-off between preserving the unique advantages of event data and ensuring compatibility with frame-based models. We address this challenge by using reconstruction as a bridge, proposing a straightforward Frame-based Reconstruction and Tokenization (FRT) method and designing an efficient Adaptive Reconstruction and Tokenization (ART) method that leverages event sparsity. For robust evaluation, we introduce EvQA, the first objective, real-world benchmark for event-based MLLMs, comprising 1,000 event-Q&A pairs from 22 public datasets. Our experiments demonstrate that our methods achieve state-of-the-art performance on EvQA, highlighting the significant potential of MLLMs in event-based vision.

📄 PDF Abstract BibTeX arXiv:2512.11510

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question AnsweringScene UnderstandingEvent-based vision

Similar Papers 제목 키워드 기반

VGAT: A Cancer Survival Analysis Framework Transitioning from Generative Visual Question Answering to Genomic Reconstruction

2025-03-25 · Zizhi Chen, Minghao Han, Xukun Zhang, Shuwei Ma 외

Multimodal learning combining pathology images and genomic sequences enhances cancer survival analysis but faces clinical implementation barriers due to limited access to genomic sequencing in under-resourced regions. To…

Generative Visual Question AnsweringQuestion AnsweringSurvival AnalysisSurvival Prediction+3

Bridge Damage Cause Estimation Using Multiple Images Based on Visual Question Answering

2023-02-18 · Tatsuro Yamane, Pang-jo Chun, Ji Dang, Takayuki Okatani

In this paper, a bridge member damage cause estimation framework is proposed by calculating the image position using Structure from Motion (SfM) and acquiring its information via Visual Question Answering (VQA). For this…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Bridge to Answer: Structure-aware Graph Interaction Network for Video Question Answering

2021-04-29 · CVPR 2021 1 · Jungin Park, Jiyoung Lee, Kwanghoon Sohn

This paper presents a novel method, termed Bridge to Answer, to infer correct answers for questions about a given video by leveraging adequate graph interactions of heterogeneous crossmodal graphs. To realize this, we le…

Question AnsweringVideo Question Answering

Cross-Modal Causal Relational Reasoning for Event-Level Visual Question Answering

2022-07-26 · Yang Liu, Guanbin Li, Liang Lin

Existing visual question answering methods often suffer from cross-modal spurious correlations and oversimplified event-level reasoning processes that fail to capture event temporality, causality, and dynamics spanning o…

Causal InferenceQuestion AnsweringRelational ReasoningVisual Question Answering+1

MM-Prompt: Cross-Modal Prompt Tuning for Continual Visual Question Answering

2025-05-26 · Xu Li, Fan Lyu

Continual Visual Question Answering (CVQA) based on pre-trained models(PTMs) has achieved promising progress by leveraging prompt tuning to enable continual multi-modal learning. However, most existing methods adopt cros…

Continual LearningQuestion AnsweringVisual Question Answering