paper-with-me

홈 › Papers

VTAgent: Agentic Keyframe Anchoring for Evidence-Aware Video TextVQA

2026-05-06 · Haibin He, Maoyuan Ye, Jing Zhang, Juhua Liu, Bo Du arxiv

Video text-based visual question answering (Video TextVQA) aims to answer questions by reasoning over visual textual content appearing in videos. Despite the strong multimodal video understanding capabilities of recent Video-LLMs, their performance on existing Video TextVQA benchmarks remains limited. To better understand this gap, we conduct an upper-bound analysis through frame-wise question answering, counting a sample as correct if any frame yields the right answer, which significantly outperforms direct video-based inference and reveals a substantial performance gap. The results suggest that the primary bottleneck lies in the localization of key question-relevant evidence, rather than in reasoning capacity itself. Building on this insight, we propose a question-guided agent framework that explicitly anchors the relevant keyframes before answering. The approach operates effectively in a training-free setting and consistently surpasses direct video inference. With additional supervised fine-tuning (SFT) and reinforcement learning (RL), it achieves an average improvement of +12.12 in accuracy and +11.15 in ANLS across benchmarks, establishing new state-of-the-art results. Our study underscores the critical role of explicit keyframe anchoring for advancing Video TextVQA. The code will be publicly released.

📄 PDF Abstract BibTeX arXiv:2605.04870

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question AnsweringReinforcement Learning

Similar Papers 제목 키워드 기반

PathoSage: Towards Multi-Source Evidence Adjudication in Pathology via Experience-Aware Agentic Workflow

2026-05-18 · Chengyang Zhang, Wenchuan Zhang, Bo Li, Mengran Li 외 arxiv

Recent advances in Multimodal Large Language Models (MLLMs) and agent workflows have shown strong promise for computational pathology, yet reliable patch-level reasoning remains challenging. End-to-end pathology MLLMs of…

Multimodal Reasoning

Lookahead Anchoring: Preserving Character Identity in Audio-Driven Human Animation

2025-10-27 · Junyoung Seo, Rodrigo Mira, Alexandros Haliassos, Stella Bounareli 외 arxiv

Audio-driven human animation models often suffer from identity drift during temporal autoregressive generation, where characters gradually lose their identity over time. One solution is to generate keyframes as intermedi…

MMSkills: Towards Multimodal Skills for General Visual Agents

2026-05-13 · Kangning Zhang, Shuai Shao, Qingyao Li, Jianghao Lin 외 arxiv

Reusable skills have become a core substrate for improving agent capabilities, yet most existing skill packages encode reusable behavior primarily as textual prompts, executable code, or learned routines. For visual agen…

Visual GroundingDecision Making

Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction

2026-06-28 · Sunqi Fan, Qingle Liu, Runqi Yin, Meng-Hao Guo 외 arxiv

Video understanding is a fundamental capability for multimodal intelligence, and recent Multimodal Large Language Models (MLLMs) have achieved remarkable performance on Video Question Answering (VideoQA) benchmarks. Howe…

Video Question Answering

RaMem: Contextual Reinstatement for Long-term Agentic Memory

2026-06-22 · Wei Yang, Bryce Kan, Shixuan Li, Li Li 외 arxiv

Long-term memory has become increasingly important for LLM agents that operate across extended interactions and evolving task contexts. Recent memory systems have made past experiences more persistent, compact, and retri…