paper-with-me

홈 › Papers

SFA: Scan, Focus, and Amplify toward Guidance-aware Answering for Video TextVQA

2025-11-25 · Haibin He, Qihuang Zhong, Juhua Liu, Bo Du, Peng Wang, Jing Zhang arxiv

Video text-based visual question answering (Video TextVQA) task aims to answer questions about videos by leveraging the visual text appearing within the videos. This task poses significant challenges, requiring models to accurately perceive and comprehend scene text that varies in scale, orientation, and clarity across frames, while effectively integrating temporal and semantic context to generate precise answers. Moreover, the model must identify question-relevant textual cues and filter out redundant or irrelevant information to ensure answering is guided by the most relevant and informative cues. To address these challenges, we propose SFA, a training-free framework and the first Video-LLM-based method tailored for Video TextVQA, motivated by the human process of answering questions. By adaptively scanning video frames, selectively focusing on key regions, and directly amplifying them, SFA effectively guides the Video-LLM's attention toward essential cues, enabling it to generate more accurate answers. SFA achieves new state-of-the-art results across several public Video TextVQA datasets and surpasses previous methods by a substantial margin, demonstrating its effectiveness and generalizability.

📄 PDF Abstract BibTeX arXiv:2511.20190

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question Answering

Similar Papers 제목 키워드 기반

Heterogeneity-Adaptive Diffusion Schrodinger Bridge for PET-Guided Whole-Body MRI Translation

2026-07-08 · Chengbo Wang, Jiacheng Yu, Linjie Bian, Ming Qi 외 arxiv

While whole-body multimodal medical imaging scanners have been increasingly recognized for more effective medical applications, the excessive long acquisition time in PET-MR scanning is a major obstacle in more efficient…

Predicting Human Scanpaths in Visual Question Answering

2021-06-19 · CVPR 2021 1 · Xianyu Chen, Ming Jiang, Qi Zhao

Attention has been an important mechanism for both humans and computer vision systems. While state-of-the-art models to predict attention focus on estimating a static probabilistic saliency map with free-viewing beha…

Deep Reinforcement LearningQuestion AnsweringScanpath predictionTemporal Sequences+2

TopoMamba: Topology-Aware Scanning and Fusion for Segmenting Heterogeneous Medical Visual Media

2026-04-28 · Fuchen Zheng, Chengpei Xu, Long Ma, Weixuan Li 외 arxiv

Visual state-space models (SSMs) have shown strong potential for medical image segmentation, yet their effectiveness is often limited by two practical issues: axis-biased scan ordering weakens the modeling of oblique and…

Medical Image Segmentation

Multimodal-GuideNet: Gaze-Probe Bidirectional Guidance in Obstetric Ultrasound Scanning

2022-07-26 · Qianhui Men, Clare Teng, Lior Drukker, Aris T. Papageorghiou 외

Eye trackers can provide visual guidance to sonographers during ultrasound (US) scanning. Such guidance is potentially valuable for less experienced operators to improve their scanning skills on how to manipulate the pro…

Motion Magnification in Robotic Sonography: Enabling Pulsation-Aware Artery Segmentation

2023-07-07 · Dianye Huang, Yuan Bi, Nassir Navab, Zhongliang Jiang

Ultrasound (US) imaging is widely used for diagnosing and monitoring arterial diseases, mainly due to the advantages of being non-invasive, radiation-free, and real-time. In order to provide additional information to ass…

Motion MagnificationSegmentation