paper-with-me

홈 › Papers

Prompting Video-Language Foundation Models with Domain-specific Fine-grained Heuristics for Video Question Answering

2024-10-12 · Ting Yu, Kunhao Fu, Shuhui Wang, Qingming Huang, Jun Yu

Video Question Answering (VideoQA) represents a crucial intersection between video understanding and language processing, requiring both discriminative unimodal comprehension and sophisticated cross-modal interaction for accurate inference. Despite advancements in multi-modal pre-trained models and video-language foundation models, these systems often struggle with domain-specific VideoQA due to their generalized pre-training objectives. Addressing this gap necessitates bridging the divide between broad cross-modal knowledge and the specific inference demands of VideoQA tasks. To this end, we introduce HeurVidQA, a framework that leverages domain-specific entity-action heuristics to refine pre-trained video-language foundation models. Our approach treats these models as implicit knowledge engines, employing domain-specific entity-action prompters to direct the model's focus toward precise cues that enhance reasoning. By delivering fine-grained heuristics, we improve the model's ability to identify and interpret key entities and actions, thereby enhancing its reasoning capabilities. Extensive evaluations across multiple VideoQA datasets demonstrate that our method significantly outperforms existing models, underscoring the importance of integrating domain-specific knowledge into video-language models for more accurate and context-aware VideoQA.

📄 PDF Abstract BibTeX arXiv:2410.09380

Code (0)

등록된 구현이 없습니다.

Tasks

Question AnsweringVideo Question AnsweringVideo Understanding

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

SASVi - Segment Any Surgical Video

2025-02-12 · Ssharvien Kumar Sivakumar, Yannik Frisch, Amin Ranem, Anirban Mukhopadhyay

Purpose: Foundation models, trained on multitudes of public datasets, often require additional fine-tuning or re-prompting mechanisms to be applied to visually distinct target domains such as surgical videos. Further, wi…

SegmentationVideo SegmentationVideo Semantic Segmentation

Prompts to Summaries: Zero-Shot Language-Guided Video Summarization

2025-06-12 · Mario Barbara, Alaa Maalouf

The explosive growth of video data intensified the need for flexible user-controllable summarization tools that can operate without domain-specific training data. Existing methods either rely on datasets, limiting genera…

GPUQuery focused video summarizationVideo Summarization

Exploring Automated Recognition of Instructional Activity and Discourse from Multimodal Classroom Data

2025-11-26 · Ivo Bueno, Ruikun Hou, Babette Bühler, Tim Fütterer 외 arxiv

Observation of classroom interactions can provide concrete feedback to teachers, but current methods rely on manual annotation, which is resource-intensive and hard to scale. This work explores AI-driven analysis of clas…

Can Language Models Laugh at YouTube Short-form Videos?

2023-10-22 · Dayoon Ko, Sangho Lee, Gunhee Kim

As short-form funny videos on social networks are gaining popularity, it becomes demanding for AI models to understand them for better communication with humans. Unfortunately, previous video humor datasets target specif…

Form

Exploring the Boundaries of GPT-4 in Radiology

2023-10-23 · Qianchu Liu, Stephanie Hyland, Shruthi Bannur, Kenza Bouzid 외

The recent success of general-domain large language models (LLMs) has significantly changed the natural language processing paradigm towards a unified foundation model across domains and applications. In this paper, we f…

Natural Language InferenceSentenceSentence Similarity