paper-with-me

Papers

FineQuest: Adaptive Knowledge-Assisted Sports Video Understanding via Agent-of-Thoughts Reasoning

2025-09-15 · Haodong Chen, Haojian Huang, XinXiang Yin, Dian Shao arxiv

Video Question Answering (VideoQA) based on Large Language Models (LLMs) has shown potential in general video understanding but faces significant challenges when applied to the inherently complex domain of sports videos. In this work, we propose FineQuest, the first training-free framework that leverages dual-mode reasoning inspired by cognitive science: i) Reactive Reasoning for straightforward sports queries and ii) Deliberative Reasoning for more complex ones. To bridge the knowledge gap between general-purpose models and domain-specific sports understanding, FineQuest incorporates SSGraph, a multimodal sports knowledge scene graph spanning nine sports, which encodes both visual instances and domain-specific terminology to enhance reasoning accuracy. Furthermore, we introduce two new sports VideoQA benchmarks, Gym-QA and Diving-QA, derived from the FineGym and FineDiving datasets, enabling diverse and comprehensive evaluation. FineQuest achieves state-of-the-art performance on these benchmarks as well as the existing SPORTU dataset, while maintains strong general VideoQA capabilities.

📄 PDF Abstract BibTeX arXiv:2509.11796

Code (0)

등록된 구현이 없습니다.

Tasks

Video Question Answering

Similar Papers 제목 키워드 기반

MuLMINet: Multi-Layer Multi-Input Transformer Network with Weighted Loss

2023-07-17 · Minwoo Seong, Jeongseok Oh, SeungJun Kim

The increasing use of artificial intelligence (AI) technology in turn-based sports, such as badminton, has sparked significant interest in evaluating strategies through the analysis of match video data. Predicting future…

RefereeBench: Are Video MLLMs Ready to be Multi-Sport Referees

2026-04-17 · Yichen Xu, Yuanhang Liu, Chuhan Wang, Zihan Zhao 외 arxiv

While Multimodal Large Language Models (MLLMs) excel at generic video understanding, their ability to support specialized, rule-grounded decision-making remains insufficiently explored. In this paper, we introduce Refere…

MeMoSORT: Memory-Assisted Filtering and Motion-Adaptive Association Metric for Multi-Person Tracking

2025-08-13 · Yingjie Wang, Zhixing Wang, Le Zheng, Tianxiao Liu 외 arxiv

Multi-object tracking (MOT) in human-dominant scenarios, which involves continuously tracking multiple people within video sequences, remains a significant challenge in computer vision due to targets' complex motion and …

Multi-Object Tracking

Subjective and Objective Quality Assessment of High-Motion Sports Videos at Low-Bitrates

2022-07-12 · Joshua P. Ebenezer, Yixu Chen, Yongjun Wu, Hai Wei 외

Videos often have to be transmitted and stored at low bitrates due to poor network connectivity during adaptive bitrate streaming. Designing optimal bitrate ladders that would select the perceptually-optimized resolution…

Video Quality AssessmentVisual Question Answering (VQA)

Distantly Supervised Semantic Text Detection and Recognition for Broadcast Sports Videos Understanding

2021-10-31 · Avijit Shah, Topojoy Biswas, Sathish Ramadoss, Deven Santosh Shah

Comprehensive understanding of key players and actions in multiplayer sports broadcast videos is a challenging problem. Unlike in news or finance videos, sports videos have limited text. While both action recognition for…

Action RecognitionText DetectionVideo Understanding