paper-with-me

Papers

SAMA: Towards Multi-Turn Referential Grounded Video Chat with Large Language Models

2025-05-24 · Ye Sun, Hao Zhang, Henghui Ding, Tiehua Zhang, Xingjun Ma, Yu-Gang Jiang

Achieving fine-grained spatio-temporal understanding in videos remains a major challenge for current Video Large Multimodal Models (Video LMMs). Addressing this challenge requires mastering two core capabilities: video referring understanding, which captures the semantics of video regions, and video grounding, which segments object regions based on natural language descriptions. However, most existing approaches tackle these tasks in isolation, limiting progress toward unified, referentially grounded video interaction. We identify a key bottleneck in the lack of high-quality, unified video instruction data and a comprehensive benchmark for evaluating referentially grounded video chat. To address these challenges, we contribute in three core aspects: dataset, model, and benchmark. First, we introduce SAMA-239K, a large-scale dataset comprising 15K videos specifically curated to enable joint learning of video referring understanding, grounding, and multi-turn video chat. Second, we propose the SAMA model, which incorporates a versatile spatio-temporal context aggregator and a Segment Anything Model to jointly enhance fine-grained video comprehension and precise grounding capabilities. Finally, we establish SAMA-Bench, a meticulously designed benchmark consisting of 5,067 questions from 522 videos, to comprehensively evaluate the integrated capabilities of Video LMMs in multi-turn, spatio-temporal referring understanding and grounded dialogue. Extensive experiments and benchmarking results show that SAMA not only achieves strong performance on SAMA-Bench but also sets a new state-of-the-art on general grounding benchmarks, while maintaining highly competitive performance on standard visual understanding benchmarks.

📄 PDF Abstract BibTeX arXiv:2505.18812

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingVideo Grounding

Similar Papers 제목 키워드 기반

SPARROW: Learning Spatial Precision and Temporal Referential Consistency in Pixel-Grounded Video MLLMs

2026-03-12 · Mohamad Alansari, Naufal Suryanto, Divya Velayudhan, Sajid Javed 외 arxiv

Multimodal large language models (MLLMs) have advanced from image-level reasoning to pixel-level grounding, but extending these capabilities to videos remains challenging as models must achieve spatial precision and temp…

Visual Grounding

GraphDiffs: Graph Modeling with Differential Sequence for Document-Grounded Conversation

2022-01-16 · ACL ARR January 2022 1 · Anonymous

Knowledge grounded dialogue systems need to incorporate natural transitions between knowledge for dialogue to flow smoothly. Current systems not only lack good structured representations for knowledge that span multiple …

SAMA: Factorized Semantic Anchoring and Motion Alignment for Instruction-Guided Video Editing

2026-03-19 · Xinyao Zhang, Wenkai Dong, Yuxin Song, Bo Fang 외 arxiv

Current instruction-guided video editing models struggle to simultaneously balance precise semantic modifications with faithful motion preservation. While existing approaches rely on injecting explicit external priors (e…

Video Restoration

CorefDiffs: Co-referential and Differential Knowledge Flow in Document Grounded Conversations

2022-10-05 · COLING 2022 10 · Lin Xu, Qixian Zhou, Jinlan Fu, Min-Yen Kan 외

Knowledge-grounded dialog systems need to incorporate smooth transitions among knowledge selected for generating responses, to ensure that dialog flows naturally. For document-grounded dialog systems, the inter- and intr…

Management

A Skill-augmented Agentic Framework and Benchmark for Multi-Video Understanding

2026-03-16 · Yue Zhang, Liqiang Jing, Jia Li, Yapeng Tian 외 arxiv

Multimodal Large Language Models have achieved strong performance in single-video understanding, yet their ability to reason across multiple videos remains limited. Existing approaches typically concatenate multiple vide…