paper-with-me

Papers

SimBase: A Simple Baseline for Temporal Video Grounding

2024-11-12 · Peijun Bao, Alex C. Kot

This paper presents SimBase, a simple yet effective baseline for temporal video grounding. While recent advances in temporal grounding have led to impressive performance, they have also driven network architectures toward greater complexity, with a range of methods to (1) capture temporal relationships and (2) achieve effective multimodal fusion. In contrast, this paper explores the question: How effective can a simplified approach be? To investigate, we design SimBase, a network that leverages lightweight, one-dimensional temporal convolutional layers instead of complex temporal structures. For cross-modal interaction, SimBase only employs an element-wise product instead of intricate multimodal fusion. Remarkably, SimBase achieves state-of-the-art results on two large-scale datasets. As a simple yet powerful baseline, we hope SimBase will spark new ideas and streamline future evaluations in temporal video grounding.

📄 PDF Abstract BibTeX arXiv:2411.07945

Code (0)

등록된 구현이 없습니다.

Tasks

Video Grounding

Similar Papers 제목 키워드 기반

SnAG: Scalable and Accurate Video Grounding

2024-04-02 · CVPR 2024 1 · Fangzhou Mu, Sicheng Mo, Yin Li

Temporal grounding of text descriptions in videos is a central problem in vision-language learning and video understanding. Existing methods often prioritize accuracy over scalability -- they have been optimized for grou…

Video GroundingVideo Understanding

VTimeCoT: Thinking by Drawing for Video Temporal Grounding and Reasoning

2025-10-16 · Jinglei Zhang, Yuanfan Guo, Rolandos Alexandros Potamias, Jiankang Deng 외 arxiv

In recent years, video question answering based on multimodal large language models (MLLM) has garnered considerable attention, due to the benefits from the substantial advancements in LLMs. However, these models have a …

Video Question AnsweringVideo Grounding

DART: Difficulty-Adaptive Routing for Zero-Shot Video Temporal Grounding

2026-07-01 · Zhengbo Zhang, Mark He Huang, Zhigang Tu, Ming-Hsuan Yang arxiv

Zero-shot video temporal grounding (VTG) localizes events in untrimmed videos from natural language queries without task-specific training. Existing methods rely on frame-query feature matching, which suffices for simple…

Natural Language Queries

Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

2024-10-04 · Haibo Wang, Zhiyang Xu, Yu Cheng, Shizhe Diao 외

Video Large Language Models (Video-LLMs) have demonstrated remarkable capabilities in coarse-grained video understanding, however, they struggle with fine-grained temporal grounding. In this paper, we introduce Grounded-…

Dense Video CaptioningSentenceTemporal Sentence GroundingVideo Captioning+1

SynopGround: A Large-Scale Dataset for Multi-Paragraph Video Grounding from TV Dramas and Synopses

2024-08-03 · Chaolei Tan, Zihang Lin, Junfu Pu, Zhongang Qi 외

Video grounding is a fundamental problem in multimodal content understanding, aiming to localize specific natural language queries in an untrimmed video. However, current video grounding datasets merely focus on simple e…

Natural Language QueriesVideo Grounding