Video Segmentation
1개 벤치마크 · 논문 487편 · 이 태스크의 논문 보기 →
Benchmarks
SegTrack v2
Most implemented
SAM 2: Segment Anything in Images and Videos
One-Shot Video Object Segmentation
Mask2Former for Video Instance Segmentation
GraphEcho: Graph-Driven Unsupervised Domain Adaptation for Echocardiogram Video Segmentation
CCNet: Criss-Cross Attention for Semantic Segmentation
YouTube-VOS: Sequence-to-Sequence Video Object Segmentation
Papers
Bootstrapping a 4D LiDAR Annotation Tool from Video Foundation Models
Progress in 4D LiDAR segmentation is bottlenecked by data. Assigning temporally consistent labels across sparse point cloud sequences is costly and hard to scale, and every new task or domain tends to demand fresh dense …
Scene UnderstandingVideo SegmentationMLLM-Assisted Audio VOS: A 3rd Place Report for the MeViS-Audio Track, 8th LSVOS Challenge
In this technical report, we present a training-free framework for audio-guided video object segmentation, which integrates Multimodal Large Language Models (MLLMs) with SAM-based segmentation models. We decompose the ta…
Video Object SegmentationMultimodal ReasoningVideo SegmentationCoarse Indexing, Fine Evidence: Decoupling Temporal Granularity in Long-Video RAG
Graph-based retrieval-augmented generation (RAG) provides a scalable paradigm for long-video understanding, but existing systems typically inherit a fixed temporal granularity from video segmentation when constructing th…
Video SegmentationAnswer GenerationStitch-Inferencer: Enhance Endoscopic Video Segmentation and Tracking via Panoramic Reconstruction
Surgical video understanding is fundamental to navigation systems. Endoscopic perception is often hindered by a limited field-of-view and frequent instrument occlusions, making spatio-temporal context essential for robus…
Video SegmentationSAM-MT: Real-Time Interactive Multi-Target Video Segmentation
Modern Video Object Segmentation (VOS) involves tracking and segmenting user-specified targets. While recent approaches have achieved remarkable performance in single-target scenarios, extending them to multi-target sett…
Video Object SegmentationVideo SegmentationG$^2$TAM: Geometry Grounded Track Anything Model
Human spatial understanding arises from jointly perceiving geometry and semantics, enabling consistent object identification and localization across viewpoints and time. Current video segmentation models depend on explic…
Video Object SegmentationVideo SegmentationSpatial Reasoning3D Reconstruction