Video Object Segmentation
13개 벤치마크 · 논문 608편 · 이 태스크의 논문 보기 →
Benchmarks
DAVIS 2016
DAVIS 2017 (val)
YouTube-VOS 2018
DAVIS 2017 (test-dev)
YouTube-VOS 2019
DAVIS 2017
M$^3$-VOS
DAVIS-2017 (test-dev)
FBMS
YouTube
FBMS-59
MOSE
SegTrack-v2
Most implemented
Emerging Properties in Self-Supervised Vision Transformers
SAM 2: Segment Anything in Images and Videos
One-Shot Video Object Segmentation
Rethinking Self-supervised Correspondence Learning: A Video Frame-level Similarity Perspective
Papers
MLLM-Assisted Audio VOS: A 3rd Place Report for the MeViS-Audio Track, 8th LSVOS Challenge
In this technical report, we present a training-free framework for audio-guided video object segmentation, which integrates Multimodal Large Language Models (MLLMs) with SAM-based segmentation models. We decompose the ta…
Video Object SegmentationMultimodal ReasoningVideo SegmentationSAM2Dual: Training-Free, Dual Memory for Long-Term Video Object Segmentation
Long-term video object segmentation (VOS) remains challenging due to error accumulation under extended occlusions, re-appearance, and scene changes. Although SAM2 provides strong zero-shot performance, its streaming memo…
Video Object SegmentationRRTrack: Robust and Recoverable Object 6D Pose Tracking for Dynamic Scenes
Robust object 6D pose tracking is critical for robotic systems operating in dynamic and occluded scenes. Per-frame estimators are accurate but computationally expensive, while current trackers struggle with fast motion a…
Video Object SegmentationPose TrackingREMIND: RE-Identification with Memory for INDoor Navigation
Mobile robots operating indoors must re-identify previously observed objects after long temporal gaps, significant viewpoint changes, and severe illumination variations. This remains a challenging problem: multi-object t…
Video Object SegmentationVehicle Re-IdentificationMulti-Object TrackingSAM-MT: Real-Time Interactive Multi-Target Video Segmentation
Modern Video Object Segmentation (VOS) involves tracking and segmenting user-specified targets. While recent approaches have achieved remarkable performance in single-target scenarios, extending them to multi-target sett…
Video Object SegmentationVideo Segmentation`Attention-Guided Cross-Temporal Clustering for Self-Supervised Video Object Segmentation
Video object segmentation (VOS) is a fundamental task in video understanding, requiring accurate delineation and consistent tracking of objects across frames. While supervised methods achieve strong performance, they rel…
Video Object SegmentationSelf-Supervised Learning