Papers Video Instance Segmentation
“Video Instance Segmentation” 태그가 달린 논문 163편 · 필터 해제
UAV-OVVIS: Unmanned Aerial Vehicles Also Need Open-Vocabulary Video Instance Segmentation
Unmanned Aerial Vehicle (UAV) videos are widely used in traffic monitoring, urban management, and emergency rescue. However, existing UAV video perception is largely limited to box-level detection and tracking over prede…
Video Instance SegmentationSegmenting, Fast and Slow: Real-Time Open-Vocabulary Video Instance Segmentation with Dual-Path Processing
Object-centric models inspired by DETR have become the dominant paradigm for open-vocabulary video instance segmentation (OV-VIS). While recent efforts have reduced the computational cost of pixel decoding, textual modal…
Video Instance SegmentationSA-VIS: Sparse frame Annotations for training Video Instance Segmentation
Recent online video instance segmentation (VIS) methods have achieved impressive results, thus becoming the preferred approach to segment instances in videos. Despite the resurgence of impressive single image models, the…
Video Instance SegmentationMind the Gap: Disentangling Performance Bottlenecks in Video Instance Segmentation
In Video Instance Segmentation (VIS), classification, segmentation, and tracking objectives are jointly evaluated, but their individual contributions to performance loss remain opaque. We introduce a diagnostic framework…
Video Instance SegmentationVideo Patch Pruning: Efficient Video Instance Segmentation via Early Token Reduction
Vision Transformers (ViTs) have demonstrated state-ofthe-art performance in several benchmarks, yet their high computational costs hinders their practical deployment. Patch Pruning offers significant savings, but existin…
Video Instance SegmentationSAMannot: A Memory-Efficient, Local, Open-source Framework for Interactive Video Instance Segmentation based on SAM2
Current research workflows for precise video segmentation are often forced into a compromise between labor-intensive manual curation, costly commercial platforms, and/or privacy-compromising cloud-based services. The dem…
Video Instance SegmentationVideo SegmentationS2D: Sparse-To-Dense Keymask Distillation for Unsupervised Video Instance Segmentation
In recent years, the state-of-the-art in unsupervised video instance segmentation has heavily relied on synthetic video data, generated from object-centric image datasets such as ImageNet. However, video synthesis by art…
Unsupervised Instance SegmentationVideo Instance SegmentationBoosting Unsupervised Video Instance Segmentation with Automatic Quality-Guided Self-Training
Video Instance Segmentation (VIS) faces significant annotation challenges due to its dual requirements of pixel-level masks and temporal consistency labels. While recent unsupervised methods like VideoCutLER eliminate op…
Video Instance SegmentationELVIS: Enhance Low-Light for Video Instance Segmentation in the Dark
Video instance segmentation (VIS) for low-light content remains highly challenging for both humans and machines alike, due to noise, blur and other adverse conditions. The lack of large-scale annotated datasets and the l…
Video Instance SegmentationDomain AdaptationAutoQ-VIS: Improving Unsupervised Video Instance Segmentation via Automatic Quality Assessment
Video Instance Segmentation (VIS) faces significant annotation challenges due to its dual requirements of pixel-level masks and temporal consistency labels. While recent unsupervised methods like VideoCutLER eliminate op…
Video Instance SegmentationGeneralized Decoupled Learning for Enhancing Open-Vocabulary Dense Perception
Dense visual perception tasks have been constrained by their reliance on predefined categories, limiting their applicability in real-world scenarios where visual concepts are unbounded. While Vision-Language Models (VLMs…
Video Instance Segmentation3D Instance SegmentationPose EstimationCRISP: Contrastive Residual Injection and Semantic Prompting for Continual Video Instance Segmentation
Continual video instance segmentation demands both the plasticity to absorb new object categories and the stability to retain previously learned ones, all while preserving temporal consistency across frames. In this work…
Video Instance SegmentationContrastive LearningHierarchical Visual Prompt Learning for Continual Video Instance Segmentation
Video instance segmentation (VIS) has gained significant attention for its capability in tracking and segmenting object instances across video frames. However, most of the existing VIS approaches unrealistically assume t…
Video Instance SegmentationLocal2Global query Alignment for Video Instance Segmentation
Online video segmentation methods excel at handling long sequences and capturing gradual changes, making them ideal for real-world applications. However, achieving temporally consistent predictions remains a challenge, e…
Video Instance SegmentationVideo SegmentationLatest Object Memory Management for Temporally Consistent Video Instance Segmentation
In this paper, we present Latest Object Memory Management (LOMM) for temporally consistent video instance segmentation that significantly improves long-term instance tracking. At the core of our method is Latest Object M…
Video Instance SegmentationBeyond Appearance: Geometric Cues for Robust Video Instance Segmentation
Video Instance Segmentation (VIS) fundamentally struggles with pervasive challenges including object occlusions, motion blur, and appearance variations during temporal association. To overcome these limitations, this wor…
Depth EstimationDepth PredictionInstance SegmentationMonocular Depth Estimation+4A Comprehensive Survey on Video Scene Parsing:Advances, Challenges, and Prospects
Video Scene Parsing (VSP) has emerged as a cornerstone in computer vision, facilitating the simultaneous segmentation, recognition, and tracking of diverse visual entities in dynamic scenes. In this survey, we present a …
BenchmarkingInstance SegmentationOpen-Vocabulary Video SegmentationPanoptic Segmentation+8SAM2Auto: Auto Annotation Using FLASH
Vision-Language Models (VLMs) lag behind Large Language Models due to the scarcity of annotated datasets, as creating paired visual-textual annotations is labor-intensive and expensive. To address this bottleneck, we int…
Instance SegmentationObjectobject-detectionObject Detection+5ThinkVideo: High-Quality Reasoning Video Segmentation with Chain of Thoughts
Reasoning Video Object Segmentation is a challenging task, which generates a mask sequence from an input video and an implicit, complex text query. Existing works probe into the problem by finetuning Multimodal Large Lan…
Image SegmentationInstance SegmentationObjectReasoning Video Object Segmentation+6FlowCut: Unsupervised Video Instance Segmentation via Temporal Mask Matching
We propose FlowCut, a simple and capable method for unsupervised video instance segmentation consisting of a three-stage framework to construct a high-quality video dataset with pseudo labels. To our knowledge, our work …
Instance SegmentationSegmentationSemantic SegmentationVideo Instance Segmentation+2