paper-with-me

Papers Video Instance Segmentation

“Video Instance Segmentation” 태그가 달린 논문 163편 · 필터 해제

UAV-OVVIS: Unmanned Aerial Vehicles Also Need Open-Vocabulary Video Instance Segmentation

2026-07-09 · Mingyu Dou, Shi Qiu, Ming Hu, Yifan Chen 외 arxiv

Unmanned Aerial Vehicle (UAV) videos are widely used in traffic monitoring, urban management, and emergency rescue. However, existing UAV video perception is largely limited to box-level detection and tracking over prede…

Video Instance Segmentation

Segmenting, Fast and Slow: Real-Time Open-Vocabulary Video Instance Segmentation with Dual-Path Processing

2026-06-30 · Luca Barsellotti, Martin Sundermeyer, Mattia Segu, Nikita Araslanov 외 arxiv

Object-centric models inspired by DETR have become the dominant paradigm for open-vocabulary video instance segmentation (OV-VIS). While recent efforts have reduced the computational cost of pixel decoding, textual modal…

Video Instance Segmentation

SA-VIS: Sparse frame Annotations for training Video Instance Segmentation

2026-06-18 · Edoardo Mello Rella, Ajad Chhatkuli, Shipra Jain, Ender Konukoglu 외 arxiv

Recent online video instance segmentation (VIS) methods have achieved impressive results, thus becoming the preferred approach to segment instances in videos. Despite the resurgence of impressive single image models, the…

Video Instance Segmentation

Mind the Gap: Disentangling Performance Bottlenecks in Video Instance Segmentation

2026-06-05 · Danial Hamdi, Fardin Ayar, Mahdi Javanmardi arxiv

In Video Instance Segmentation (VIS), classification, segmentation, and tracking objectives are jointly evaluated, but their individual contributions to performance loss remain opaque. We introduce a diagnostic framework…

Video Instance Segmentation

Video Patch Pruning: Efficient Video Instance Segmentation via Early Token Reduction

2026-04-01 · Patrick Glandorf, Thomas Norrenbrock, Bodo Rosenhahn arxiv

Vision Transformers (ViTs) have demonstrated state-ofthe-art performance in several benchmarks, yet their high computational costs hinders their practical deployment. Patch Pruning offers significant savings, but existin…

Video Instance Segmentation

SAMannot: A Memory-Efficient, Local, Open-source Framework for Interactive Video Instance Segmentation based on SAM2

2026-01-16 · Gergely Dinya, András Gelencsér, Krisztina Kupán, Clemens Küpper 외 arxiv

Current research workflows for precise video segmentation are often forced into a compromise between labor-intensive manual curation, costly commercial platforms, and/or privacy-compromising cloud-based services. The dem…

Video Instance SegmentationVideo Segmentation

S2D: Sparse-To-Dense Keymask Distillation for Unsupervised Video Instance Segmentation

2025-12-16 · Leon Sick, Lukas Hoyer, Dominik Engel, Pedro Hermosilla 외 arxiv

In recent years, the state-of-the-art in unsupervised video instance segmentation has heavily relied on synthetic video data, generated from object-centric image datasets such as ImageNet. However, video synthesis by art…

Unsupervised Instance SegmentationVideo Instance Segmentation

Boosting Unsupervised Video Instance Segmentation with Automatic Quality-Guided Self-Training

2025-12-07 · Kaixuan Lu, Mehmet Onurcan Kaya, Dim P. Papadopoulos arxiv

Video Instance Segmentation (VIS) faces significant annotation challenges due to its dual requirements of pixel-level masks and temporal consistency labels. While recent unsupervised methods like VideoCutLER eliminate op…

Video Instance Segmentation

ELVIS: Enhance Low-Light for Video Instance Segmentation in the Dark

2025-12-01 · Joanne Lin, Ruirui Lin, Yini Li, David Bull 외 arxiv

Video instance segmentation (VIS) for low-light content remains highly challenging for both humans and machines alike, due to noise, blur and other adverse conditions. The lack of large-scale annotated datasets and the l…

Video Instance SegmentationDomain Adaptation

AutoQ-VIS: Improving Unsupervised Video Instance Segmentation via Automatic Quality Assessment

2025-08-27 · Kaixuan Lu, Mehmet Onurcan Kaya, Dim P. Papadopoulos arxiv

Video Instance Segmentation (VIS) faces significant annotation challenges due to its dual requirements of pixel-level masks and temporal consistency labels. While recent unsupervised methods like VideoCutLER eliminate op…

Video Instance Segmentation

Generalized Decoupled Learning for Enhancing Open-Vocabulary Dense Perception

2025-08-15 · Junjie Wang, Keyu Chen, Yulin Li, Bin Chen 외 arxiv

Dense visual perception tasks have been constrained by their reliance on predefined categories, limiting their applicability in real-world scenarios where visual concepts are unbounded. While Vision-Language Models (VLMs…

Video Instance Segmentation3D Instance SegmentationPose Estimation

CRISP: Contrastive Residual Injection and Semantic Prompting for Continual Video Instance Segmentation

2025-08-14 · Baichen Liu, Qi Lyu, Xudong Wang, Jiahua Dong 외 arxiv

Continual video instance segmentation demands both the plasticity to absorb new object categories and the stability to retain previously learned ones, all while preserving temporal consistency across frames. In this work…

Video Instance SegmentationContrastive Learning

Hierarchical Visual Prompt Learning for Continual Video Instance Segmentation

2025-08-12 · Jiahua Dong, Hui Yin, Wenqi Liang, Hanbin Zhao 외 arxiv

Video instance segmentation (VIS) has gained significant attention for its capability in tracking and segmenting object instances across video frames. However, most of the existing VIS approaches unrealistically assume t…

Video Instance Segmentation

Local2Global query Alignment for Video Instance Segmentation

2025-07-27 · Rajat Koner, Zhipeng Wang, Srinivas Parthasarathy, Chinghang Chen arxiv

Online video segmentation methods excel at handling long sequences and capturing gradual changes, making them ideal for real-world applications. However, achieving temporally consistent predictions remains a challenge, e…

Video Instance SegmentationVideo Segmentation

Latest Object Memory Management for Temporally Consistent Video Instance Segmentation

2025-07-26 · Seunghun Lee, Jiwan Seo, Minwoo Choi, Kiljoon Han 외 arxiv

In this paper, we present Latest Object Memory Management (LOMM) for temporally consistent video instance segmentation that significantly improves long-term instance tracking. At the core of our method is Latest Object M…

Video Instance Segmentation

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation

2025-07-08 · Quanzhu Niu, Yikang Zhou, Shihao Chen, Tao Zhang 외

Video Instance Segmentation (VIS) fundamentally struggles with pervasive challenges including object occlusions, motion blur, and appearance variations during temporal association. To overcome these limitations, this wor…

Depth EstimationDepth PredictionInstance SegmentationMonocular Depth Estimation+4

A Comprehensive Survey on Video Scene Parsing:Advances, Challenges, and Prospects

2025-06-16 · Guohuan Xie, Syed Ariff Syed Hesham, Wenya Guo, Bing Li 외

Video Scene Parsing (VSP) has emerged as a cornerstone in computer vision, facilitating the simultaneous segmentation, recognition, and tracking of diverse visual entities in dynamic scenes. In this survey, we present a …

BenchmarkingInstance SegmentationOpen-Vocabulary Video SegmentationPanoptic Segmentation+8

SAM2Auto: Auto Annotation Using FLASH

2025-06-09 · Arash Rocky, Q. M. Jonathan Wu

Vision-Language Models (VLMs) lag behind Large Language Models due to the scarcity of annotated datasets, as creating paired visual-textual annotations is labor-intensive and expensive. To address this bottleneck, we int…

Instance SegmentationObjectobject-detectionObject Detection+5

ThinkVideo: High-Quality Reasoning Video Segmentation with Chain of Thoughts

2025-05-24 · Shiu-hong Kao, Yu-Wing Tai, Chi-Keung Tang

Reasoning Video Object Segmentation is a challenging task, which generates a mask sequence from an input video and an implicit, complex text query. Existing works probe into the problem by finetuning Multimodal Large Lan…

Image SegmentationInstance SegmentationObjectReasoning Video Object Segmentation+6

FlowCut: Unsupervised Video Instance Segmentation via Temporal Mask Matching

2025-05-19 · Alp Eren Sari, Paolo Favaro

We propose FlowCut, a simple and capable method for unsupervised video instance segmentation consisting of a three-stage framework to construct a high-quality video dataset with pseudo labels. To our knowledge, our work …

Instance SegmentationSegmentationSemantic SegmentationVideo Instance Segmentation+2
1–20 / 163 다음 →