paper-with-me

Papers Semi-Supervised Video Object Segmentation

“Semi-Supervised Video Object Segmentation” 태그가 달린 논문 154편 · 필터 해제

TinySAM 2: Extreme Memory Compression for Efficient Track Anything Model

2026-05-18 · Zhaoyuan Ding, Yijing Yang, Han Shu, Xinghao Chen arxiv

Segment Anything Model 2 (SAM 2) serves as a core foundation model in the field of video segmentation. Building upon the original SAM model, it introduces a memory bank mechanism and demonstrates outstanding performance …

Semi-Supervised Video Object SegmentationVideo Segmentation

Re-Prompting SAM 3 via Object Retrieval: 3rd of the 5th PVUW MOSE Track

2026-03-24 · Mingqi Gao, Sijie Li, Jungong Han arxiv

This technical report explores the MOSEv2 track of the PVUW 2026 Challenge, which targets complex semi-supervised video object segmentation. Built on SAM~3, we develop an automatic re-prompting framework to improve robus…

Semi-Supervised Video Object Segmentation

Segment Anything Across Shots: A Method and Benchmark

2025-11-17 · Hengrui Hu, Kaining Ying, Henghui Ding arxiv

This work focuses on multi-shot semi-supervised video object segmentation (MVOS), which aims at segmenting the target object indicated by an initial mask throughout a video with multiple shots. The existing VOS methods m…

Semi-Supervised Video Object SegmentationData Augmentation

2nd Place Report of MOSEv2 Challenge 2025: Concept Guided Video Object Segmentation via SeC

2025-09-28 · Zhixiong Zhang, Shuangrui Ding, Xiaoyi Dong, Yuhang Zang 외 arxiv

Semi-supervised Video Object Segmentation aims to segment a specified target throughout a video sequence, initialized by a first-frame mask. Previous methods rely heavily on appearance-based pattern matching and thus exh…

Semi-Supervised Video Object Segmentation

The 1st Solution for MOSEv2 Challenge 2025: Long-term and Concept-aware Video Segmentation via SeC

2025-09-23 · Mingqi Gao, Jingkun Chen, Yunqi Miao, Gengshen Wu 외 arxiv

This technical report explores the MOSEv2 track of the LSVOS Challenge, which targets complex semi-supervised video object segmentation. By analysing and adapting SeC, an enhanced SAM-2 framework, we conduct a detailed s…

Semi-Supervised Video Object SegmentationVideo Segmentation

VoCap: Video Object Captioning and Segmentation from Any Prompt

2025-08-29 · Jasper Uijlings, Xingyi Zhou, Xiuye Gu, Arsha Nagrani 외 arxiv

Understanding objects in videos in terms of fine-grained localization masks and detailed semantic properties is a fundamental task in video understanding. In this paper, we propose VoCap, a flexible video model that cons…

Semi-Supervised Video Object SegmentationReferring Expression Segmentation

Structure Matters: Revisiting Boundary Refinement in Video Object Segmentation

2025-07-25 · Guanyi Qin, Ziyue Wang, Daiyun Shen, Haofeng Liu 외 arxiv

Given an object mask, Semi-supervised Video Object Segmentation (SVOS) technique aims to track and segment the object across video frames, serving as a fundamental task in computer vision. Although recent memory-based me…

Semi-Supervised Video Object Segmentation

THU-Warwick Submission for EPIC-KITCHEN Challenge 2025: Semi-Supervised Video Object Segmentation

2025-06-07 · Mingqi Gao, Haoran Duan, Tianlu Zhang, Jungong Han

In this report, we describe our approach to egocentric video object segmentation. Our method combines large-scale visual pretraining from SAM2 with depth-based geometric cues to handle complex scenes and long-term tracki…

SegmentationSemantic SegmentationSemi-Supervised Video Object SegmentationVideo Object Segmentation+1

Exploring Enhanced Contextual Information for Video-Level Object Tracking

2024-12-15 · AAAI2025 2024 12 · Ben Kang, Xin Chen, Simiao Lai, Yang Liu 외

Contextual information at the video level has become increasingly crucial for visual object tracking. However, existing methods typically use only a few tokens to convey this information, which can lead to information lo…

ObjectObject TrackingSemi-Supervised Video Object SegmentationVideo Object Tracking+2

A Distractor-Aware Memory for Visual Object Tracking with SAM2

2024-11-26 · CVPR 2025 1 · Jovana Videnovic, Alan Lukezic, Matej Kristan

Memory-based trackers are video object segmentation methods that form the target model by concatenating recently tracked frames into a memory buffer and localize the target by attending the current image to the buffered …

Object TrackingSemi-Supervised Video Object SegmentationVisual Object TrackingVisual Tracking

LiVOS: Light Video Object Segmentation with Gated Linear Matching

2024-11-05 · CVPR 2025 1 · Qin Liu, JianFeng Wang, Zhengyuan Yang, Linjie Li 외

Semi-supervised video object segmentation (VOS) has been largely driven by space-time memory (STM) networks, which store past frame features in a spatiotemporal memory to segment the current frame via softmax attention. …

GPUSemantic SegmentationSemi-Supervised Video Object SegmentationVideo Object Segmentation+1

Memory Matching is not Enough: Jointly Improving Memory Matching and Decoding for Video Object Segmentation

2024-09-22 · Jintu Zheng, Yun Liang, Yuqing Zhang, Wanchao Su

Memory-based video object segmentation methods model multiple objects over long temporal-spatial spans by establishing memory bank, which achieve the remarkable performance. However, they struggle to overcome the false m…

Semantic SegmentationSemi-Supervised Video Object SegmentationVideo Object SegmentationVideo Semantic Segmentation

SAM 2: Segment Anything in Images and Videos

2024-08-01 · Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang Hu 외

We present Segment Anything Model 2 (SAM 2), a foundation model towards solving promptable visual segmentation in images and videos. We build a data engine, which improves model and data via user interaction, to collect …

Image SegmentationRobot Manipulation GeneralizationSegmentationSemantic Segmentation+5

Global Motion Understanding in Large-Scale Video Object Segmentation

2024-05-11 · Volodymyr Fedynyak, Yaroslav Romanus, Oles Dobosevych, Igor Babin 외

In this paper, we show that transferring knowledge from other domains of video understanding combined with large-scale learning can improve robustness of Video Object Segmentation (VOS) under complex circumstances. Namel…

Instance SegmentationOptical Flow EstimationSegmentationSemantic Segmentation+4

Spatial-Temporal Multi-level Association for Video Object Segmentation

2024-04-09 · Deshui Miao, Xin Li, Zhenyu He, Huchuan Lu 외

Existing semi-supervised video object segmentation methods either focus on temporal feature matching or spatial-temporal feature modeling. However, they do not address the issues of sufficient target interaction and effi…

ObjectSegmentationSemantic SegmentationSemi-Supervised Video Object Segmentation+2

Efficient Video Object Segmentation via Modulated Cross-Attention Memory

2024-03-26 · Abdelrahman Shaker, Syed Talal Wasim, Martin Danelljan, Salman Khan 외

Recently, transformer-based approaches have shown promising results for semi-supervised video object segmentation. However, these approaches typically struggle on long videos due to increased GPU memory demands, as they …

GPUObjectSegmentationSemantic Segmentation+3

Video Object Segmentation with Dynamic Query Modulation

2024-03-18 · Hantao Zhou, Runze Hu, Xiu Li

Storing intermediate frame segmentations as memory for long-range context modeling, spatial-temporal memory-based methods have recently showcased impressive results in semi-supervised video object segmentation (SVOS). Ho…

ObjectSegmentationSemantic SegmentationSemi-Supervised Video Object Segmentation+2

Augmenting Efficient Real-time Surgical Instrument Segmentation in Video with Point Tracking and Segment Anything

2024-03-12 · Zijian Wu, Adam Schmidt, Peter Kazanzides, Septimiu E. Salcudean

The Segment Anything Model (SAM) is a powerful vision foundation model that is revolutionizing the traditional paradigm of segmentation. Despite this, a reliance on prompting each frame and large computational cost limit…

GPUPoint TrackingSegmentationSemantic Segmentation+4

Lester: rotoscope animation through video object segmentation and tracking

2024-02-15 · Ruben Tous

This article introduces Lester, a novel method to automatically synthetise retro-style 2D animations from videos. The method approaches the challenge mainly as an object segmentation and tracking problem. Video frames ar…

3D Human Pose EstimationObjectPose EstimationSemantic Segmentation+3

ODTrack: Online Dense Temporal Token Learning for Visual Tracking

2024-01-03 · Yaozong Zheng, Bineng Zhong, Qihua Liang, Zhiyi Mo 외

Online contextual reasoning and association across consecutive video frames are critical to perceive instances in visual tracking. However, most current top-performing trackers persistently lean on sparse temporal relati…

Semi-Supervised Video Object SegmentationVideo Object TrackingVisual Object TrackingVisual Tracking
1–20 / 154 다음 →