paper-with-me

Papers Video Object Segmentation

“Video Object Segmentation” 태그가 달린 논문 608편 · 필터 해제

MLLM-Assisted Audio VOS: A 3rd Place Report for the MeViS-Audio Track, 8th LSVOS Challenge

2026-08-24 · Liangtao Shi, Jinxia Xie, Xiantao Hu, Ting Liu arxiv

In this technical report, we present a training-free framework for audio-guided video object segmentation, which integrates Multimodal Large Language Models (MLLMs) with SAM-based segmentation models. We decompose the ta…

Video Object SegmentationMultimodal ReasoningVideo Segmentation

SAM2Dual: Training-Free, Dual Memory for Long-Term Video Object Segmentation

2026-08-19 · JeongRae Kim, Changwon Lim arxiv

Long-term video object segmentation (VOS) remains challenging due to error accumulation under extended occlusions, re-appearance, and scene changes. Although SAM2 provides strong zero-shot performance, its streaming memo…

Video Object Segmentation

RRTrack: Robust and Recoverable Object 6D Pose Tracking for Dynamic Scenes

2026-07-26 · Junyue Li, Ye Zheng, Yifan Chen, Zhe Sun 외 arxiv

Robust object 6D pose tracking is critical for robotic systems operating in dynamic and occluded scenes. Per-frame estimators are accurate but computationally expensive, while current trackers struggle with fast motion a…

Video Object SegmentationPose Tracking

REMIND: RE-Identification with Memory for INDoor Navigation

2026-07-10 · Pablo Diaz-Pereda, Alejandro Rodriguez-Ramos, David Perez-Saura, Pascual Campoy arxiv

Mobile robots operating indoors must re-identify previously observed objects after long temporal gaps, significant viewpoint changes, and severe illumination variations. This remains a challenging problem: multi-object t…

Video Object SegmentationVehicle Re-IdentificationMulti-Object Tracking

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation

2026-07-09 · Ruiqi Shen, Chang Liu, Henghui Ding arxiv

Modern Video Object Segmentation (VOS) involves tracking and segmenting user-specified targets. While recent approaches have achieved remarkable performance in single-target scenarios, extending them to multi-target sett…

Video Object SegmentationVideo Segmentation

`Attention-Guided Cross-Temporal Clustering for Self-Supervised Video Object Segmentation

2026-07-08 · Waqas Arshid, Mohammad Awrangjeb, Alan Wee-Chung Liew, Yongsheng Gao arxiv

Video object segmentation (VOS) is a fundamental task in video understanding, requiring accurate delineation and consistent tracking of objects across frames. While supervised methods achieve strong performance, they rel…

Video Object SegmentationSelf-Supervised Learning

G$^2$TAM: Geometry Grounded Track Anything Model

2026-07-04 · Chenming Zhu, Peizhou Cao, Jingli Lin, Wenbo Hu 외 arxiv

Human spatial understanding arises from jointly perceiving geometry and semantics, enabling consistent object identification and localization across viewpoints and time. Current video segmentation models depend on explic…

Video Object SegmentationVideo SegmentationSpatial Reasoning3D Reconstruction

Selective Mask Propagation for Multi-Object Tracking

2026-06-11 · Alexander Holmberg arxiv

In multi-object tracking, most frames are easy for a lightweight base tracker while a small fraction is intrinsically hard. Video object segmentation (VOS) models can often preserve identity through the hard frames where…

Video Object SegmentationMulti-Object Tracking

VideoSEG-O3: A Multi-turn Reinforcement Learning Framework for Reasoning Video Object Segmentation

2026-06-05 · Ming Dai, Sen Yang, Boqiang Duan, Boyuan Tong 외 arxiv

Reasoning Video Object Segmentation (RVOS) demands a sophisticated integration of temporal dynamics, spatial details, and linguistic reasoning to achieve precise pixel-level localization. Existing methods are limited to …

Video Object SegmentationReinforcement Learning

GMOS: Grounding Moving Object Segmentation in 3D Space and Time

2026-05-28 · Junyu Xie, Tengda Han, Weidi Xie, Andrew Zisserman arxiv

Moving Object Segmentation (MOS) aims to discover, segment, and track objects that move independently of the camera. Current MOS methods, however, exhibit two fundamental limitations: they rely on pre-computed 2D auxilia…

Video Object Segmentation

Weighted Reverse Convolution for Feature Upsampling

2026-05-17 · Wentong Li, Zhiyuan Qi, Zichen Zhao, Kai Zhang 외 arxiv

Pre-trained vision foundation models (VFMs) provide strong semantic representations, yet their patch-level features are inherently coarse, limiting their effectiveness on tasks requiring fine-grained localization, dense …

Video Object SegmentationComputational EfficiencyFeature UpsamplingDepth Estimation

Robust Promptable Video Object Segmentation

2026-05-12 · Sohyun Lee, Yeho Gwon, Lukas Hoyer, Konrad Schindler 외 arxiv

The performance of promptable video object segmentation (PVOS) models substantially degrades under input corruptions, which prevents PVOS deployment in safety-critical domains. This paper offers the first comprehensive s…

Video Object Segmentation

Featurising Pixels from Dynamic 3D Scenes with Linear In-Context Learners

2026-04-29 · Nikita Araslanov, Martin Sundermeyer, Hidenobu Matsuki, David Joseph Tan 외 arxiv

One of the most exciting applications of vision models involve pixel-level reasoning. Despite the abundance of vision foundation models, we still lack representations that effectively embed spatio-temporal properties of …

Video Object SegmentationSemantic Segmentation

2nd of the 5th PVUW MeViS-Audio Track: ASR-SaSaSa2VA

2026-04-27 · Zhiyu Wang, Xudong Kang, Shutao Li arxiv

Audio-based video object segmentation aims to locate and segment objects in videos conditioned on audio cues, requiring precise understanding of both appearance and motion. Recent audio-driven video segmentation methods …

Video Object SegmentationSpeech RecognitionVideo Segmentation

SynMulti: Synthetic-to-Real Learning for Multimodal Video Understanding

2026-04-14 · Tanzila Rahman, Renjie Liao, Leonid Sigal arxiv

Training multimodal large language models (MLLMs) for video understanding requires large-scale annotated data spanning diverse tasks such as object counting, question answering, and segmentation. However, collecting and …

Video Object SegmentationSynthetic Data GenerationVisual Question AnsweringVisual Grounding

Online Reasoning Video Object Segmentation

2026-04-13 · Jinyuan Liu, Yang Wang, Zeyu Zhao, Weixin Li 외 arxiv

Reasoning video object segmentation predicts pixel-level masks in videos from natural-language queries that may involve implicit and temporally grounded references. However, existing methods are developed and evaluated i…

Video Object Segmentation

PanoSAM2: Lightweight Distortion- and Memory-aware Adaptions of SAM2 for 360 Video Object Segmentation

2026-04-09 · Dingwen Xiao, Weiming Zhang, Shiqi Wen, Lin Wang arxiv

360 video object segmentation (360VOS) aims to predict temporally-consistent masks in 360 videos, offering full-scene coverage, benefiting applications, such as VR/AR and embodied AI. Learning 360VOS model is nontrivial …

Video Object Segmentation

UniSurgSAM: A Unified Promptable Model for Reliable Surgical Video Segmentation

2026-04-04 · Haofeng Liu, Ziyue Wang, Alex Y. W. Kong, Guanyi Qin 외 arxiv

Surgical video segmentation is fundamental to computer-assisted surgery. In practice, surgeons need to dynamically specify targets throughout extended procedures, using heterogeneous cues such as visual selections, textu…

Video Object SegmentationVideo Segmentation

Advancing Complex Video Object Segmentation via Tracking-Enhanced Prompt: The 1st Winner for 5th PVUW MOSE Challenge

2026-04-01 · Jinrong Zhang, Canyang Wu, Xusheng He, Weili Guan 외 arxiv

In the Complex Video Object Segmentation task, researchers are required to track and segment specific targets within cluttered environments, which rigorously tests a method's capability for target comprehension and envir…

Video Object Segmentation

Colon-Bench: An Agentic Workflow for Scalable Dense Lesion Annotation in Full-Procedure Colonoscopy Videos

2026-03-26 · Abdullah Hamdi, Changchun Yang, Xin Gao arxiv

Early screening via colonoscopy is critical for colon cancer prevention, yet developing robust AI systems for this domain is hindered by the lack of densely annotated, long-sequence video datasets. Existing datasets pred…

Video Object SegmentationVisual Question Answering
1–20 / 608 다음 →