Papers Video Segmentation
“Video Segmentation” 태그가 달린 논문 487편 · 필터 해제
Bootstrapping a 4D LiDAR Annotation Tool from Video Foundation Models
Progress in 4D LiDAR segmentation is bottlenecked by data. Assigning temporally consistent labels across sparse point cloud sequences is costly and hard to scale, and every new task or domain tends to demand fresh dense …
Scene UnderstandingVideo SegmentationMLLM-Assisted Audio VOS: A 3rd Place Report for the MeViS-Audio Track, 8th LSVOS Challenge
In this technical report, we present a training-free framework for audio-guided video object segmentation, which integrates Multimodal Large Language Models (MLLMs) with SAM-based segmentation models. We decompose the ta…
Video Object SegmentationMultimodal ReasoningVideo SegmentationCoarse Indexing, Fine Evidence: Decoupling Temporal Granularity in Long-Video RAG
Graph-based retrieval-augmented generation (RAG) provides a scalable paradigm for long-video understanding, but existing systems typically inherit a fixed temporal granularity from video segmentation when constructing th…
Video SegmentationAnswer GenerationStitch-Inferencer: Enhance Endoscopic Video Segmentation and Tracking via Panoramic Reconstruction
Surgical video understanding is fundamental to navigation systems. Endoscopic perception is often hindered by a limited field-of-view and frequent instrument occlusions, making spatio-temporal context essential for robus…
Video SegmentationSAM-MT: Real-Time Interactive Multi-Target Video Segmentation
Modern Video Object Segmentation (VOS) involves tracking and segmenting user-specified targets. While recent approaches have achieved remarkable performance in single-target scenarios, extending them to multi-target sett…
Video Object SegmentationVideo SegmentationG$^2$TAM: Geometry Grounded Track Anything Model
Human spatial understanding arises from jointly perceiving geometry and semantics, enabling consistent object identification and localization across viewpoints and time. Current video segmentation models depend on explic…
Video Object SegmentationVideo SegmentationSpatial Reasoning3D ReconstructionDistilling Temporal Coherence into 2D Networks for Transrectal Ultrasound Prostate Video Segmentation
Real-time video segmentation of the prostate in Transrectal Ultrasound (TRUS) is essential for image-guided interventions. While conventional 2D methods suffer from inter-frame inconsistencies by disregarding temporal co…
Knowledge DistillationVideo SegmentationEvent-Aware Instructed Assistant for Referring Video Segmentation
Existing referring video segmentation methods often treat a video as a single event consisting of multiple images, overlooking the fact that a video typically contains multiple distinct events. Under such a mechanism, th…
Video SegmentationBoosting Text-Driven Video Segmentation via Geometry-Aware Distillation
Text-driven Referring Video Object Segmentation (RVOS) aims to locate and segment target objects in videos given natural language. However, existing models are typically trained on 2D image or video datasets with naive s…
Referring Video Object SegmentationZero-shot GeneralizationImage SegmentationVideo SegmentationScene-Centric Unsupervised Video Panoptic Segmentation
Video panoptic segmentation (VPS) aims to jointly detect, segment, and track all objects while partitioning the video into semantically consistent regions. We introduce the task setting of unsupervised VPS, omitting any …
Video Panoptic SegmentationScene UnderstandingVideo SegmentationImage SegmentationEchoPilot: Training-Free Ultrasound Video Segmentation via Scale-Space Semantic Prompting and Reliability-Gated Memory
Ultrasound video segmentation is clinically valuable yet difficult due to speckle noise, weak boundaries, and rapid anatomical deformation. Recent promptable foundation models enable point-guided segmentation, but their …
Video SegmentationTinySAM 2: Extreme Memory Compression for Efficient Track Anything Model
Segment Anything Model 2 (SAM 2) serves as a core foundation model in the field of video segmentation. Building upon the original SAM model, it introduces a memory bank mechanism and demonstrates outstanding performance …
Semi-Supervised Video Object SegmentationVideo SegmentationSteerSeg: Attention Steering for Reasoning Video Segmentation
Video reasoning segmentation requires localizing objects across video frames from natural language expressions, often involving spatial reasoning and implicit references. Recent approaches leverage frozen large vision-la…
Video SegmentationSpatial ReasoningText GenerationCross-Modal-Domain Generalization Through Semantically Aligned Discrete Representations
Multimodal learning seeks to integrate information across diverse sensory sources, yet current approaches struggle to balance cross-modal generalizability with modality-specific structure. Continuous (implicit) methods p…
Representation LearningDomain GeneralizationVideo SegmentationLightweight Distillation of SAM 3 and DINOv3 for Edge-Deployable Individual-Level Livestock Monitoring and Longitudinal Visual Analytics
Foundation-model pipelines for individual-level livestock monitoring -- combining open-vocabulary detection, promptable video segmentation, and self-supervised visual embeddings -- have raised the accuracy ceiling of pre…
Video Segmentation2nd of the 5th PVUW MeViS-Audio Track: ASR-SaSaSa2VA
Audio-based video object segmentation aims to locate and segment objects in videos conditioned on audio cues, requiring precise understanding of both appearance and motion. Recent audio-driven video segmentation methods …
Video Object SegmentationSpeech RecognitionVideo SegmentationX2SAM: Any Segmentation in Images and Videos
Multimodal Large Language Models (MLLMs) have demonstrated strong image-level visual understanding and reasoning, yet their pixel-level perception across both images and videos remains limited. Foundation segmentation mo…
Image SegmentationVideo SegmentationAmodal SAM: A Unified Amodal Segmentation Framework with Generalization
Amodal segmentation is a challenging task that aims to predict the complete geometric shape of objects, including their occluded regions. Although existing methods primarily focus on amodal segmentation within the traini…
Video SegmentationInference-Time Temporal Probability Smoothing for Stable Video Segmentation with SAM2 under Weak Prompts
Interactive video segmentation models such as SAM2 have demonstrated strong generalization across diverse visual domains. However, under weak user supervision, for example, when sparse point prompts are provided on a sin…
Video SegmentationUniSurgSAM: A Unified Promptable Model for Reliable Surgical Video Segmentation
Surgical video segmentation is fundamental to computer-assisted surgery. In practice, surgeons need to dynamically specify targets throughout extended procedures, using heterogeneous cues such as visual selections, textu…
Video Object SegmentationVideo Segmentation