Papers Video Semantic Segmentation
“Video Semantic Segmentation” 태그가 달린 논문 905편 · 필터 해제
Surgical Anatomy Recognition with Context Learning using Foundation Representations
Accurate recognition of anatomical structures is essential for safe and effective minimally invasive surgery (MIS), yet it remains underexplored in surgical computer vision due to limited annotated data and methods tailo…
Video Semantic SegmentationScene UnderstandingObject TrackingZero-Parameter Geometric Gating for Temporally Stable Low-Altitude UAV Video Semantic Segmentation
Video semantic segmentation for low-altitude UAVs requires temporal consistency, yet dense optical flow introduces spatially structured noise in the planar regions that dominate aerial imagery. We propose a zero-paramete…
Video Semantic SegmentationSemantic SimilaritySeeing Realism from Simulation: Efficient Video Transfer for Vision-Language-Action Data Augmentation
Vision-language-action (VLA) models typically rely on large-scale real-world videos, whereas simulated data, despite being inexpensive and highly parallelizable to collect, often suffers from a substantial visual domain …
Video Semantic SegmentationData AugmentationVideo CaptioningBootstrapping Video Semantic Segmentation Model via Distillation-assisted Test-Time Adaptation
Fully supervised Video Semantic Segmentation (VSS) relies heavily on densely annotated video data, limiting practical applicability. Alternatively, applying pre-trained Image Semantic Segmentation (ISS) models frame-by-f…
Video Semantic SegmentationTest-time AdaptationCan Unsupervised Segmentation Reduce Annotation Costs for Video Semantic Segmentation?
Present-day deep neural networks for video semantic segmentation require a large number of fine-grained pixel-level annotations to achieve the best possible results. Obtaining such annotations, however, is very expensive…
Video Semantic SegmentationVideo SegmentationRS-SSM: Refining Forgotten Specifics in State Space Model for Video Semantic Segmentation
Recently, state space models have demonstrated efficient video segmentation through linear-complexity state space compression. However, Video Semantic Segmentation (VSS) requires pixel-level spatiotemporal modeling capab…
Video Semantic SegmentationComputational EfficiencyVideo SegmentationInterpretable Motion-Attentive Maps: Spatio-Temporally Localizing Concepts in Video Diffusion Transformers
Video Diffusion Transformers (DiTs) have been synthesizing high-quality video with high fidelity from given text descriptions involving motion. However, understanding how Video DiTs convert motion words into video remain…
Video Semantic SegmentationTime2General: Learning Spatiotemporal Invariant Representations for Domain-Generalization Video Semantic Segmentation
Domain Generalized Video Semantic Segmentation (DGVSS) is trained on a single labeled driving domain and is directly deployed on unseen domains without target labels and test-time adaptation while maintaining temporally …
Video Semantic SegmentationTest-time AdaptationSpatio-Temporal Attention for Consistent Video Semantic Segmentation in Automated Driving
Deep neural networks, especially transformer-based architectures, have achieved remarkable success in semantic segmentation for environmental perception. However, existing models process video frames independently, thus …
Video Semantic SegmentationComputational EfficiencyEvaluating SAM2 for Video Semantic Segmentation
The Segmentation Anything Model 2 (SAM2) has proven to be a powerful foundation model for promptable visual object segmentation in both images and videos, capable of storing object-aware memories and transferring them te…
Video Semantic SegmentationVideo Object SegmentationSeC: Advancing Complex Video Object Segmentation via Progressive Concept Construction
Video Object Segmentation (VOS) is a core task in computer vision, requiring models to track and segment target objects across video frames. Despite notable advances with recent efforts, current techniques still lag behi…
ObjectSegmentationSemantic SegmentationVideo Object Segmentation+1Memory-Augmented SAM2 for Training-Free Surgical Video Segmentation
Surgical video segmentation is a critical task in computer-assisted surgery, essential for enhancing surgical quality and patient outcomes. Recently, the Segment Anything Model 2 (SAM2) framework has demonstrated remarka…
SegmentationSemantic SegmentationVideo Object SegmentationVideo Segmentation+1MUVOD: A Novel Multi-view Video Object Segmentation Dataset and A Benchmark for 3D Segmentation
The application of methods based on Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3D GS) have steadily gained popularity in the field of 3D object segmentation in static scenes. These approaches demonstrate ef…
NeRFObjectScene UnderstandingSegmentation+4Decoupled Seg Tokens Make Stronger Reasoning Video Segmenter and Grounder
Existing video segmenter and grounder approaches, exemplified by Sa2VA, directly fuse features within segmentation models. This often results in an undesirable entanglement of dynamic visual information and static semant…
Image SegmentationLarge Language ModelQuestion AnsweringSegmentation+5CogGen: A Learner-Centered Generative AI Architecture for Intelligent Tutoring with Programming Video
We introduce CogGen, a learner-centered AI architecture that transforms programming videos into interactive, adaptive learning experiences by integrating student modeling with generative AI tutoring based on the Cognitiv…
Knowledge TracingVideo SegmentationVideo Semantic SegmentationLeader360V: The Large-scale, Real-world 360 Video Dataset for Multi-task Learning in Diverse Environment
360 video captures the complete surrounding scenes with the ultra-large field of view of 360X180. This makes 360 scene understanding tasks, eg, segmentation and tracking, crucial for appications, such as autonomous drivi…
Autonomous DrivingInstance SegmentationMulti-Task LearningScene Understanding+3A Comprehensive Survey on Video Scene Parsing:Advances, Challenges, and Prospects
Video Scene Parsing (VSP) has emerged as a cornerstone in computer vision, facilitating the simultaneous segmentation, recognition, and tracking of diverse visual entities in dynamic scenes. In this survey, we present a …
BenchmarkingInstance SegmentationOpen-Vocabulary Video SegmentationPanoptic Segmentation+8M^3-VOS: Multi-Phase, Multi-Transition, and Multi-Scenery Video Object Segmentation
Intelligent robots need to interact with diverse objects across various environments. The appearance and state of objects frequently undergo complex transformations depending on the object properties, e.g., phase transit…
ObjectSemantic SegmentationVideo Object SegmentationVideo Semantic Segmentation+2Q-SAM2: Accurate Quantization for Segment Anything Model 2
The Segment Anything Model 2 (SAM2) has gained significant attention as a foundational approach for promptable image and video segmentation. However, its expensive computational and memory consumption poses a severe chal…
QuantizationVideo SegmentationVideo Semantic SegmentationTHU-Warwick Submission for EPIC-KITCHEN Challenge 2025: Semi-Supervised Video Object Segmentation
In this report, we describe our approach to egocentric video object segmentation. Our method combines large-scale visual pretraining from SAM2 with depth-based geometric cues to handle complex scenes and long-term tracki…
SegmentationSemantic SegmentationSemi-Supervised Video Object SegmentationVideo Object Segmentation+1