paper-with-me

Papers Video Semantic Segmentation

“Video Semantic Segmentation” 태그가 달린 논문 905편 · 필터 해제

Surgical Anatomy Recognition with Context Learning using Foundation Representations

2026-06-20 · Ronald L. P. D. de Jong, Tim J. M. Jaspers, Raf A. H. Vervoort, Aron F. H. A. Bakker 외 arxiv

Accurate recognition of anatomical structures is essential for safe and effective minimally invasive surgery (MIS), yet it remains underexplored in surgical computer vision due to limited annotated data and methods tailo…

Video Semantic SegmentationScene UnderstandingObject Tracking

Zero-Parameter Geometric Gating for Temporally Stable Low-Altitude UAV Video Semantic Segmentation

2026-06-08 · Jingpu Yang, Fengxian Ji, Zhengzhao Lai, Juanfan Wu 외 arxiv

Video semantic segmentation for low-altitude UAVs requires temporal consistency, yet dense optical flow introduces spatially structured noise in the planar regions that dominate aerial imagery. We propose a zero-paramete…

Video Semantic SegmentationSemantic Similarity

Seeing Realism from Simulation: Efficient Video Transfer for Vision-Language-Action Data Augmentation

2026-05-04 · Chenyu Hui, Xiaodi Huang, Siyu Xu, Yunke Wang 외 arxiv

Vision-language-action (VLA) models typically rely on large-scale real-world videos, whereas simulated data, despite being inexpensive and highly parallelizable to collect, often suffers from a substantial visual domain …

Video Semantic SegmentationData AugmentationVideo Captioning

Bootstrapping Video Semantic Segmentation Model via Distillation-assisted Test-Time Adaptation

2026-04-13 · Jihun Kim, Hoyong Kwon, Hyeokjun Kweon, Kuk-Jin Yoon arxiv

Fully supervised Video Semantic Segmentation (VSS) relies heavily on densely annotated video data, limiting practical applicability. Alternatively, applying pre-trained Image Semantic Segmentation (ISS) models frame-by-f…

Video Semantic SegmentationTest-time Adaptation

Can Unsupervised Segmentation Reduce Annotation Costs for Video Semantic Segmentation?

2026-03-29 · Samik Some, Vinay P. Namboodiri arxiv

Present-day deep neural networks for video semantic segmentation require a large number of fine-grained pixel-level annotations to achieve the best possible results. Obtaining such annotations, however, is very expensive…

Video Semantic SegmentationVideo Segmentation

RS-SSM: Refining Forgotten Specifics in State Space Model for Video Semantic Segmentation

2026-03-25 · Kai Zhu, Zhenyu Cui, Zehua Zang, Jiahuan Zhou arxiv

Recently, state space models have demonstrated efficient video segmentation through linear-complexity state space compression. However, Video Semantic Segmentation (VSS) requires pixel-level spatiotemporal modeling capab…

Video Semantic SegmentationComputational EfficiencyVideo Segmentation

Interpretable Motion-Attentive Maps: Spatio-Temporally Localizing Concepts in Video Diffusion Transformers

2026-03-03 · Youngjun Jun, Seil Kang, Woojung Han, Seong Jae Hwang arxiv

Video Diffusion Transformers (DiTs) have been synthesizing high-quality video with high fidelity from given text descriptions involving motion. However, understanding how Video DiTs convert motion words into video remain…

Video Semantic Segmentation

Time2General: Learning Spatiotemporal Invariant Representations for Domain-Generalization Video Semantic Segmentation

2026-02-10 · Siyu Chen, Ting Han, Haoling Huang, Chaolei Wang 외 arxiv

Domain Generalized Video Semantic Segmentation (DGVSS) is trained on a single labeled driving domain and is directly deployed on unseen domains without target labels and test-time adaptation while maintaining temporally …

Video Semantic SegmentationTest-time Adaptation

Spatio-Temporal Attention for Consistent Video Semantic Segmentation in Automated Driving

2026-02-10 · Serin Varghese, Kevin Ross, Fabian Hueger, Kira Maag arxiv

Deep neural networks, especially transformer-based architectures, have achieved remarkable success in semantic segmentation for environmental perception. However, existing models process video frames independently, thus …

Video Semantic SegmentationComputational Efficiency

Evaluating SAM2 for Video Semantic Segmentation

2025-12-01 · Syed Hesham Syed Ariff, Yun Liu, Guolei Sun, Jing Yang 외 arxiv

The Segmentation Anything Model 2 (SAM2) has proven to be a powerful foundation model for promptable visual object segmentation in both images and videos, capable of storing object-aware memories and transferring them te…

Video Semantic SegmentationVideo Object Segmentation

SeC: Advancing Complex Video Object Segmentation via Progressive Concept Construction

2025-07-21 · Zhixiong Zhang, Shuangrui Ding, Xiaoyi Dong, Songxin He 외

Video Object Segmentation (VOS) is a core task in computer vision, requiring models to track and segment target objects across video frames. Despite notable advances with recent efforts, current techniques still lag behi…

ObjectSegmentationSemantic SegmentationVideo Object Segmentation+1

Memory-Augmented SAM2 for Training-Free Surgical Video Segmentation

2025-07-13 · Ming Yin, Fu Wang, Xujiong Ye, Yanda Meng 외

Surgical video segmentation is a critical task in computer-assisted surgery, essential for enhancing surgical quality and patient outcomes. Recently, the Segment Anything Model 2 (SAM2) framework has demonstrated remarka…

SegmentationSemantic SegmentationVideo Object SegmentationVideo Segmentation+1

MUVOD: A Novel Multi-view Video Object Segmentation Dataset and A Benchmark for 3D Segmentation

2025-07-10 · Bangning Wei, Joshua Maraval, Meriem Outtas, Kidiyo Kpalma 외

The application of methods based on Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3D GS) have steadily gained popularity in the field of 3D object segmentation in static scenes. These approaches demonstrate ef…

NeRFObjectScene UnderstandingSegmentation+4

Decoupled Seg Tokens Make Stronger Reasoning Video Segmenter and Grounder

2025-06-28 · Dang Jisheng, Wu Xudong, Wang Bimei, Lv Ning 외

Existing video segmenter and grounder approaches, exemplified by Sa2VA, directly fuse features within segmentation models. This often results in an undesirable entanglement of dynamic visual information and static semant…

Image SegmentationLarge Language ModelQuestion AnsweringSegmentation+5

CogGen: A Learner-Centered Generative AI Architecture for Intelligent Tutoring with Programming Video

2025-06-25 · Wengxi Li, Roy Pea, Nick Haber, Hari Subramonyam

We introduce CogGen, a learner-centered AI architecture that transforms programming videos into interactive, adaptive learning experiences by integrating student modeling with generative AI tutoring based on the Cognitiv…

Knowledge TracingVideo SegmentationVideo Semantic Segmentation

Leader360V: The Large-scale, Real-world 360 Video Dataset for Multi-task Learning in Diverse Environment

2025-06-17 · Weiming Zhang, Dingwen Xiao, Aobotao Dai, Yexin Liu 외

360 video captures the complete surrounding scenes with the ultra-large field of view of 360X180. This makes 360 scene understanding tasks, eg, segmentation and tracking, crucial for appications, such as autonomous drivi…

Autonomous DrivingInstance SegmentationMulti-Task LearningScene Understanding+3

A Comprehensive Survey on Video Scene Parsing:Advances, Challenges, and Prospects

2025-06-16 · Guohuan Xie, Syed Ariff Syed Hesham, Wenya Guo, Bing Li 외

Video Scene Parsing (VSP) has emerged as a cornerstone in computer vision, facilitating the simultaneous segmentation, recognition, and tracking of diverse visual entities in dynamic scenes. In this survey, we present a …

BenchmarkingInstance SegmentationOpen-Vocabulary Video SegmentationPanoptic Segmentation+8

M^3-VOS: Multi-Phase, Multi-Transition, and Multi-Scenery Video Object Segmentation

2025-06-15 · CVPR 2025 1 · Zixuan Chen, Jiaxin Li, Liming Tan, Yejie Guo 외

Intelligent robots need to interact with diverse objects across various environments. The appearance and state of objects frequently undergo complex transformations depending on the object properties, e.g., phase transit…

ObjectSemantic SegmentationVideo Object SegmentationVideo Semantic Segmentation+2

Q-SAM2: Accurate Quantization for Segment Anything Model 2

2025-06-11 · Nicola Farronato, Florian Scheidegger, Mattia Rigotti, Cristiano Malossi 외

The Segment Anything Model 2 (SAM2) has gained significant attention as a foundational approach for promptable image and video segmentation. However, its expensive computational and memory consumption poses a severe chal…

QuantizationVideo SegmentationVideo Semantic Segmentation

THU-Warwick Submission for EPIC-KITCHEN Challenge 2025: Semi-Supervised Video Object Segmentation

2025-06-07 · Mingqi Gao, Haoran Duan, Tianlu Zhang, Jungong Han

In this report, we describe our approach to egocentric video object segmentation. Our method combines large-scale visual pretraining from SAM2 with depth-based geometric cues to handle complex scenes and long-term tracki…

SegmentationSemantic SegmentationSemi-Supervised Video Object SegmentationVideo Object Segmentation+1
1–20 / 905 다음 →