paper-with-me

홈 › Papers

We're Not Using Videos Effectively: An Updated Domain Adaptive Video Segmentation Baseline

2024-02-01 · Simar Kareer, Vivek Vijaykumar, Harsh Maheshwari, Prithvijit Chattopadhyay, Judy Hoffman, Viraj Prabhu

There has been abundant work in unsupervised domain adaptation for semantic segmentation (DAS) seeking to adapt a model trained on images from a labeled source domain to an unlabeled target domain. While the vast majority of prior work has studied this as a frame-level Image-DAS problem, a few Video-DAS works have sought to additionally leverage the temporal signal present in adjacent frames. However, Video-DAS works have historically studied a distinct set of benchmarks from Image-DAS, with minimal cross-benchmarking. In this work, we address this gap. Surprisingly, we find that (1) even after carefully controlling for data and model architecture, state-of-the-art Image-DAS methods (HRDA and HRDA+MIC) outperform Video-DAS methods on established Video-DAS benchmarks (+14.5 mIoU on Viper$\rightarrow$CityscapesSeq, +19.0 mIoU on Synthia$\rightarrow$CityscapesSeq), and (2) naive combinations of Image-DAS and Video-DAS techniques only lead to marginal improvements across datasets. To avoid siloed progress between Image-DAS and Video-DAS, we open-source our codebase with support for a comprehensive set of Video-DAS and Image-DAS methods on a common benchmark. Code available at https://github.com/SimarKareer/UnifiedVideoDA

📄 PDF Abstract BibTeX arXiv:2402.00868

Code (1)

simarkareer/unifiedvideoda 공식 구현 pytorch

Tasks

BenchmarkingDomain AdaptationSemantic SegmentationUnsupervised Domain AdaptationVideo SegmentationVideo Semantic Segmentation

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

One to Many: Adaptive Instrument Segmentation via Meta Learning and Dynamic Online Adaptation in Robotic Surgical Video

2021-03-24 · Zixu Zhao, Yueming Jin, Bo Lu, Chi-Fai Ng 외

Surgical instrument segmentation in robot-assisted surgery (RAS) - especially that using learning-based models - relies on the assumption that training and testing videos are sampled from the same domain. However, it is …

General KnowledgeMeta-Learning

Domain Adaptive Video Segmentation via Temporal Consistency Regularization

2021-07-23 · ICCV 2021 10 · Dayan Guan, Jiaxing Huang, Aoran Xiao, Shijian Lu

Video semantic segmentation is an essential task for the analysis and understanding of videos. Recent efforts largely focus on supervised video segmentation by learning from fully annotated data, but the learnt models of…

SegmentationUnsupervised Domain AdaptationVideo Semantic Segmentation

Selective Structured State-Spaces for Long-Form Video Understanding

2023-03-25 · CVPR 2023 1 · Jue Wang, Wentao Zhu, Pichao Wang, Xiang Yu 외

Effective modeling of complex spatiotemporal dependencies in long-form videos remains an open problem. The recently proposed Structured State-Space Sequence (S4) model with its linear complexity offers a promising direct…

Contrastive LearningFormToken ReductionVideo Classification+1

LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

2024-10-22 · Xiaoqian Shen, Yunyang Xiong, Changsheng Zhao, Lemeng Wu 외

Multimodal Large Language Models (MLLMs) have shown promising progress in understanding and analyzing video content. However, processing long videos remains a significant challenge constrained by LLM's context size. To a…

Token ReductionVideo Question AnsweringVideo UnderstandingZero-Shot Video Question Answer

NeuSaver: Neural Adaptive Power Consumption Optimization for Mobile Video Streaming

2021-07-15 · Kyoungjun Park, Myungchul Kim, Laihyuk Park

Video streaming services strive to support high-quality videos at higher resolutions and frame rates to improve the quality of experience (QoE). However, high-quality videos consume considerable amounts of energy on mobi…

Reinforcement Learning (RL)