paper-with-me

홈 › Papers

Leveraging Endo- and Exo-Temporal Regularization for Black-box Video Domain Adaptation

2022-08-10 · Yuecong Xu, Jianfei Yang, Haozhi Cao, Min Wu, XiaoLi Li, Lihua Xie, Zhenghua Chen

To enable video models to be applied seamlessly across video tasks in different environments, various Video Unsupervised Domain Adaptation (VUDA) methods have been proposed to improve the robustness and transferability of video models. Despite improvements made in model robustness, these VUDA methods require access to both source data and source model parameters for adaptation, raising serious data privacy and model portability issues. To cope with the above concerns, this paper firstly formulates Black-box Video Domain Adaptation (BVDA) as a more realistic yet challenging scenario where the source video model is provided only as a black-box predictor. While a few methods for Black-box Domain Adaptation (BDA) are proposed in image domain, these methods cannot apply to video domain since video modality has more complicated temporal features that are harder to align. To address BVDA, we propose a novel Endo and eXo-TEmporal Regularized Network (EXTERN) by applying mask-to-mix strategies and video-tailored regularizations: endo-temporal regularization and exo-temporal regularization, performed across both clip and temporal features, while distilling knowledge from the predictions obtained from the black-box predictor. Empirical results demonstrate the state-of-the-art performance of EXTERN across various cross-domain closed-set and partial-set action recognition benchmarks, which even surpassed most existing video domain adaptation methods with source data accessibility.

📄 PDF Abstract BibTeX arXiv:2208.05187

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionDomain AdaptationUnsupervised Domain Adaptation

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

VSTAR: Generative Temporal Nursing for Longer Dynamic Video Synthesis

2024-03-20 · Yumeng Li, William Beluch, Margret Keuper, Dan Zhang 외

Despite tremendous progress in the field of text-to-video (T2V) synthesis, open-sourced T2V diffusion models struggle to generate longer videos with dynamically varying and evolving content. They tend to synthesize quasi…

Generative Temporal NursingText-to-Video GenerationVideo GenerationVideo Synopsis

EndoMamba: An Efficient Foundation Model for Endoscopic Videos via Hierarchical Pre-training

2025-02-26 · Qingyao Tian, Huai Liao, Xinyan Huang, Bingyu Yang 외

Endoscopic video-based tasks, such as visual navigation and surgical phase recognition, play a crucial role in minimally invasive surgeries by providing real-time assistance. While recent video foundation models have sho…

MambaRepresentation LearningState Space ModelsSurgical phase recognition+1

EndoGS: Deformable Endoscopic Tissues Reconstruction with Gaussian Splatting

2024-01-21 · Lingting Zhu, Zhao Wang, Jiahao Cui, Zhenchao Jin 외

Surgical 3D reconstruction is a critical area of research in robotic surgery, with recent works adopting variants of dynamic radiance fields to achieve success in 3D reconstruction of deformable tissues from single-viewp…

3D Reconstruction

Graph Convolution Neural Network For Weakly Supervised Abnormality Localization In Long Capsule Endoscopy Videos

2021-10-18 · Sodiq Adewole, Philip Fernandes, James Jablonski, Andrew Copland 외

Temporal activity localization in long videos is an important problem. The cost of obtaining frame level label for long Wireless Capsule Endoscopy (WCE) videos is prohibitive. In this paper, we propose an end-to-end temp…

Change Point DetectionGraph ClassificationSpecificity

MeshBrush: Painting the Anatomical Mesh with Neural Stylization for Endoscopy

2024-04-03 · John J. Han, Ayberk Acar, Nicholas Kavoussi, Jie Ying Wu

Style transfer is a promising approach to close the sim-to-real gap in medical endoscopy. Rendering synthetic endoscopic videos by traversing pre-operative scans (such as MRI or CT) can generate structurally accurate sim…

Neural StylizationStyle TransferVideo-to-Video Synthesis