paper-with-me

홈 › Papers

OmViD: Omni-supervised active learning for video action detection

2025-08-19 · Aayush Rana, Akash Kumar, Vibhav Vineet, Yogesh S Rawat arxiv

Video action detection requires dense spatio-temporal annotations, which are both challenging and expensive to obtain. However, real-world videos often vary in difficulty and may not require the same level of annotation. This paper analyzes the appropriate annotation types for each sample and their impact on spatio-temporal video action detection. It focuses on two key aspects: 1) how to obtain varying levels of annotation for videos, and 2) how to learn action detection from different annotation types. The study explores video-level tags, points, scribbles, bounding boxes, and pixel-level masks. First, a simple active learning strategy is proposed to estimate the necessary annotation type for each video. Then, a novel spatio-temporal 3D-superpixel approach is introduced to generate pseudo-labels from these annotations, enabling effective training. The approach is validated on UCF101-24 and JHMDB-21 datasets, significantly cutting annotation costs with minimal performance loss.

📄 PDF Abstract BibTeX arXiv:2508.13983

Code (0)

등록된 구현이 없습니다.

Tasks

Action DetectionActive Learning

Similar Papers 제목 키워드 기반

CustomVideoX: 3D Reference Attention Driven Dynamic Adaptation for Zero-Shot Customized Video Diffusion Transformers

2025-02-10 · D. She, Mushui Liu, Jingxuan Pang, Jin Wang 외

Customized generation has achieved significant progress in image synthesis, yet personalized video generation remains challenging due to temporal inconsistencies and quality degradation. In this paper, we introduce Custo…

Image GenerationVideo Generation

Native Active Perception as Reasoning for Omni-Modal Understanding

2026-06-17 · Zhenghao Xing, Ruiyang Xu, Yuxuan Wang, Jinzheng He 외 arxiv

Passive models for long video understanding typically rely on a "watch-it-all" paradigm, processing frames uniformly regardless of query difficulty, causing computational cost to grow with video duration. Although intera…

Reinforcement Learning

Omni-Interactive Universal Embedder

2026-08-27 · Wei-Yao Wang, Kazuya Tateishi, Shuyang Cui, Christian Simon 외 arxiv

Multimodal representation learning has been shifting from traditional two-tower architectures to large language model (LLM)-based embedders due to their strong instruction-following capabilities. Despite this progress, e…

Representation Learning

LoomVideo: Unifying Multimodal Inputs into Video Generation and Editing

2026-06-04 · Jianzong Wu, Hao Lian, Jiongfan Yang, Dachao Hao 외 arxiv

Developing unified video generation and editing models capable of interpreting interleaved multimodal inputs is a promising yet challenging frontier field. Existing unified frameworks predominantly rely on massive models…

Video Generation

OmniMMI: A Comprehensive Multi-modal Interaction Benchmark in Streaming Video Contexts

2025-03-29 · CVPR 2025 1 · Yuxuan Wang, Yueqian Wang, Bo Chen, Tong Wu 외

The rapid advancement of multi-modal language models (MLLMs) like GPT-4o has propelled the development of Omni language models, designed to process and proactively respond to continuous streams of multi-modal data. Despi…

Streaming video understandingVideo Understanding