paper-with-me

홈 › Papers

Multi-View Video-Based Learning: Leveraging Weak Labels for Frame-Level Perception

2024-03-18 · Vijay John, Yasutomo Kawanishi

For training a video-based action recognition model that accepts multi-view video, annotating frame-level labels is tedious and difficult. However, it is relatively easy to annotate sequence-level labels. This kind of coarse annotations are called as weak labels. However, training a multi-view video-based action recognition model with weak labels for frame-level perception is challenging. In this paper, we propose a novel learning framework, where the weak labels are first used to train a multi-view video-based base model, which is subsequently used for downstream frame-level perception tasks. The base model is trained to obtain individual latent embeddings for each view in the multi-view input. For training the model using the weak labels, we propose a novel latent loss function. We also propose a model that uses the view-specific latent embeddings for downstream frame-level action recognition and detection tasks. The proposed framework is evaluated using the MM Office dataset by comparing several baseline algorithms. The results show that the proposed base model is effectively trained using weak labels and the latent embeddings help the downstream models improve accuracy.

📄 PDF Abstract BibTeX arXiv:2403.11616

Code (0)

등록된 구현이 없습니다.

Tasks

Action Recognition

Methods 이 논문이 사용한 방법론

BASE 설명 없음

Similar Papers 제목 키워드 기반

Action Selection Learning for Multi-label Multi-view Action Recognition

2024-10-04 · Trung Thanh Nguyen, Yasutomo Kawanishi, Takahiro Komamizu, Ichiro Ide

Multi-label multi-view action recognition aims to recognize multiple concurrent or sequential actions from untrimmed videos captured by multiple cameras. Existing work has focused on multi-view action recognition in a na…

Action Recognition

Graph Convolution Neural Network For Weakly Supervised Abnormality Localization In Long Capsule Endoscopy Videos

2021-10-18 · Sodiq Adewole, Philip Fernandes, James Jablonski, Andrew Copland 외

Temporal activity localization in long videos is an important problem. The cost of obtaining frame level label for long Wireless Capsule Endoscopy (WCE) videos is prohibitive. In this paper, we propose an end-to-end temp…

Change Point DetectionGraph ClassificationSpecificity

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Videos

2024-11-13 · Sagnik Majumder, Tushar Nagarajan, Ziad Al-Halah, Reina Pradhan 외

Given a multi-view video, which viewpoint is most informative for a human observer? Existing methods rely on heuristics or expensive "best-view" supervision to answer this question, limiting their applicability. We propo…

Weakly Supervised Person Re-Identification

2019-04-08 · CVPR 2019 6 · Jingke Meng, Sheng Wu, Wei-Shi Zheng

In the conventional person re-id setting, it is assumed that the labeled images are the person images within the bounding box for each individual; this labeling across multiple nonoverlapping camera views from raw video …

Multi-Label LearningPerson Re-Identification

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos

2025-01-01 · CVPR 2025 1 · Sagnik Majumder, Tushar Nagarajan, Ziad Al-Halah, Reina Pradhan 외

Given a multi-view video, which viewpoint is most informative for a human observer? Existing methods rely on heuristics or expensive "best-view" supervision to answer this question, limiting their applicability. We p…