paper-with-me

홈 › Papers

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Videos

2024-11-13 · Sagnik Majumder, Tushar Nagarajan, Ziad Al-Halah, Reina Pradhan, Kristen Grauman

Given a multi-view video, which viewpoint is most informative for a human observer? Existing methods rely on heuristics or expensive "best-view" supervision to answer this question, limiting their applicability. We propose a weakly supervised approach that leverages language accompanying an instructional multi-view video as a means to recover its most informative viewpoint(s). Our key hypothesis is that the more accurately an individual view can predict a view-agnostic text summary, the more informative it is. To put this into action, we propose a framework that uses the relative accuracy of view-dependent caption predictions as a proxy for best view pseudo-labels. Then, those pseudo-labels are used to train a view selector, together with an auxiliary camera pose predictor that enhances view-sensitivity. During inference, our model takes as input only a multi-view video--no language or camera poses--and returns the best viewpoint to watch at each timestep. On two challenging datasets comprised of diverse multi-camera setups and how-to activities, our model consistently outperforms state-of-the-art baselines, both with quantitative metrics and human evaluation.

📄 PDF Abstract BibTeX arXiv:2411.08753

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos

2025-01-01 · CVPR 2025 1 · Sagnik Majumder, Tushar Nagarajan, Ziad Al-Halah, Reina Pradhan 외

Given a multi-view video, which viewpoint is most informative for a human observer? Existing methods rely on heuristics or expensive "best-view" supervision to answer this question, limiting their applicability. We p…

Next-Best-View Estimation based on Deep Reinforcement Learning for Active Object Classification

2021-10-13 · Christian Korbach, Markus D. Solbach, Raphael Memmesheimer, Dietrich Paulus 외

The presentation and analysis of image data from a single viewpoint are often not sufficient to solve a task. Several viewpoints are necessary to obtain more information. The next-best-view problem attempts to find the o…

Deep Reinforcement LearningObjectReinforcement Learning (RL)

Iterative Hybrid Discrete-Continuous Viewpoint Planning for UAV Photogrammetry

2026-08-06 · Alan Grech, Daniel Pisani, Andre Grima, Carl James Debono 외 arxiv

Unmanned aerial vehicle (UAV) photogrammetry requires camera networks that provide sufficient surface coverage, image overlap, parallax, and resolution, yet conventional flight patterns are often poorly adapted to scene …

Pluggable Weakly-Supervised Cross-View Learning for Accurate Vehicle Re-Identification

2021-03-09 · Lu Yang, Hongbang Liu, Jinghao Zhou, Lingqiao Liu 외

Learning cross-view consistent feature representation is the key for accurate vehicle Re-identification (ReID), since the visual appearance of vehicles changes significantly under different viewpoints. To this end, most …

Vehicle Re-Identification

Real-time Active Vision for a Humanoid Soccer Robot Using Deep Reinforcement Learning

2020-11-27 · Soheil Khatibi, Meisam Teimouri, Mahdi Rezaei

In this paper, we present an active vision method using a deep reinforcement learning approach for a humanoid soccer-playing robot. The proposed method adaptively optimises the viewpoint of the robot to acquire the most …

Deep Reinforcement LearningQ-Learningreinforcement-learningReinforcement Learning+1