paper-with-me

Papers

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding

2025-10-30 · Minjoon Jung, Junbin Xiao, Junghyun Kim, Byoung-Tak Zhang, Angela Yao arxiv

Do Video-LLMs have consistent temporal understanding when videos capture the same event from different viewpoints? To study this question, we introduce EgoExo-Con(sistency), a benchmark of synchronized egocentric and exocentric video pairs with human-refined queries that ensure all concepts are visible in both viewpoints. EgoExo-Con emphasizes two temporal understanding tasks: Temporal Verification and Temporal Grounding. It evaluates not only correctness but consistency across viewpoints. Our analysis reveals two critical limitations of existing Video-LLMs: (1) models often fail to maintain consistency, with results far worse than their single-view performances. (2) When naively finetuned with synchronized videos of both viewpoints, the models show improved consistency but often underperform those trained on a single view. For improvements, we propose View-GRPO, a novel reinforcement learning framework that effectively strengthens view-specific temporal reasoning while encouraging consistent comprehension across viewpoints. Our method demonstrates its superior temporal understanding capabilities, especially for improving cross-view consistency. All resources have been made available at https://minjoong507.github.io/projects/EgoExo-Con/

📄 PDF Abstract BibTeX arXiv:2510.26113

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

PRISM: Predictive Recomposition via Semantic Latent Decomposition for View-invariant Video Representation Learning

2026-08-31 · Youngchae Chee, Hosu Lee, Sungjune Park, Junho Kim 외 arxiv

Cross-view video representation learning aims to capture viewpoint-invariant action semantics despite substantial appearance changes across egocentric and exocentric videos. However, existing methods encode each video as…

Representation Learning

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs

2025-07-24 · Yuping He, Yifei Huang, Guo Chen, Baoqi Pei 외 arxiv

Transferring and integrating knowledge across first-person (egocentric) and third-person (exocentric) viewpoints is intrinsic to human intelligence, enabling humans to learn from others and convey insights from their own…

EgoExoMem: Cross-View Memory Reasoning over Synchronized Egocentric and Exocentric Videos

2026-05-18 · Ruiping Liu, Junwei Zheng, Yufan Chen, Di Wen 외 arxiv

Egocentric memory is widely used in embodied intelligence, but it may be insufficient for comprehensive spatial-temporal reasoning. Inspired by human recall from both field and observer perspectives, we introduce EgoExoM…

EgoExo-Fitness: Towards Egocentric and Exocentric Full-Body Action Understanding

2024-06-13 · Yuan-Ming Li, Wei-Jin Huang, An-Lan Wang, Ling-An Zeng 외

We present EgoExo-Fitness, a new full-body action understanding dataset, featuring fitness sequence videos recorded from synchronized egocentric and fixed exocentric (third-person) cameras. Compared with existing full-bo…

Action ClassificationAction LocalizationAction Understanding

EgoExo-Gen: Ego-centric Video Prediction by Watching Exo-centric Videos

2025-04-16 · Jilan Xu, Yifei HUANG, Baoqi Pei, Junlin Hou 외

Generating videos in the first-person perspective has broad application prospects in the field of augmented reality and embodied intelligence. In this work, we explore the cross-view video prediction task, where given an…

PredictionVideo Prediction