paper-with-me

홈 › Papers

Gaze-Enabled Egocentric Video Summarization via Constrained Submodular Maximization

2015-06-01 · CVPR 2015 6 · Jia Xu, Lopamudra Mukherjee, Yin Li, Jamieson Warner, James M. Rehg, Vikas Singh

With the proliferation of wearable cameras, the number of videos of users documenting their personal lives using such devices is rapidly increasing. Since such videos may span hours, there is an important need for mechanisms that represent the information content in a compact form (i.e., shorter videos which are more easily browsable/sharable). Motivated by these applications, this paper focuses on the problem of egocentric video summarization. Such videos are usually continuous with significant camera shake and other quality issues. Because of these reasons, there is growing consensus that direct application of standard video summarization tools to such data yields unsatisfactory performance. In this paper, we demonstrate that using gaze tracking information (such as fixation and saccade) significantly helps the summarization task. It allows meaningful comparison of different image frames and enables deriving personalized summaries (gaze provides a sense of the camera wearer's intent). We formulate a summarization model which captures common-sense properties of a good summary, and show that it can be solved as a submodular function maximization with partition matroid constraints, opening the door to a rich body of work from combinatorial optimization. We evaluate our approach on a new gaze-enabled egocentric video dataset (over 15 hours), which will be a valuable standalone resource.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Combinatorial OptimizationCommon Sense ReasoningVideo Summarization

Similar Papers 제목 키워드 기반

Predicting Important Objects for Egocentric Video Summarization

2015-05-18 · Yong Jae Lee, Kristen Grauman

We present a video summarization approach for egocentric or "wearable" camera data. Given hours of video, the proposed method produces a compact storyboard summary of the camera wearer's day. In contrast to traditional k…

Event DetectionVideo Summarization

Temporal Segmentation of Egocentric Videos

2014-06-01 · CVPR 2014 6 · Yair Poleg, Chetan Arora, Shmuel Peleg

The use of wearable cameras makes it possible to record life logging egocentric videos. Browsing such long unstructured videos is time consuming and tedious. Segmentation into meaningful chapters is an important first st…

SegmentationVideo SegmentationVideo Semantic Segmentation

Knowledge Guided Learning: Towards Open Domain Egocentric Action Recognition with Zero Supervision

2020-09-16 · Sathyanarayanan N. Aakur, Sanjoy Kundu, Nikhil Gunti

Advances in deep learning have enabled the development of models that have exhibited a remarkable tendency to recognize and even localize actions in videos. However, they tend to experience errors when faced with scenes …

Action RecognitionDomain AdaptationNovel Object Detectionobject-detection+2

In the Eye of Transformer: Global-Local Correlation for Egocentric Gaze Estimation

2022-08-08 · Bolin Lai, Miao Liu, Fiona Ryan, James M. Rehg

In this paper, we present the first transformer-based model to address the challenging problem of egocentric gaze estimation. We observe that the connection between the global scene context and local visual information i…

Gaze Estimation

In the Eye of MLLM: Benchmarking Egocentric Video Intent Understanding with Gaze-Guided Prompting

2025-09-09 · Taiying Peng, Jiacheng Hua, Miao Liu, Feng Lu arxiv

The emergence of advanced multimodal large language models (MLLMs) has significantly enhanced AI assistants' ability to process complex information across modalities. Recently, egocentric videos, by directly capturing us…

Video Question AnsweringGaze Estimation