paper-with-me

Papers

COM Kitchens: An Unedited Overhead-view Video Dataset as a Vision-Language Benchmark

2024-08-05 · Koki Maeda, Tosho Hirasawa, Atsushi Hashimoto, Jun Harashima, Leszek Rybicki, Yusuke Fukasawa, Yoshitaka Ushiku

Procedural video understanding is gaining attention in the vision and language community. Deep learning-based video analysis requires extensive data. Consequently, existing works often use web videos as training resources, making it challenging to query instructional contents from raw video observations. To address this issue, we propose a new dataset, COM Kitchens. The dataset consists of unedited overhead-view videos captured by smartphones, in which participants performed food preparation based on given recipes. Fixed-viewpoint video datasets often lack environmental diversity due to high camera setup costs. We used modern wide-angle smartphone lenses to cover cooking counters from sink to cooktop in an overhead view, capturing activity without in-person assistance. With this setup, we collected a diverse dataset by distributing smartphones to participants. With this dataset, we propose the novel video-to-text retrieval task Online Recipe Retrieval (OnRR) and new video captioning domain Dense Video Captioning on unedited Overhead-View videos (DVC-OV). Our experiments verified the capabilities and limitations of current web-video-based SOTA methods in handling these tasks.

📄 PDF Abstract BibTeX arXiv:2408.02272

Code (1)

omron-sinicx/com_kitchens 공식 구현 pytorch

Tasks

Dense Video CaptioningDiversityRetrievalText RetrievalVideo CaptioningVideo to Text RetrievalVideo Understanding

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음

Similar Papers 제목 키워드 기반

Object Aware Egocentric Online Action Detection

2024-06-03 · Joungbin An, Yunsu Park, Hyolim Kang, Seon Joo Kim

Advancements in egocentric video datasets like Ego4D, EPIC-Kitchens, and Ego-Exo4D have enriched the study of first-person human interactions, which is crucial for applications in augmented reality and assisted living. D…

Action DetectionObjectOnline Action DetectionScene Understanding

Scaling Egocentric Vision: The EPIC-KITCHENS Dataset

2018-04-08 · ECCV 2018 9 · Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler 외

First-person vision is gaining interest as it offers a unique viewpoint on people's interaction with objects, their attention, and even intention. However, progress in this challenging domain has been relatively slow due…

Action Anticipation

Sketch3DVE: Sketch-based 3D-Aware Scene Video Editing

2025-08-19 · Feng-Lin Liu, Shi-Yang Li, Yan-Pei Cao, Hongbo Fu 외 arxiv

Recent video editing methods achieve attractive results in style transfer or appearance modification. However, editing the structural content of 3D scenes in videos remains challenging, particularly when dealing with sig…

Style TransferImage Editing

The EPIC-KITCHENS Dataset: Collection, Challenges and Baselines

2020-04-29 · Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler 외

Since its introduction in 2018, EPIC-KITCHENS has attracted attention as the largest egocentric video benchmark, offering a unique viewpoint on people's interaction with objects, their attention, and even intention. In t…

Object

Temporal Saliency Adaptation in Egocentric Videos

2018-08-28 · Panagiotis Linardos, Eva Mohedano, Monica Cherto, Cathal Gurrin 외

This work adapts a deep neural model for image saliency prediction to the temporal domain of egocentric video. We compute the saliency map for each video frame, firstly with an off-the-shelf model trained from static ima…

Saliency PredictionVideo Saliency Prediction