paper-with-me

Papers

EPIC-Fusion: Audio-Visual Temporal Binding for Egocentric Action Recognition

2019-08-22 · ICCV 2019 10 · Evangelos Kazakos, Arsha Nagrani, Andrew Zisserman, Dima Damen

We focus on multi-modal fusion for egocentric action recognition, and propose a novel architecture for multi-modal temporal-binding, i.e. the combination of modalities within a range of temporal offsets. We train the architecture with three modalities -- RGB, Flow and Audio -- and combine them with mid-level fusion alongside sparse temporal sampling of fused representations. In contrast with previous works, modalities are fused before temporal aggregation, with shared modality and fusion weights over time. Our proposed architecture is trained end-to-end, outperforming individual modalities as well as late-fusion of modalities. We demonstrate the importance of audio in egocentric vision, on per-class basis, for identifying actions as well as interacting objects. Our method achieves state of the art results on both the seen and unseen test sets of the largest egocentric dataset: EPIC-Kitchens, on all metrics using the public leaderboard.

📄 PDF Abstract BibTeX arXiv:1908.08498

Code (1)

ekazakos/temporal-binding-network 공식 구현 tf

Tasks

Action RecognitionEgocentric Activity Recognition

Similar Papers 제목 키워드 기반

Dynamique temporelle du liage dans la fusion de la parole audiovisuelle (Temporal dynamics of binding in audiovisual speech fusion) [in French]

2012-06-01 · JEPTALNRECITAL 2012 6 · Olha Nahorna, Fr{\'e}d{\'e}ric Berthommier, Jean-Luc Schwartz

Epic-Sounds: A Large-scale Dataset of Actions That Sound

2023-02-01 · Jaesung Huh, Jacob Chalk, Evangelos Kazakos, Dima Damen 외

We introduce Epic-Sounds, a large-scale dataset of audio annotations capturing temporal extents and class labels within the audio stream of the egocentric videos. We propose an annotation pipeline where annotators tempor…

Action RecognitionSound Classification

Seeing and Hearing Egocentric Actions: How Much Can We Learn?

2019-10-15 · Alejandro Cartas, Jordi Luque, Petia Radeva, Carlos Segura 외

Our interaction with the world is an inherently multimodal experience. However, the understanding of human-to-object interactions has historically been addressed focusing on a single modality. In particular, a limited nu…

Action Recognition

TIM: A Time Interval Machine for Audio-Visual Action Recognition

2024-04-08 · CVPR 2024 1 · Jacob Chalk, Jaesung Huh, Evangelos Kazakos, Andrew Zisserman 외

Diverse actions give rise to rich audio-visual signals in long videos. Recent works showcase that the two modalities of audio and video exhibit different temporal extents of events and distinct labels. We address the int…

Action DetectionAction Recognition

Temporal and Cross-Modal Alignment for Enhanced Audiovisual Video Captioning

2026-07-02 · Chen Zhao, Jiajun Ma, Qilong Huang, Tiehan Fan 외 arxiv

While Multimodal Large Language Models (MLLMs) have advanced video understanding, achieving precise temporal and cross-modal alignment in audiovisual video captioning remains a formidable challenge. Most existing approac…

Relational ReasoningVideo Captioning