How Much Does Audio Matter to Recognize Egocentric Object Interactions?
Sounds are an important source of information on our daily interactions with objects. For instance, a significant amount of people can discern the temperature of water that it is being poured just by using the sense of hearing. However, only a few works have explored the use of audio for the classification of object interactions in conjunction with vision or as single modality. In this preliminary work, we propose an audio model for egocentric action recognition and explore its usefulness on the parts of the problem (noun, verb, and action classification). Our model achieves a competitive result in terms of verb classification (34.26% accuracy) on a standard benchmark with respect to vision-based state of the art systems, using a comparatively lighter architecture.
Code (0)
등록된 구현이 없습니다.
Tasks
Action ClassificationAction RecognitionClassificationGeneral ClassificationSimilar Papers 제목 키워드 기반
SoundingActions: Learning How Actions Sound from Narrated Egocentric Videos
We propose a novel self-supervised embedding to learn how actions sound from narrated in-the-wild egocentric videos. Whereas existing methods rely on curated data with known audio-visual correspondence, our multimodal co…
Seeing and Hearing Egocentric Actions: How Much Can We Learn?
Our interaction with the world is an inherently multimodal experience. However, the understanding of human-to-object interactions has historically been addressed focusing on a single modality. In particular, a limited nu…
Action RecognitionIs Sharing of Egocentric Video Giving Away Your Biometric Signature?
Easy availability of wearable egocentric cameras, and the sense of privacy propagated by the fact that the wearer is never seen in the captured videos, has led to a tremendous rise in public sharing of such videos. Unlik…
Optical Flow EstimationTrajectory Aligned Features For First Person Action Recognition
Egocentric videos are characterised by their ability to have the first person view. With the popularity of Google Glass and GoPro, use of egocentric videos is on the rise. Recognizing action of the wearer from egocentric…
Action RecognitionPoint TrackingTemporal Action LocalizationActive Audio-Visual Separation of Dynamic Sound Sources
We explore active audio-visual separation for dynamic sound sources, where an embodied agent moves intelligently in a 3D environment to continuously isolate the time-varying audio stream being emitted by an object of int…