paper-with-me

Papers

TriBERT: Human-centric Audio-visual Representation Learning

2021-12-01 · NeurIPS 2021 12 · Tanzila Rahman, Mengyu Yang, Leonid Sigal

The recent success of transformer models in language, such as BERT, has motivated the use of such architectures for multi-modal feature learning and tasks. However, most multi-modal variants (e.g., ViLBERT) have limited themselves to visual-linguistic data. Relatively few have explored its use in audio-visual modalities, and none, to our knowledge, illustrate them in the context of granular audio-visual detection or segmentation tasks such as sound source separation and localization. In this work, we introduce TriBERT -- a transformer-based architecture, inspired by ViLBERT, which enables contextual feature learning across three modalities: vision, pose, and audio, with the use of flexible co-attention. The use of pose keypoints is inspired by recent works that illustrate that such representations can significantly boost performance in many audio-visual scenarios where often one or more persons are responsible for the sound explicitly (e.g., talking) or implicitly (e.g., sound produced as a function of human manipulating an object). From a technical perspective, as part of the TriBERT architecture, we introduce a learned visual tokenization scheme based on spatial attention and leverage weak-supervision to allow granular cross-modal interactions for visual and pose modalities. Further, we supplement learning with sound-source separation loss formulated across all three streams. We pre-train our model on the large MUSIC21 dataset and demonstrate improved performance in audio-visual sound source separation on that dataset as well as other datasets through fine-tuning. In addition, we show that the learned TriBERT representations are generic and significantly improve performance on other audio-visual tasks such as cross-modal audio-visual-pose retrieval by as much as 66.7% in top-1 accuracy.

📄 PDF Abstract BibTeX

Code (1)

ubc-vision/tribert 공식 구현 pytorch

Tasks

Pose RetrievalRepresentation LearningRetrieval

Methods 이 논문이 사용한 방법론

Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Attention Dropout Attention Dropout is a type of dropout used in attention-based architectures, where elements are randomly dropped out of the…
Weight Decay 설명 없음
Adam 설명 없음
Multi-Head Attention 설명 없음
Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Linear Warmup With Linear Decay Linear Warmup With Linear Decay is a learning rate schedule in which we increase the learning rate linearly for $n$ updates and then linearly decay afterwards.

Similar Papers 제목 키워드 기반

TriBERT: Full-body Human-centric Audio-visual Representation Learning for Visual Sound Separation

2021-10-26 · Tanzila Rahman, Mengyu Yang, Leonid Sigal

The recent success of transformer models in language, such as BERT, has motivated the use of such architectures for multi-modal feature learning and tasks. However, most multi-modal variants (e.g., ViLBERT) have limited …

Pose RetrievalRepresentation LearningRetrieval

Egocentric Audio-Visual Object Localization

2023-03-23 · CVPR 2023 1 · Chao Huang, Yapeng Tian, Anurag Kumar, Chenliang Xu

Humans naturally perceive surrounding scenes by unifying sound and sight in a first-person view. Likewise, machines are advanced to approach human intelligence by learning with multisensory inputs from an egocentric pers…

ObjectObject Localization

Sound Bridge: Associating Egocentric and Exocentric Videos via Audio Cues

2025-01-01 · CVPR 2025 1 · Sihong Huang, Jiaxin Wu, XiaoYong Wei, Yi Cai 외

Understanding human behavior and the environmental information in the egocentric video is very challenging due to the invisibility of some actions (e.g., laughing and sneezing) and the local nature of the first-perso…

Action RecognitionScene RecognitionVideo Alignment

AVTrack: Audio-Visual Tracking in Human-centric Complex Scenes

2026-06-01 · Yaoting Wang, Yun Zhou, Zipei Zhang, Henghui Ding arxiv

Audio-visual speaker tracking aims to localize and track active speakers by leveraging auditory and visual cues, enabling fine-grained, human-centric scene understanding. This capability is essential for real-world appli…

Instance SegmentationScene UnderstandingVisual Tracking

Listen to Look into the Future: Audio-Visual Egocentric Gaze Anticipation

2023-05-06 · Bolin Lai, Fiona Ryan, Wenqi Jia, Miao Liu 외

Egocentric gaze anticipation serves as a key building block for the emerging capability of Augmented Reality. Notably, gaze behavior is driven by both visual cues and audio signals during daily activities. Motivated by t…

Representation Learning