paper-with-me

홈 › Papers

FATE: Frame-Level Audio-Visual Temporal Embedding

2026-08-02 · Kaisi Guan, Bingzi Zhang, Xihua Wang, Ying Ba, Xin Cheng, Yijing Chen, Ruihua Song hf

When a dog opens its mouth and barks, humans naturally recognize what the sound is and when it occurs. Building audio-visual models with this same ability requires representations that capture both semantic and temporal alignment. Current approaches fall short on one side or the other: embedding models match semantic but lose temporal information; synchronization models capture temporal offsets but lack semantic understanding. To bridge this gap, we propose FATE, Frame-level Audio-visual Temporal Embedding. Unlike prior embedding models that pool each modality into a single embedding and discard temporal information, FATE retains frame-level sequences, aligns them on the physical timeline, and computes similarity over strictly aligned frame pairs. Unlike synchronization models that output only an offset prediction, FATE encodes synchronization in a reusable embedding space, trained with a joint objective combining cross-video semantic and within-video temporal contrastive learning to capture both what sounds and when it occurs. Across three tasks, FATE surpasses the strongest baseline on temporal and semantic retrieval by a large margin, matches fully supervised methods on event localization in a zero-shot setting, and achieves the best correlation with human judgments as a generation evaluation metric. The source code can be found at https://github.com/guankaisi/FATE.

📄 PDF Abstract BibTeX arXiv:2608.01310

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive LearningSemantic Retrieval

Similar Papers 제목 키워드 기반

What's Making That Sound Right Now? Video-centric Audio-Visual Localization

2025-07-07 · Hahyeon Choi, Junhoo Lee, Nojun Kwak arxiv

Audio-Visual Localization (AVL) aims to identify sound-emitting sources within a visual scene. However, existing studies focus on image-level audio-visual associations, failing to capture temporal dynamics. Moreover, the…

Visual Localization

Extending Segment Anything Model into Auditory and Temporal Dimensions for Audio-Visual Segmentation

2024-06-10 · Juhyeong Seon, Woobin Im, Sebin Lee, Jumin Lee 외

Audio-visual segmentation (AVS) aims to segment sound sources in the video sequence, requiring a pixel-level understanding of audio-visual correspondence. As the Segment Anything Model (SAM) has strongly impacted extensi…

Decoder

Discovering Sounding Objects by Audio Queries for Audio Visual Segmentation

2023-09-18 · Shaofei Huang, Han Li, Yuqing Wang, Hongji Zhu 외

Audio visual segmentation (AVS) aims to segment the sounding objects for each frame of a given video. To distinguish the sounding objects from silent ones, both audio-visual semantic correspondence and temporal interacti…

ObjectSemantic correspondence

Understanding Cell Fate Decisions with Temporal Attention

2026-03-17 · Florian Bürger, Martim Dias Gomes, Adrián E. Granada, Noémie Moreau 외 arxiv

Understanding non-genetic determinants of cell fate is critical for developing and improving cancer therapies, as genetically identical cells can exhibit divergent outcomes under the same treatment conditions. In this wo…

CATR: Combinatorial-Dependence Audio-Queried Transformer for Audio-Visual Video Segmentation

2023-09-18 · Kexin Li, Zongxin Yang, Lei Chen, Yi Yang 외

Audio-visual video segmentation~(AVVS) aims to generate pixel-level maps of sound-producing objects within image frames and ensure the maps faithfully adhere to the given audio, such as identifying and segmenting a singi…

Video SegmentationVideo Semantic Segmentation