paper-with-me

Papers

TB-AVA: Text as a Semantic Bridge for Audio-Visual Parameter Efficient Finetuning

2026-05-12 · Seongah Kim, Dinh Phu Tran, Hyeontaek Hwang, Saad Wazir, Duc Do Minh, Daeyoung Kim arxiv

Audio-visual understanding requires effective alignment between heterogeneous modalities, yet cross-modal correspondence remains challenging when temporally aligned audio and visual signals lack clear semantic correspondence. We propose to use text as a semantic anchor for audio-visual representation learning. To this end, we introduce a parameter-efficient adaptation framework built on frozen audio and visual encoders, centered on Text-Bridged Audio-Visual Adapter (TB-AVA), which enables text-mediated interaction between audio and visual streams. At the core of TB-AVA, Gated Semantic Modulation (GSM) selectively modulates feature channels based on text-inferred semantic relevance. We evaluate the proposed approach on multiple benchmarks, including AVE, AVS, and AVVP, where the proposed framework achieves state-of-the-art performance, demonstrating text as an effective semantic anchor for parameter-efficient fine-tuning (PEFT) in audio-visual learning.

📄 PDF Abstract BibTeX arXiv:2605.11572

Code (0)

등록된 구현이 없습니다.

Tasks

parameter-efficient fine-tuningRepresentation LearningSemantic correspondence

Similar Papers 제목 키워드 기반

TAViS: Text-bridged Audio-Visual Segmentation with Foundation Models

2025-06-13 · Ziyang Luo, Nian Liu, Xuguang Yang, Salman Khan 외

Audio-Visual Segmentation (AVS) faces a fundamental challenge of effectively aligning audio and visual modalities. While recent approaches leverage foundation models to address data scarcity, they often rely on single-mo…

cross-modal alignmentSegmentation

Audio-Visual Semantic Graph Network for Audio-Visual Event Localization

2025-01-01 · CVPR 2025 1 · Liang Liu, Shuaiyong Li, Yongqiang Zhu

Audio-visual event localization (AVEL) aims to identify both the category and temporal boundaries of events that are both audible and visible in unconstrained videos. However, the inherent semantic gap between hetero…

audio-visual event localizationcross-modal alignment

Foley Control: Aligning a Frozen Latent Text-to-Audio Model to Video

2025-10-24 · Ciara Rowles, Varun Jampani, Simon Donné, Shimon Vainer 외 arxiv

Foley Control is a lightweight approach to video-guided Foley that keeps pretrained single-modality models frozen and learns only a small cross-attention bridge between them. We connect V-JEPA2 video embeddings to a froz…

Discovering Sounding Objects by Audio Queries for Audio Visual Segmentation

2023-09-18 · Shaofei Huang, Han Li, Yuqing Wang, Hongji Zhu 외

Audio visual segmentation (AVS) aims to segment the sounding objects for each frame of a given video. To distinguish the sounding objects from silent ones, both audio-visual semantic correspondence and temporal interacti…

ObjectSemantic correspondence

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport

2025-01-16 · Kyeongha Rho, Hyeongkeun Lee, Valentio Iverson, Joon Son Chung

Automated audio captioning is a task that generates textual descriptions for audio content, and recent studies have explored using visual information to enhance captioning quality. However, current methods often fail to …

AudioCapsAudio captioningAudio-Visual CaptioningImage Captioning+3