paper-with-me

홈 › Papers

AuralSAM2: Enabling SAM2 Hear Through Pyramid Audio-Visual Feature Prompting

2025-06-01 · Yuyuan Liu, Yuanhong Chen, Chong Wang, Junlin Han, Junde Wu, Can Peng, Jingkun Chen, Yu Tian, Gustavo Carneiro

Segment Anything Model 2 (SAM2) exhibits strong generalisation for promptable segmentation in video clips; however, its integration with the audio modality remains underexplored. Existing approaches mainly follow two directions: (1) injecting adapters into the image encoder to receive audio signals, which incurs efficiency costs during prompt engineering, and (2) leveraging additional foundation models to generate visual prompts for the sounding objects, which are often imprecisely localised, leading to misguidance in SAM2. Moreover, these methods overlook the rich semantic interplay between hierarchical visual features and other modalities, resulting in suboptimal cross-modal fusion. In this work, we propose AuralSAM2, comprising the novel AuralFuser module, which externally attaches to SAM2 to integrate features from different modalities and generate feature-level prompts, guiding SAM2's decoder in segmenting sounding targets. Such integration is facilitated by a feature pyramid, further refining semantic understanding and enhancing object awareness in multimodal scenarios. Additionally, the audio-guided contrastive learning is introduced to explicitly align audio and visual representations and to also mitigate biases caused by dominant visual patterns. Results on public benchmarks show that our approach achieves remarkable improvements over the previous methods in the field. Code is available at https://github.com/yyliu01/AuralSAM2.

📄 PDF Abstract BibTeX arXiv:2506.01015

Code (1)

yyliu01/auralsam2 공식 구현

Tasks

Contrastive LearningDecoderPrompt Engineering

Methods 이 논문이 사용한 방법론

Contrastive Learning 설명 없음
ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Learning to Hear by Seeing: It's Time for Vision Language Models to Understand Artistic Emotion from Sight and Sound

2025-11-15 · Dengming Zhang, Weitao You, Jingxiong Li, Weishen Lin 외 arxiv

Emotion understanding is critical for making Large Language Models (LLMs) more general, reliable, and aligned with humans. Art conveys emotion through the joint design of visual and auditory elements, yet most prior work…

HEAR: Holistic Evaluation of Audio Representations

2022-03-06 · Joseph Turian, Jordie Shier, Humair Raj Khan, Bhiksha Raj 외

What audio embedding approach generalizes best to a wide range of downstream tasks across a variety of everyday domains without fine-tuning? The aim of the HEAR benchmark is to develop a general-purpose audio representat…

Open-Ended Question Answering

Interpreting Audiograms with Multi-stage Neural Networks

2021-12-17 · Shufan Li, Congxi Lu, Linkai Li, Jirong Duan 외

Audiograms are a particular type of line charts representing individuals' hearing level at various frequencies. They are used by audiologists to diagnose hearing loss, and further select and tune appropriate hearing aids…

Hear-Your-Click: Interactive Object-Specific Video-to-Audio Generation

2025-07-07 · Yingshan Liang, Keyu Fan, Zhicheng Du, Yiran Wang 외 arxiv

Video-to-audio (V2A) generation shows great potential in fields such as film production. Despite significant advances, current V2A methods relying on global video information struggle with complex scenes and generating a…

Data AugmentationAudio Generation

EarCough: Enabling Continuous Subject Cough Event Detection on Hearables

2023-03-18 · Xiyuxing Zhang, Yuntao Wang, Jingru Zhang, Yaqing Yang 외

Cough monitoring can enable new individual pulmonary health applications. Subject cough event detection is the foundation for continuous cough monitoring. Recently, the rapid growth in smart hearables has opened new oppo…

Edge-computingEvent Detection