paper-with-me

Papers

Separating the "Chirp" from the "Chat": Self-supervised Visual Grounding of Sound and Language

2024-06-09 · CVPR 2024 1 · Mark Hamilton, Andrew Zisserman, John R. Hershey, William T. Freeman

We present DenseAV, a novel dual encoder grounding architecture that learns high-resolution, semantically meaningful, and audio-visually aligned features solely through watching videos. We show that DenseAV can discover the `meaning'' of words and the location'' of sounds without explicit localization supervision. Furthermore, it automatically discovers and distinguishes between these two types of associations without supervision. We show that DenseAV's localization abilities arise from a new multi-head feature aggregation operator that directly compares dense image and audio representations for contrastive learning. In contrast, many other systems that learn `global'' audio and video representations cannot localize words and sound. Finally, we contribute two new datasets to improve the evaluation of AV representations through speech and sound prompted semantic segmentation. On these and other datasets we show DenseAV dramatically outperforms the prior art on speech and sound prompted semantic segmentation. DenseAV outperforms the previous state-of-the-art, ImageBind, on cross-modal retrieval using fewer than half of the parameters. Project Page: \href{https://aka.ms/denseav}{https://aka.ms/denseav}

📄 PDF Abstract BibTeX arXiv:2406.05629

Code (1)

mhamilton723/DenseAV 공식 구현 pytorch

Tasks

Contrastive LearningCross-Modal RetrievalSemantic SegmentationSound Prompted Semantic SegmentationSpeech Prompted Semantic SegmentationVisual Grounding

Similar Papers 제목 키워드 기반

Classifying Clinical Outcome of Epilepsy Patients with Ictal Chirp Embeddings

2025-08-19 · Nooshin Bahador, Milad Lankarany arxiv

This study presents a pipeline leveraging t-Distributed Stochastic Neighbor Embedding (t-SNE) for interpretable visualizations of chirp features across diverse outcome scenarios. The dataset, comprising chirp-based tempo…

Feature Importance

Neural Generation Meets Real People: Building a Social, Informative Open-Domain Dialogue Agent

2022-07-25 · SIGDIAL (ACL) 2022 9 · Ethan A. Chi, Ashwin Paranjape, Abigail See, Caleb Chiam 외

We present Chirpy Cardinal, an open-domain social chatbot. Aiming to be both informative and conversational, our bot chats with users in an authentic, emotionally intelligent way. By integrating controlled neural generat…

Chatbot

Chirpy3D: Creative Fine-grained 3D Object Fabrication via Part Sampling

2025-01-07 · Kam Woh Ng, Jing Yang, Jia Wei Sii, Jiankang Deng 외

We present Chirpy3D, a novel approach for fine-grained 3D object generation, tackling the challenging task of synthesizing creative 3D objects in a zero-shot setting, with access only to unposed 2D images of seen categor…

3D Generation

Glottal Source Estimation using an Automatic Chirp Decomposition

2020-05-16 · Thomas Drugman, Baris Bozkurt, Thierry Dutoit

In a previous work, we showed that the glottal source can be estimated from speech signals by computing the Zeros of the Z-Transform (ZZT). Decomposition was achieved by separating the roots inside (causal contribution) …

Chirp Complex Cepstrum-based Decomposition for Asynchronous Glottal Analysis

2020-05-10 · Thomas Drugman, Thierry Dutoit

It was recently shown that complex cepstrum can be effectively used for glottal flow estimation by separating the causal and anticausal components of speech. In order to guarantee a correct estimation, some constraints o…