paper-with-me

Speech Prompted Semantic Segmentation

1개 벤치마크 · 논문 4편 · 이 태스크의 논문 보기 →

Benchmarks

ADE20K

결과 8개

Most implemented

Papers

Separating the "Chirp" from the "Chat": Self-supervised Visual Grounding of Sound and Language

2024-06-09 · CVPR 2024 1 · Mark Hamilton, Andrew Zisserman, John R. Hershey, William T. Freeman

We present DenseAV, a novel dual encoder grounding architecture that learns high-resolution, semantically meaningful, and audio-visually aligned features solely through watching videos. We show that DenseAV can discover …

Contrastive LearningCross-Modal RetrievalSemantic SegmentationSound Prompted Semantic Segmentation+2

ImageBind: One Embedding Space To Bind Them All

2023-05-09 · CVPR 2023 1 · Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh 외

We present ImageBind, an approach to learn a joint embedding across six different modalities - images, text, audio, depth, thermal, and IMU data. We show that all combinations of paired data are not necessary to train su…

AllCross-Modal RetrievalMultimodal Deep LearningRetrieval+10

Contrastive Audio-Visual Masked Autoencoder

2022-10-02 · Yuan Gong, Andrew Rouditchenko, Alexander H. Liu, David Harwath 외

In this paper, we first extend the recent Masked Auto-Encoder (MAE) model from a single modality to audio-visual multi-modalities. Subsequently, we propose the Contrastive Audio-Visual Masked Auto-Encoder (CAV-MAE) by co…

Audio ClassificationAudio TaggingContrastive LearningMulti-modal Classification+4

Jointly Discovering Visual Objects and Spoken Words from Raw Sensory Input

2018-04-04 · ECCV 2018 9 · David Harwath, Adrià Recasens, Dídac Surís, Galen Chuang 외

In this paper, we explore neural network models that learn to associate segments of spoken audio captions with the semantically relevant portions of natural images that they refer to. We demonstrate that these audio-visu…

RetrievalSound Prompted Semantic SegmentationSpeech Prompted Semantic Segmentation