Sound Prompted Semantic Segmentation 벤치마크
Sound Prompted Semantic Segmentation on ADE20K
mAP
- 2018-04-04 — DAVENet: mAP 16.8
- 2022-10-02 — CAVMAE: mAP 26.0
- 2024-06-09 — DenseAV: mAP 32.7
| Rank | Model | mAP | mIoU | Paper | Code | Year |
|---|---|---|---|---|---|---|
| 1 | DenseAV | 32.7 | 24.7 | Separating the "Chirp" from the "Chat": Self-supervised Visual Grounding of Sound and Language | mhamilton723/DenseAV | 2024 |
| 2 | CAVMAE | 26.0 | 17.0 | Contrastive Audio-Visual Masked Autoencoder | yuangongnd/cav-mae | 2022 |
| 3 | ImageBIND | 19.7 | 20.5 | ImageBind: One Embedding Space To Bind Them All | facebookresearch/imagebind · klemens-floege/oneprot · ginihumer/amumo | 2023 |
| 4 | DAVENet | 16.8 | 18.1 | Jointly Discovering Visual Objects and Spoken Words from Raw Sensory Input | 2018 | |
| 5 | DenseAV | 32.7 | 24.7 | Separating the "Chirp" from the "Chat": Self-supervised Visual Grounding of Sound and Language | mhamilton723/DenseAV | 2024 |
| 6 | CAVMAE | 26.0 | 17.0 | Contrastive Audio-Visual Masked Autoencoder | yuangongnd/cav-mae | 2022 |
| 7 | ImageBIND | 19.7 | 20.5 | ImageBind: One Embedding Space To Bind Them All | facebookresearch/imagebind · klemens-floege/oneprot · ginihumer/amumo | 2023 |
| 8 | DAVENet | 16.8 | 18.1 | Jointly Discovering Visual Objects and Spoken Words from Raw Sensory Input | 2018 |