| Rank | Model |
Test mAP | AUC | d-prime |
Extra Training Data |
Paper | Code | Year |
| 20 |
M2D-CLAP/0.7 |
0.485 | – | – |
|
M2D-CLAP: Masked Modeling Duo Meets CLAP for Learning General-purpose Audio-Language Representation
|
nttcslab/m2d · nttcslab/eval-audio-repr |
2024 |
| 20 |
M2D-AS/0.7 |
0.485 | – | – |
|
Masked Modeling Duo: Towards a Universal Audio Pre-training Framework
|
nttcslab/m2d · nttcslab/eval-audio-repr |
2024 |
| 23 |
MAViL (Audio-only, single) |
0.484 | – | – |
✓ |
|
|
|
| 24 |
mn40_as (Single) |
0.483 | – | – |
✓ |
Efficient Large-scale Audio Tagging via Transformer-to-CNN Knowledge Distillation
|
fschmid56/efficientat · fschmid56/efficientat_hear |
2022 |
| 25 |
MAX-AST (Single) |
0.481 | – | – |
|
MAX-AST: COMBINING CONVOLUTION, LOCAL AND GLOBAL SELF-ATTENTIONS FOR AUDIO EVENT CLASSIFICATION
|
ta012/MaxAST |
2024 |
| 26 |
ATST-Frame |
0.480 | – | – |
|
Self-supervised Audio Teacher-Student Transformer for Both Clip-level and Frame-level Tasks
|
Audio-WestlakeU/ATST-SED · audio-westlakeu/audiossl |
2023 |
| 27 |
M2D/0.7 |
0.479 | – | – |
|
Masked Modeling Duo: Towards a Universal Audio Pre-training Framework
|
nttcslab/m2d · nttcslab/eval-audio-repr |
2024 |
| 28 |
PlayItBackX3 |
0.477 | – | – |
|
Play It Back: Iterative Attention for Audio Recognition
|
alexandrosstergiou/PlayItBack |
2022 |
| 29 |
DASS-Medium (Audio-only, single) |
0.476 | – | – |
|
DASS: Distilled Audio State Space Models Are Stronger and More Duration-Scalable Learners
|
Saurabhbhati/DASS |
2024 |
| 30 |
PSLA (Ensemble) |
0.474 | 0.981 | 2.936 |
✓ |
PSLA: Improving Audio Tagging with Pretraining, Sampling, Labeling, and Aggregation
|
YuanGongND/psla |
2021 |
| 31 |
DASS-Small (Audio-only, single) |
0.472 | – | – |
|
DASS: Distilled Audio State Space Models Are Stronger and More Duration-Scalable Learners
|
Saurabhbhati/DASS |
2024 |
| 32 |
PaSST-S (Single) |
0.471 | – | – |
✓ |
Efficient Training of Audio Transformers with Patchout
|
kkoutini/passt · kkoutini/passt_hear21 |
2021 |
| 32 |
MaskSpec (AS-2M) |
0.471 | – | – |
|
|
|
|
| 34 |
CAV-MAE (Audio-Only) |
0.466 | – | – |
✓ |
Contrastive Audio-Visual Masked Autoencoder
|
yuangongnd/cav-mae |
2022 |
| 34 |
Audiovisual Masked Autoencoder (Audio-only, Single) |
0.466 | – | – |
|
Audiovisual Masked Autoencoders
|
google-research/scenic · google-research/scenic |
2022 |
| 36 |
AudioVisual Fusion Net |
0.462 | 0.975 | – |
|
Large Scale Audiovisual Learning of Sounds with Weakly Labeled Data
|
|
2020 |
| 37 |
AST (Single) |
0.459 | – | – |
✓ |
AST: Audio Spectrogram Transformer
|
YuanGongND/ast · nttcslab/composing-general-audio-repr · pxaris/ccml
· +2 |
2021 |
| 38 |
ERANN-1-6 |
0.450 | 0.976 | 2.804 |
|
ERANNs: Efficient Residual Audio Neural Networks for Audio Pattern Recognition
|
|
2021 |
| 39 |
Perceiver |
0.449 | – | – |
|
Perceiver: General Perception with Iterative Attention
|
deepmind/deepmind-research · towhee-io/towhee · lucidrains/perceiver-pytorch
· +9 |
2021 |
| 40 |
PSLA (Single) |
0.443 | 0.975 | 2.778 |
✓ |
PSLA: Improving Audio Tagging with Pretraining, Sampling, Labeling, and Aggregation
|
YuanGongND/psla |
2021 |