paper-with-me

Papers

Sound Check: Auditing Audio Datasets

2024-10-17 · William Agnew, Julia Barnett, Annie Chu, Rachel Hong, Michael Feffer, Robin Netzorg, Harry H. Jiang, Ezra Awumey, Sauvik Das

Generative audio models are rapidly advancing in both capabilities and public utilization -- several powerful generative audio models have readily available open weights, and some tech companies have released high quality generative audio products. Yet, while prior work has enumerated many ethical issues stemming from the data on which generative visual and textual models have been trained, we have little understanding of similar issues with generative audio datasets, including those related to bias, toxicity, and intellectual property. To bridge this gap, we conducted a literature review of hundreds of audio datasets and selected seven of the most prominent to audit in more detail. We found that these datasets are biased against women, contain toxic stereotypes about marginalized communities, and contain significant amounts of copyrighted work. To enable artists to see if they are in popular audio datasets and facilitate exploration of the contents of these datasets, we developed a web tool audio datasets exploration tool at https://audio-audit.vercel.app.

📄 PDF Abstract BibTeX arXiv:2410.13114

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Effective Pre-Training of Audio Transformers for Sound Event Detection

2024-09-14 · Florian Schmid, Tobias Morocutti, Francesco Foscarin, Jan Schlüter 외

We propose a pre-training pipeline for audio spectrogram transformers for frame-level sound event detection tasks. On top of common pre-training steps, we add a meticulously designed training routine on AudioSet frame-le…

Data AugmentationEvent DetectionKnowledge DistillationSound Event Detection

A Checklist for Trustworthy, Safe, and User-Friendly Mental Health Chatbots

2026-01-21 · Shreya Haran, Samiha Thatikonda, Dong Whi Yoo, Koustuv Saha arxiv

Mental health concerns are rising globally, prompting increased reliance on technology to address the demand-supply gap in mental health services. In particular, mental health chatbots are emerging as a promising solutio…

Benchmarks and leaderboards for sound demixing tasks

2023-05-12 · Roman Solovyev, Alexander Stempkovskiy, Tatiana Habruseva

Music demixing is the task of separating different tracks from the given single audio signal into components, such as drums, bass, and vocals from the rest of the accompaniment. Separation of sources is useful for a rang…

GLAP: General contrastive audio-text pretraining across domains and languages

2025-06-12 · Heinrich Dinkel, Zhiyong Yan, Tianzi Wang, Yongqing Wang 외

Contrastive Language Audio Pretraining (CLAP) is a widely-used method to bridge the gap between audio and text domains. Current CLAP methods enable sound and music retrieval in English, ignoring multilingual spoken conte…

AudioCapsKeyword SpottingRetrievalText Retrieval

No Free Lunch from Audio Pretraining in Bioacoustics: A Benchmark Study of Embeddings

2025-08-13 · Chenggang Chen, Zhiyu Yang arxiv

Bioacoustics, the study of animal sounds, offers a non-invasive method to monitor ecosystems. Extracting embeddings from audio-pretrained deep learning (DL) models without fine-tuning has become popular for obtaining bio…