paper-with-me

홈 › Papers

Representation-Based Data Quality Audits for Audio

2025-09-30 · Alvaro Gonzalez-Jimenez, Fabian Gröger, Linda Wermelinger, Andrin Bürli, Iason Kastanis, Simone Lionetti, Marc Pouly arxiv

Data quality issues such as off-topic samples, near duplicates, and label errors often limit the performance of audio-based systems. This paper addresses these issues by adapting SelfClean, a representation-to-rank data auditing framework, from the image to the audio domain. This approach leverages self-supervised audio representations to identify common data quality issues, creating ranked review lists that surface distinct issues within a single, unified process. The method is benchmarked on the ESC-50, GTZAN, and a proprietary industrial dataset, using both synthetic and naturally occurring corruptions. The results demonstrate that this framework achieves state-of-the-art ranking performance, often outperforming issue-specific baselines and enabling significant annotation savings by efficiently guiding human review.

📄 PDF Abstract BibTeX arXiv:2509.26291

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Benchmarking Single-Factor Physical Video-to-Audio Generation

2026-05-28 · Tingle Li, Siddharth Gururani, Kevin J. Shih, Gantavya Bhatt 외 arxiv

Generative video-to-audio (V2A) models produce highly plausible soundtracks, but it remains unclear whether they capture the underlying physical processes. Existing evaluations emphasize perceptual realism and overlook p…

Audio Generation

Addressing Pitfalls in Auditing Practices of Automatic Speech Recognition Technologies: A Case Study of People with Aphasia

2025-06-10 · Katelyn Xiaoying Mei, Anna Seo Gyeong Choi, Hilke Schellmann, Mona Sloane 외

Automatic Speech Recognition (ASR) has transformed daily tasks from video transcription to workplace hiring. ASR systems' growing use warrants robust and standardized auditing approaches to ensure automated transcription…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Navigatespeech-recognition+1

Who Decides if AI is Fair? The Labels Problem in Algorithmic Auditing

2021-11-16 · Abhilash Mishra, Yash Gorana

Labelled "ground truth" datasets are routinely used to evaluate and audit AI algorithms applied in high-stakes settings. However, there do not exist widely accepted benchmarks for the quality of labels in these datasets.…

Ad Headline Generation using Self-Critical Masked Language Model

2026-07-07 · Yashal Shakti Kanungo, Sumit Negi, Aruna Rajan arxiv

For any E-commerce website it is a nontrivial problem to build enduring advertisements that attract shoppers. It is hard to pass the creative quality bar of the website, especially at a large scale. We thus propose a pro…

Reinforcement LearningHeadline Generation

Ad Headline Generation using Self-Critical Masked Language Model

2021-06-01 · NAACL 2021 4 · Yashal Shakti Kanungo, Sumit Negi, Aruna Rajan

For any E-commerce website it is a nontrivial problem to build enduring advertisements that attract shoppers. It is hard to pass the creative quality bar of the website, especially at a large scale. We thus propose a pro…

Headline GenerationLanguage ModelingLanguage ModellingPolicy Gradient Methods+1