paper-with-me

Papers

Visually Exploring Multi-Purpose Audio Data

2021-10-09 · David Heise, Helen L. Bear

We analyse multi-purpose audio using tools to visualise similarities within the data that may be observed via unsupervised methods. The success of machine learning classifiers is affected by the information contained within system inputs, so we investigate whether latent patterns within the data may explain performance limitations of such classifiers. We use the visual assessment of cluster tendency (VAT) technique on a well known data set to observe how the samples naturally cluster, and we make comparisons to the labels used for audio geotagging and acoustic scene classification. We demonstrate that VAT helps to explain and corroborate confusions observed in prior work to classify this audio, yielding greater insight into the performance - and limitations - of supervised classification systems. While this exploratory analysis is conducted on data for which we know the "ground truth" labels, this method of visualising the natural groupings as dictated by the data leads to important questions about unlabelled data that can help the evaluation and realistic expectations of future (including self-supervised) classification systems.

📄 PDF Abstract BibTeX arXiv:2110.04584

Code (0)

등록된 구현이 없습니다.

Tasks

Acoustic Scene ClassificationClassificationScene Classification

Similar Papers 제목 키워드 기반

Exploring Federated Self-Supervised Learning for General Purpose Audio Understanding

2024-02-05 · Yasar Abbas Ur Rehman, Kin Wai Lau, Yuyang Xie, Lan Ma 외

The integration of Federated Learning (FL) and Self-supervised Learning (SSL) offers a unique and synergetic combination to exploit the audio data for general-purpose audio understanding, without compromising user data p…

Federated LearningRetrievalSelf-Supervised Learning

M2D2: Exploring General-purpose Audio-Language Representations Beyond CLAP

2025-03-28 · Daisuke Niizumi, Daiki Takeuchi, Masahiro Yasuda, Binh Thien Nguyen 외

Contrastive language-audio pre-training (CLAP) has addressed audio-language tasks such as audio-text retrieval by aligning audio and text in a common feature space. While CLAP addresses general audio-language tasks, its …

Audio captioningAudio ClassificationAudio TaggingAudio to Text Retrieval+13

Exploring Pre-trained General-purpose Audio Representations for Heart Murmur Detection

2024-04-26 · Daisuke Niizumi, Daiki Takeuchi, Yasunori Ohishi, Noboru Harada 외

To reduce the need for skilled clinicians in heart sound interpretation, recent studies on automating cardiac auscultation have explored deep learning approaches. However, despite the demands for large data for deep lear…

Classify murmursGPUSelf-Supervised LearningTransfer Learning

Exploring Efficient-Tuned Learning Audio Representation Method from BriVL

2023-03-08 · Sen Fang, Yangjian Wu, Bowen Gao, Jingwen Cai 외

Recently, researchers have gradually realized that in some cases, the self-supervised pre-training on large-scale Internet data is better than that of high-quality/manually labeled data sets, and multimodal/large models …

Image GenerationRepresentation Learning

A virtual reality-based method for examining audiovisual prosody perception

2022-09-13 · Hartmut Meister, Isa Samira Winter, Moritz Waeachtler, Pascale Sandmann 외

Prosody plays a vital role in verbal communication. Acoustic cues of prosody have been examined extensively. However, prosodic characteristics are not only perceived auditorily, but also visually based on head and facial…