paper-with-me

홈 › Papers

DDFAD: Dataset Distillation Framework for Audio Data

2024-07-15 · Wenbo Jiang, Rui Zhang, Hongwei Li, Xiaoyuan Liu, Haomiao Yang, Shui Yu

Deep neural networks (DNNs) have achieved significant success in numerous applications. The remarkable performance of DNNs is largely attributed to the availability of massive, high-quality training datasets. However, processing such massive training data requires huge computational and storage resources. Dataset distillation is a promising solution to this problem, offering the capability to compress a large dataset into a smaller distilled dataset. The model trained on the distilled dataset can achieve comparable performance to the model trained on the whole dataset. While dataset distillation has been demonstrated in image data, none have explored dataset distillation for audio data. In this work, for the first time, we propose a Dataset Distillation Framework for Audio Data (DDFAD). Specifically, we first propose the Fused Differential MFCC (FD-MFCC) as extracted features for audio data. After that, the FD-MFCC is distilled through the matching training trajectory distillation method. Finally, we propose an audio signal reconstruction algorithm based on the Griffin-Lim Algorithm to reconstruct the audio signal from the distilled FD-MFCC. Extensive experiments demonstrate the effectiveness of DDFAD on various audio datasets. In addition, we show that DDFAD has promising application prospects in many applications, such as continual learning and neural architecture search.

📄 PDF Abstract BibTeX arXiv:2407.10446

Code (0)

등록된 구현이 없습니다.

Tasks

Continual LearningDataset DistillationNeural Architecture Search

Methods 이 논문이 사용한 방법론

Griffin-Lim Algorithm The Griffin-Lim Algorithm (GLA) is a phase reconstruction method based on the redundancy of the short-time Fourier transform. It promotes the consistency of a spectrogram by…

Similar Papers 제목 키워드 기반

Teaching Audio Models to Reason: A Unified Framework for Source- and Layer-wise Distillation

2025-09-23 · Runyan Yang, Yuke Si, Yingying Gao, Junlan Feng 외 arxiv

While large audio language models excel at tasks like ASR and emotion recognition, they still struggle with complex reasoning due to the modality gap between audio and text as well as the lack of structured intermediate …

Knowledge DistillationEmotion Recognition

Feature-Rich Audio Model Inversion for Data-Free Knowledge Distillation Towards General Sound Classification

2023-03-14 · Zuheng Kang, Yayun He, Jianzong Wang, Junqing Peng 외

Data-Free Knowledge Distillation (DFKD) has recently attracted growing attention in the academic community, especially with major breakthroughs in computer vision. Despite promising results, the technique has not been we…

Data-free Knowledge DistillationKnowledge DistillationSound Classification

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation

2026-06-30 · Binh Mai, Tran Quoc Bao Le, Hung Dinh, Cong Tran arxiv

Diffusion-based text-to-audio (TTA) models achieve impressive synthesis quality but suffer from high inference latency due to iterative multi-step denoising. Existing one-step approaches alleviate this issue but still re…

USAD: Universal Speech and Audio Representation via Distillation

2025-06-23 · Heng-Jui Chang, Saurabhchand Bhati, James Glass, Alexander H. Liu

Self-supervised learning (SSL) has revolutionized audio representations, yet models often remain domain-specific, focusing on either speech or non-speech tasks. In this work, we present Universal Speech and Audio Distill…

Audio TaggingRepresentation LearningSelf-Supervised LearningSound Classification

SightSound-R1: Cross-Modal Reasoning Distillation from Vision to Audio Language Models

2025-09-19 · Qiaolin Wang, Xilin Jiang, Linyang He, Junkai Wu 외 arxiv

While large audio-language models (LALMs) have demonstrated state-of-the-art audio understanding, their reasoning capability in complex soundscapes still falls behind large vision-language models (LVLMs). Compared to the…

Audio-visual Question Answering