paper-with-me

Papers

ADIFF: Explaining audio difference using natural language

2025-02-06 · Soham Deshmukh, Shuo Han, Rita Singh, Bhiksha Raj

Understanding and explaining differences between audio recordings is crucial for fields like audio forensics, quality assessment, and audio generation. This involves identifying and describing audio events, acoustic scenes, signal characteristics, and their emotional impact on listeners. This paper stands out as the first work to comprehensively study the task of explaining audio differences and then propose benchmark, baselines for the task. First, we present two new datasets for audio difference explanation derived from the AudioCaps and Clotho audio captioning datasets. Using Large Language Models (LLMs), we generate three levels of difference explanations: (1) concise descriptions of audio events and objects, (2) brief sentences about audio events, acoustic scenes, and signal properties, and (3) comprehensive explanations that include semantics and listener emotions. For the baseline, we use prefix tuning where audio embeddings from two audio files are used to prompt a frozen language model. Our empirical analysis and ablation studies reveal that the naive baseline struggles to distinguish perceptually similar sounds and generate detailed tier 3 explanations. To address these limitations, we propose ADIFF, which introduces a cross-projection module, position captioning, and a three-step training process to enhance the model's ability to produce detailed explanations. We evaluate our model using objective metrics and human evaluation and show our model enhancements lead to significant improvements in performance over naive baseline and SoTA Audio-Language Model (ALM) Qwen Audio. Lastly, we conduct multiple ablation studies to study the effects of cross-projection, language model parameters, position captioning, third stage fine-tuning, and present our findings. Our benchmarks, findings, and strong baseline pave the way for nuanced and human-like explanations of audio differences.

📄 PDF Abstract BibTeX arXiv:2502.04476

Code (1)

soham97/adiff 공식 구현 pytorch

Tasks

AudioCapsAudio captioningAudio GenerationLanguage ModelingLanguage ModellingPosition

Similar Papers 제목 키워드 기반

TaxaDiffusion: Progressively Trained Diffusion Model for Fine-Grained Species Generation

2025-06-02 · Amin Karimi Monsefi, Mridul Khurana, Rajiv Ramnath, Anuj Karpatne 외

We propose TaxaDiffusion, a taxonomy-informed training framework for diffusion models to generate fine-grained animal images with high morphological and identity accuracy. Unlike standard approaches that treat each speci…

Image GenerationTransfer Learning

Remote Sensing Image Super-Resolution for Imbalanced Textures: A Texture-Aware Diffusion Framework

2026-04-15 · Enzhuo Zhang, Sijie Zhao, Dilxat Muhtar, Zhenshi Li 외 arxiv

Generative diffusion priors have recently achieved state-of-the-art performance in natural image super-resolution, demonstrating a powerful capability to synthesize photorealistic details. However, their direct applicati…

Image Super-Resolution

AADiff: Audio-Aligned Video Synthesis with Text-to-Image Diffusion

2023-05-06 · Seungwoo Lee, Chaerin Kong, Donghyeon Jeon, Nojun Kwak

Recent advances in diffusion models have showcased promising results in the text-to-video (T2V) synthesis task. However, as these T2V models solely employ text as the guidance, they tend to struggle in modeling detailed …

TraDiffusion: Trajectory-Based Training-Free Image Generation

2024-08-19 · Mingrui Wu, Oucheng Huang, Jiayi Ji, Jiale Li 외

In this work, we propose a training-free, trajectory-based controllable T2I approach, termed TraDiffusion. This novel method allows users to effortlessly guide image generation via mouse trajectories. To achieve precise …

Image Generation

NatADiff: Adversarial Boundary Guidance for Natural Adversarial Diffusion

2025-05-27 · Max Collins, Jordan Vice, Tim French, Ajmal Mian

Adversarial samples exploit irregularities in the manifold ``learned'' by deep learning models to cause misclassifications. The study of these adversarial samples provides insight into the features a model uses to classi…

Denoising