paper-with-me

홈 › Papers

Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound

2025-02-07 · Andros Tjandra, Yi-Chiao Wu, Baishan Guo, John Hoffman, Brian Ellis, Apoorv Vyas, Bowen Shi, Sanyuan Chen, Matt Le, Nick Zacharov, Carleigh Wood, Ann Lee, Wei-Ning Hsu

The quantification of audio aesthetics remains a complex challenge in audio processing, primarily due to its subjective nature, which is influenced by human perception and cultural context. Traditional methods often depend on human listeners for evaluation, leading to inconsistencies and high resource demands. This paper addresses the growing need for automated systems capable of predicting audio aesthetics without human intervention. Such systems are crucial for applications like data filtering, pseudo-labeling large datasets, and evaluating generative audio models, especially as these models become more sophisticated. In this work, we introduce a novel approach to audio aesthetic evaluation by proposing new annotation guidelines that decompose human listening perspectives into four distinct axes. We develop and train no-reference, per-item prediction models that offer a more nuanced assessment of audio quality. Our models are evaluated against human mean opinion scores (MOS) and existing methods, demonstrating comparable or superior performance. This research not only advances the field of audio aesthetics but also provides open-source models and datasets to facilitate future work and benchmarking. We release our code and pre-trained model at: https://github.com/facebookresearch/audiobox-aesthetics

📄 PDF Abstract BibTeX arXiv:2502.05139

Code (1)

facebookresearch/audiobox-aesthetics 공식 구현 pytorch

Tasks

Benchmarking

Similar Papers 제목 키워드 기반

Audiobox: Unified Audio Generation with Natural Language Prompts

2023-12-25 · Apoorv Vyas, Bowen Shi, Matthew Le, Andros Tjandra 외

Audio is an essential part of our life, but creating it often requires expertise and is time-consuming. Research communities have made great progress over the past year advancing the performance of large scale audio gene…

AudioCapsAudio GenerationFAD

CoComposer: LLM Multi-agent Collaborative Music Composition

2025-08-29 · Peiwen Xing, Aske Plaat, Niki van Stein arxiv

Existing AI Music composition tools are limited in generation duration, musical quality, and controllability. We introduce CoComposer, a multi-agent system that consists of five collaborating agents, each with a task bas…

QAMRO: Quality-aware Adaptive Margin Ranking Optimization for Human-aligned Assessment of Audio Generation Systems

2025-08-12 · Chien-Chun Wang, Kuan-Tang Huang, Cheng-Yeh Yang, Hung-Shin Lee 외 arxiv

Evaluating audio generation systems, including text-to-music (TTM), text-to-speech (TTS), and text-to-audio (TTA), remains challenging due to the subjective and multi-dimensional nature of human perception. Existing meth…

Audio Generation

DRASP: A Dual-Resolution Attentive Statistics Pooling Framework for Automatic MOS Prediction

2025-08-29 · Cheng-Yeh Yang, Kuan-Tang Huang, Chien-Chun Wang, Hung-Shin Lee 외 arxiv

A pooling mechanism is essential for mean opinion score (MOS) prediction, facilitating the transformation of variable-length audio features into a concise fixed-size representation that effectively encodes speech quality…

Audio Generation

Hierarchical Layout-Aware Graph Convolutional Network for Unified Aesthetics Assessment

2021-06-19 · CVPR 2021 1 · Dongyu She, Yu-Kun Lai, Gaoxiong Yi, Kun Xu

Learning computational models of image aesthetics can have a substantial impact on visual art and graphic design. Although automatic image aesthetics assessment is a challenging topic by its subjective nature, psycho…