paper-with-me

홈 › Papers

UNQA: Unified No-Reference Quality Assessment for Audio, Image, Video, and Audio-Visual Content

2024-07-29 · Yuqin Cao, Xiongkuo Min, Yixuan Gao, Wei Sun, Weisi Lin, Guangtao Zhai

As multimedia data flourishes on the Internet, quality assessment (QA) of multimedia data becomes paramount for digital media applications. Since multimedia data includes multiple modalities including audio, image, video, and audio-visual (A/V) content, researchers have developed a range of QA methods to evaluate the quality of different modality data. While they exclusively focus on addressing the single modality QA issues, a unified QA model that can handle diverse media across multiple modalities is still missing, whereas the latter can better resemble human perception behaviour and also have a wider range of applications. In this paper, we propose the Unified No-reference Quality Assessment model (UNQA) for audio, image, video, and A/V content, which tries to train a single QA model across different media modalities. To tackle the issue of inconsistent quality scales among different QA databases, we develop a multi-modality strategy to jointly train UNQA on multiple QA databases. Based on the input modality, UNQA selectively extracts the spatial features, motion features, and audio features, and calculates a final quality score via the four corresponding modality regression modules. Compared with existing QA methods, UNQA has two advantages: 1) the multi-modality training strategy makes the QA model learn more general and robust quality-aware feature representation as evidenced by the superior performance of UNQA compared to state-of-the-art QA methods. 2) UNQA reduces the number of models required to assess multimedia data across different modalities. and is friendly to deploy to practical applications.

📄 PDF Abstract BibTeX arXiv:2407.19704

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

AudioJudge: Understanding What Works in Large Audio Model Based Speech Evaluation

2025-07-17 · Potsawee Manakul, Woody Haosheng Gan, Michael J. Ryan, Ali Sartaz Khan 외 arxiv

Current speech evaluation suffers from two critical limitations: the need and difficulty of designing specialized systems targeting individual audio characteristics, and poor correlation between automatic evaluation meth…

Speaker IdentificationPrompt Engineering

A novel fuzzy logic-based metric for audio quality assessment: Objective audio quality assessment

2019-10-29

ITU-R BS.1387 states a method for objective assessment of perceived audio quality. This Recommendation, known also as PEAQ (Perceptual Evaluation of Audio Quality) is based on a psychoacoustic model of the human ear and …

Audio Quality Assessment

Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound

2025-02-07 · Andros Tjandra, Yi-Chiao Wu, Baishan Guo, John Hoffman 외

The quantification of audio aesthetics remains a complex challenge in audio processing, primarily due to its subjective nature, which is influenced by human perception and cultural context. Traditional methods often depe…

Benchmarking

Perceiving Music Quality with GANs

2020-06-11 · Agrin Hilmkil, Carl Thomé, Anders Arpteg

Several methods have been developed to assess the perceptual quality of audio under transforms like lossy compression. However, they require paired reference signals of the unaltered content, limiting their use in applic…

Audio GenerationAudio Quality AssessmentMusic GenerationStyle Transfer

MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation

2026-07-15 · Xiaohan Zhang, Yuqing Wen, Junlin Chen, Yuqi Tang 외 hf

Multi-reference-to-audio-video (MR2AV) generation aims to generate coherent audio-video content conditioned on multiple references and textual instructions. Existing benchmarks mainly focus on text-driven generation, sin…

Instruction FollowingVideo GenerationVideo Alignment