paper-with-me

홈 › Papers

DeRA-MOS: Optimizing Text-to-Music Evaluation via Decoupled Listwise Ranking and Modality Alignment

2026-06-08 · Chien-Chun Wang, Hung-Shin Lee, Hsin-Min Wang, Berlin Chen arxiv

Evaluating text-to-music (TTM) systems remains expensive because music impression (MI) and text alignment (TA) scores rely on human mean opinion scores (MOS). Most automatic MOS estimators are trained with point-wise regression or distributional classification. These objectives do not directly optimize rank-based metrics and provide weak geometric constraints for cross-modal coherence. To address these gaps, we propose DeRA-MOS, a decoupled optimization framework for TTM evaluation. For MI, we introduce a batch-aware listwise ranking loss that models relative order within each mini-batch and better aligns with evaluation based on Spearman's rank correlation coefficient (SRCC). For TA, we introduce a score-anchored modality alignment loss that maps human scores to target audio-text similarity and regularizes the latent space before fusion. By effectively mitigating the point-wise training mismatch and modality drift, experiments on MusicEval demonstrate that our decoupled framework yields substantial improvements in both MI and TA ranking metrics, establishing a robust paradigm for large-scale TTM evaluation.

📄 PDF Abstract BibTeX arXiv:2606.10010

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Aligning Text-to-Music Evaluation with Human Preferences

2025-03-20 · Yichen Huang, Zachary Novack, Koichi Saito, Jiatong Shi 외

Despite significant recent advances in generative acoustic text-to-music (TTM) modeling, robust evaluation of these models lags behind, relying in particular on the popular Fr\'echet Audio Distance (FAD). In this work, w…

FAD

YuE: Scaling Open Foundation Models for Long-Form Music Generation

2025-03-11 · Ruibin Yuan, Hanfeng Lin, Shuyue Guo, Ge Zhang 외

We tackle the task of long-form music generation--particularly the challenging \textbf{lyrics-to-song} problem--by introducing YuE, a family of open foundation models based on the LLaMA2 architecture. Specifically, YuE s…

FormIn-Context LearningMusic GenerationStyle Transfer

Text2midi-InferAlign: Improving Symbolic Music Generation with Inference-Time Alignment

2025-05-19 · Abhinaba Roy, Geeta Puri, Dorien Herremans

We present Text2midi-InferAlign, a novel technique for improving symbolic music generation at inference time. Our method leverages text-to-audio alignment and music structural alignment rewards during inference to encour…

Music Generation

Perception of AI-Generated Music -- The Role of Composer Identity, Personality Traits, Music Preferences, and Perceived Humanness

2025-12-02 · David Stammer, Hannah Strauss, Peter Knees arxiv

The rapid rise of AI-generated art has sparked debate about potential biases in how audiences perceive and evaluate such works. This study investigates how composer information and listener characteristics shape the perc…

Caring Trouble and Musical AI: Considerations towards a Feminist Musical AI

2023-11-14 · Kelsey Cotton, Kıvanç Tatar

The ethics of AI as both material and medium for interaction remains in murky waters within the context of musical and artistic practice. The interdisciplinarity of the field is revealing matters of concern and care, whi…

AI AgentEthics