paper-with-me

홈 › Papers

MixAssist: An Audio-Language Dataset for Co-Creative AI Assistance in Music Mixing

2025-07-08 · Michael Clemens, Ana Marasović arxiv

While AI presents significant potential for enhancing music mixing and mastering workflows, current research predominantly emphasizes end-to-end automation or generation, often overlooking the collaborative and instructional dimensions vital for co-creative processes. This gap leaves artists, particularly amateurs seeking to develop expertise, underserved. To bridge this, we introduce MixAssist, a novel audio-language dataset capturing the situated, multi-turn dialogue between expert and amateur music producers during collaborative mixing sessions. Comprising 431 audio-grounded conversational turns derived from 7 in-depth sessions involving 12 producers, MixAssist provides a unique resource for training and evaluating audio-language models that can comprehend and respond to the complexities of real-world music production dialogues. Our evaluations, including automated LLM-as-a-judge assessments and human expert comparisons, demonstrate that fine-tuning models such as Qwen-Audio on MixAssist can yield promising results, with Qwen significantly outperforming other tested models in generating helpful, contextually relevant mixing advice. By focusing on co-creative instruction grounded in audio context, MixAssist enables the development of intelligent AI assistants designed to support and augment the creative process in music mixing.

📄 PDF Abstract BibTeX arXiv:2507.06329

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Creative Writing with an AI-Powered Writing Assistant: Perspectives from Professional Writers

2022-11-09 · Daphne Ippolito, Ann Yuan, Andy Coenen, Sehmon Burnam

Recent developments in natural language generation (NLG) using neural language models have brought us closer than ever to the goal of building AI-powered creative writing tools. However, most prior work on human-AI colla…

Text Generation

MIDI-DDSP: Detailed Control of Musical Performance via Hierarchical Modeling

2021-12-17 · ICLR 2022 4 · Yusong Wu, Ethan Manilow, Yi Deng, Rigel Swavely 외

Musical expression requires control of both what notes are played, and how they are performed. Conventional audio synthesizers provide detailed expressive controls, but at the cost of realism. Black-box neural audio synt…

Audio Synthesis

The Song Describer Dataset: a Corpus of Audio Captions for Music-and-Language Evaluation

2023-11-16 · Ilaria Manco, Benno Weck, Seungheon Doh, Minz Won 외

We introduce the Song Describer dataset (SDD), a new crowdsourced corpus of high-quality audio-caption pairs, designed for the evaluation of music-and-language models. The dataset consists of 1.1k human-written natural l…

Music CaptioningMusic GenerationRetrievalText to Audio Retrieval+1

ETTA: Elucidating the Design Space of Text-to-Audio Models

2024-12-26 · Sang-gil Lee, Zhifeng Kong, Arushi Goel, Sungwon Kim 외

Recent years have seen significant progress in Text-To-Audio (TTA) synthesis, enabling users to enrich their creative workflows with synthetic audio generated from natural language prompts. Despite this progress, the eff…

AudioCapsAudio captioningAudio GenerationLanguage Modelling+2

How Does the Disclosure of AI Assistance Affect the Perceptions of Writing?

2024-10-06 · Zhuoyan Li, Chen Liang, Jing Peng, Ming Yin

Recent advances in generative AI technologies like large language models have boosted the incorporation of AI assistance in writing workflows, leading to the rise of a new paradigm of human-AI co-creation in writing. To …