paper-with-me

Papers

Art2Mus: Bridging Visual Arts and Music through Cross-Modal Generation

2024-10-07 · Ivan Rinaldi, Nicola Fanelli, Giovanna Castellano, Gennaro Vessio

Artificial Intelligence and generative models have revolutionized music creation, with many models leveraging textual or visual prompts for guidance. However, existing image-to-music models are limited to simple images, lacking the capability to generate music from complex digitized artworks. To address this gap, we introduce $\mathcal{A}\textit{rt2}\mathcal{M}\textit{us}$, a novel model designed to create music from digitized artworks or text inputs. $\mathcal{A}\textit{rt2}\mathcal{M}\textit{us}$ extends the AudioLDM~2 architecture, a text-to-audio model, and employs our newly curated datasets, created via ImageBind, which pair digitized artworks with music. Experimental results demonstrate that $\mathcal{A}\textit{rt2}\mathcal{M}\textit{us}$ can generate music that resonates with the input stimuli. These findings suggest promising applications in multimedia art, interactive installations, and AI-driven creative tools.

📄 PDF Abstract BibTeX arXiv:2410.04906

Code (1)

justivanr/art2mus_ 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Mozart's Touch: A Lightweight Multi-modal Music Generation Framework Based on Pre-Trained Large Models

2024-05-05 · Jiajun Li, Tianze Xu, Xuesong Chen, Xinrui Yao 외

In recent years, AI-Generated Content (AIGC) has witnessed rapid advancements, facilitating the creation of music, images, and other artistic forms across a wide range of industries. However, current models for image- an…

DescriptiveLanguage ModelingLanguage ModellingLarge Language Model+2

Bridging Paintings and Music -- Exploring Emotion based Music Generation through Paintings

2024-09-12 · Tanisha Hisariya, huan zhang, Jinhua Liang

Rapid advancements in artificial intelligence have significantly enhanced generative tasks involving music and images, employing both unimodal and multimodal approaches. This research develops a model capable of generati…

FADImage CaptioningMusic Generationtext similarity

Crossing You in Style: Cross-modal Style Transfer from Music to Visual Arts

2020-09-17 · Cheng-Che Lee, Wan-Yi Lin, Yen-Ting Shih, Pei-Yi Patricia Kuo 외

Music-to-visual style transfer is a challenging yet important cross-modal learning problem in the practice of creativity. Its major difference from the traditional image style transfer problem is that the style informati…

Generative Adversarial NetworkStyle Transfer

Art2Mus: Artwork-to-Music Generation via Visual Conditioning and Large-Scale Cross-Modal Alignment

2026-02-19 · Ivan Rinaldi, Matteo Mendula, Nicola Fanelli, Florence Levé 외 arxiv

Music generation has advanced markedly through multimodal deep learning, enabling models to synthesize audio from text and, more recently, from images. However, existing image-conditioned systems suffer from two fundamen…

Multimodal Deep LearningMusic Generation

Vis2Mus: Exploring Multimodal Representation Mapping for Controllable Music Generation

2022-11-10 · Runbang Zhang, Yixiao Zhang, Kai Shao, Ying Shan 외

In this study, we explore the representation mapping from the domain of visual arts to the domain of music, with which we can use visual arts as an effective handle to control music generation. Unlike most studies in mul…

Music GenerationRepresentation LearningStyle Transfer